A robotic hybrid visual servoing positioning method, apparatus and device
By combining a camera and a laser sensor in the Jacobian matrix fusion method at the robot end effector, the problems of feature point extraction error on complex curved surfaces and low efficiency of deep learning in traditional visual servoing are solved. This achieves high-precision recognition and positioning of pre-assembled holes, improving the robot's assembly efficiency and success rate.
Patent Information
- Application Number
- CN202511526520.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2045-10-24
AI Technical Summary
Traditional visual servoing methods suffer from large feature point extraction errors and low efficiency in deep reinforcement learning when dealing with complex curved surface structures, resulting in unstable robot servo control and difficulty in achieving high-precision and highly adaptable recognition and positioning of pre-assembled holes.
The robot employs a camera and multiple laser sensors mounted on its end effector. By combining image processing and deep reinforcement learning, visual and laser data are fused using a Jacobi matrix to construct a joint feature Jacobi matrix, which is then input into a pre-trained decision model for real-time control.
It significantly improves the target detection accuracy and attitude recognition robustness of robots on complex curved surfaces, achieves high-precision position and attitude adjustment, enhances the adaptability and task success rate of robot systems, and meets the needs of high-requirement assembly tasks on complex curved surfaces.
Smart Images

Figure CN121004614B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of robots, in particular to a robot hybrid visual servo positioning method, device and equipment. BACKGROUND
[0002] Robots play an increasingly important role in the field of high-end equipment manufacturing, and robot technology has become a key force to improve production efficiency and ensure assembly precision. In manufacturing, due to the large size, small rigidity and high processing difficulty of structural parts, how to quickly and accurately complete the recognition and positioning of pre-assembly holes has become one of the key bottlenecks for manufacturing efficiency and quality improvement.
[0003] Now it has been tried to use a monocular vision-based servo positioning method, to collect images through a camera, extract image features of pre-assembly holes on a workpiece surface, and combine image processing and servo control algorithms to drive the robot end effector to complete pose adjustment and realize operations such as inserting a pin. At the same time, laser ranging technology is also gradually applied to the normal posture correction process, by measuring the distance between multiple laser points at the end and the surface of the workpiece, calculating the normal of the curved surface, and realizing the fine adjustment of the end posture.
[0004] However, the traditional visual servo method still has great limitations when facing complex curved surface structural parts, pre-assembly holes with variable orientations or unstable lighting environments, etc. First, the reflection of curved surfaces and irregular textures can easily cause feature point extraction errors, making the control results driven by image features unstable. Second, existing deep reinforcement learning has low learning efficiency and parameters are not easy to converge when dealing with high-dimensional continuous action planning of robots, affecting the servo control effect. Therefore, there is an urgent need for a servo positioning method that integrates the perception capabilities of multiple sensors and combines advanced control strategies to meet the high-precision and high-adaptability requirements of automatic assembly of complex curved surface structural parts. SUMMARY
[0005] Therefore, the present application provides a robot hybrid visual servo positioning method, device and equipment to significantly improve the success rate and operation efficiency of robots during assembly.
[0006] Specifically, the present application is realized by the following technical solutions:
[0007] The first aspect of the present application provides a robot hybrid visual servo positioning method, which is applied to a robot, and the end of the robot is loaded with a camera and multiple laser sensors; the method comprises:
[0008] Under the current pose, a visual image of a curved workpiece is collected based on the camera, and distance data from the robot end to the curved workpiece is collected based on the multiple laser sensors;
[0009] identify feature points of the pre-assembly hole of the curved workpiece from the visual image, and determine pixel coordinates of a center of the pre-assembly hole based on the feature points;
[0010] determine depth information of the robot end to the curved workpiece according to the distance data, and construct a first Jacobian matrix of the image feature according to the pixel coordinates and the depth information;
[0011] For each laser sensor, determine a local Jacobian matrix of a laser ranging point corresponding to the laser sensor according to distance data obtained by the laser sensor;
[0012] determine a second Jacobian matrix of the distance feature according to the local Jacobian matrices of all laser ranging points, and determine a joint feature Jacobian matrix according to the first Jacobian matrix and the second Jacobian matrix;
[0013] determine a feature point error matrix according to pixel coordinates of the feature points in the visual image and pixel coordinates of the feature points of the pre-assembly hole in a visual image of the curved workpiece collected under a target pose;
[0014] determine a pose error matrix according to the current pose and the target pose, and determine an overall error matrix according to the pose error matrix and the feature point error matrix;
[0015] input the overall error matrix and the joint feature Jacobian matrix into a pre-trained decision model, determine a current control quantity by the decision model, and control the robot according to the current control quantity.
[0016] The second aspect of the present application provides a robot hybrid visual servo positioning device, which is applied to a robot, and an end of the robot is loaded with a camera and a plurality of laser sensors; the device comprises an acquisition module, a determination module and a control module; wherein,
[0017] The acquisition module is configured to acquire a visual image of a curved workpiece based on the camera under a current pose, and acquire distance data of a robot end to the curved workpiece based on the plurality of laser sensors;
[0018] The determination module is configured to identify feature points of a pre-assembly hole of the curved workpiece from the visual image, and determine pixel coordinates of a center of the pre-assembly hole based on the feature points;
[0019] The determination module is further configured to determine depth information of the robot end to the curved workpiece according to the distance data, and construct a first Jacobian matrix of the image feature according to the pixel coordinates and the depth information;
[0020] The determining module is further configured to determine, for each laser sensor, a local Jacobian matrix of a laser ranging point corresponding to the laser sensor according to distance data obtained by the laser sensor.
[0021] The determining module is further configured to determine a second Jacobian matrix of the distance feature according to the local Jacobian matrices of all the laser ranging points, and determine a joint feature Jacobian matrix according to the first Jacobian matrix and the second Jacobian matrix.
[0022] The determining module is further configured to determine a feature point error matrix according to pixel coordinates of the feature points in the visual image and pixel coordinates of feature points of the pre-assembly hole in the visual image of the curved workpiece acquired under the target pose.
[0023] The determining module is further configured to determine a pose error matrix according to the current pose and the target pose, and determine a total error matrix according to the pose error matrix and the feature point error matrix.
[0024] The control module is configured to input the total error matrix and the joint feature Jacobian matrix into a pre-trained decision model, determine a current control quantity by the decision model, and control the robot according to the current control quantity.
[0025] The third aspect of the present application provides a robot hybrid visual servo positioning device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the steps of any one of the method provided in the first aspect of the present application when executing the program.
[0026] The robot hybrid visual servo positioning method, device and equipment provided by the application first acquire a visual image of a curved workpiece based on a camera under a current pose, and acquire distance data of a robot end to the curved workpiece based on a plurality of laser sensors; secondly, feature points of a pre-assembly hole of the curved workpiece are identified from the visual image, and pixel coordinates of a center of the pre-assembly hole are determined based on the feature points; then, depth information of the robot end to the curved workpiece is determined according to the distance data, and a first Jacobian matrix of an image feature is constructed according to the pixel coordinates and the depth information; again, for each laser sensor, a local Jacobian matrix of a laser ranging point corresponding to the laser sensor is determined according to distance data acquired by the laser sensor; then, a second Jacobian matrix of a distance feature is determined according to the local Jacobian matrices of all laser ranging points, and a joint feature Jacobian matrix is determined according to the first Jacobian matrix and the second Jacobian matrix; again, a feature point error matrix is determined according to pixel coordinates of the feature points in the visual image and pixel coordinates of feature points of the pre-assembly hole in a visual image acquired on the curved workpiece under a target pose; further, a pose error matrix is determined according to the current pose and the target pose, and an overall error matrix is determined according to the pose error matrix and the feature point error matrix; finally, the overall error matrix and the joint feature Jacobian matrix are input into a pre-trained decision model, a current control quantity is determined by the decision model, and the robot is controlled according to the current control quantity.Thus, the first aspect: in the current pose of the robot, the camera is used to collect the image of the curved workpiece in real time, and the multi-laser ranging sensor is used to obtain the distance information of the end to the workpiece surface, effectively combining two-dimensional visual features and three-dimensional distance perception ability. Compared with the traditional method which only relies on image processing, the multi-source perception significantly improves the target detection and pose recognition accuracy on the unstructured and significantly varying curvature workpiece surface, and enhances the robustness and adaptability of the robot system in complex manufacturing scenarios. The second aspect: through the image processing process, including image denoising, edge extraction and fitting, the geometric profile of the pre-assembled hole in the curved workpiece is accurately identified, and the image pixel coordinates of the center are obtained, providing a key plane reference for subsequent three-dimensional positioning. Thus, when the workpiece surface has certain reflection and light changes, the recognition can also be completed, significantly improving the stability of the round hole recognition and solving the problem of feature mismatch of traditional visual servo on complex curved surfaces. The third aspect: combining the recognized pixel coordinates of the center and the depth data provided by the multiple laser range finders, the two-dimensional image features can be effectively converted into three-dimensional space information, and the image feature Jacobian matrix (i.e. the first Jacobian matrix) can be constructed, which can reflect the influence relationship of image feature error on the end velocity of the robot, and realize the accurate modeling of the visual servo system for end control, thereby improving the responsiveness and stability of the robot pose adjustment. The fourth aspect: for each laser sensor, the mapping relationship between the spatial coordinate change of the corresponding ranging point and the end velocity can be calculated by combining its reading and installation pose, forming a local Jacobian matrix. The local Jacobian matrix provides a feedback basis for the normal error between the end and the workpiece surface from multiple angles, so that the robot can not only adjust the position, but also realize high-precision pose adjustment, especially suitable for normal correction in the pre-assembled hole plug task. The fifth aspect: by integrating the first Jacobian matrix of the image feature and the second Jacobian matrix of the laser ranging, a joint feature Jacobian matrix (i.e. the joint feature Jacobian matrix) is constructed, which considers the control requirements of both position and plane normal dimensions, can improve the control integrity of the system for complex curved surface positioning, and solve the limitation of single sensor method in control dimension, so that the robot can maintain smooth motion while maintaining high precision. The sixth aspect: the overall error matrix and the joint feature Jacobian matrix are taken as inputs and sent to the pre-trained TD3 deep reinforcement learning decision model to dynamically generate the current control quantity, guide the robot to make real-time motion adjustment, avoid the problem of manual parameter tuning and fixed parameters in traditional models, realize high adaptability to different workpiece poses, positions and curvatures, significantly improve the task success rate and generalization ability of the robot system, meet the needs of high-demand assembly tasks such as automatic rivet insertion on complex curved surfaces, and the above can significantly improve the success rate and operation efficiency of the robot during assembly, ensure that the workpiece can be accurately aligned with the plug hole, and realize the plug into the hole. BRIEF DESCRIPTION OF DRAWINGS
[0027] Figure 1 A flow chart of the robot hybrid visual servo positioning method embodiment one provided in the present application is shown in the following figure.
[0028] Figure 2 A schematic diagram of a robot is shown in the present application.
[0029] Figure 3 A flow chart of the robot hybrid visual servo positioning method embodiment two provided in the present application is shown in the following figure.
[0030] Figure 4 A hardware structure diagram of the robot hybrid visual servo positioning device provided in the present application is shown in the following figure.
[0031] Figure 5 A structural schematic diagram of the robot hybrid visual servo positioning device embodiment one provided in the present application is shown in the following figure. DETAILED DESCRIPTION
[0032] The example embodiments will be described in detail herein with reference to the accompanying drawings. When the description below refers to accompanying drawings, unless otherwise specified, the same numbers in different drawings refer to the same or similar elements. The embodiments described in the following example embodiments do not represent all embodiments consistent with the present application.
[0033] The terms used in the present application are merely for the purpose of describing particular embodiments and are not intended to limit the present application. As used in the present application, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "and / or," as used herein, refer to and encompass any or all possible combinations of one or more of the associated listed items.
[0034] It should be understood that, although the terms first, second, third, etc. can be employed in this application to describe various information, such information should not be limited to these terms. These terms are only used to differentiate one piece of information from another piece of information. For example, without departing from the scope of the present application, first information can also be referred to as second information, and similarly, second information can also be referred to as first information. Depending on the context, the word "if" as used herein can be interpreted as "when" or "upon" or "in response to determining".
[0035] The following specific embodiments are given to describe the technical solutions of the present application in detail.
[0036] Figure 1 A flow chart of the robot hybrid visual servo positioning method embodiment one provided in the present application is shown in the following figure. Figure 1The method provided by the embodiment is applied to a robot, an end of the robot is loaded with a camera and a plurality of laser sensors, and the method comprises the following steps:
[0037] S101, acquiring a visual image of a curved workpiece based on the camera and acquiring distance data from the robot end to the curved workpiece based on the plurality of laser sensors in a current pose.
[0038] It should be noted that the method introduced in the embodiment can be used for automatic assembly of large-size thin-walled components in aircraft manufacturing. For example, a six-degree-of-freedom industrial robot can be used to realize the positioning task in the nail insertion positioning stage before the drilling and riveting process. Since the workpiece types are tail wings, fuselage sections, stiffeners and the like, these workpieces are often composed of plates and profiles, and have the characteristics of large size, poor rigidity, complex curvature and the like; in order to ensure subsequent riveting and other steps, the nail insertion alignment of the robot is particularly important.
[0039] Specifically, the current pose refers to the real-time position and attitude of the end effector of the robot in the task execution process.
[0040] Figure 2 A schematic diagram of the robot shown in an exemplary embodiment of the present application is shown. Please refer to Figure 2 The camera refers to a monocular camera installed at the end of the robot arm, which is used to acquire image information of the workpiece. In specific implementation, the camera can adopt an "eye on hand" configuration, that is, it is fixedly installed at the end of the robot and moves with it.
[0041] Further, please continue to refer to Figure 2 The laser sensor is a plurality of laser range finders installed at the end of the robot. In specific implementation, the laser sensor can be four laser range finders installed at the end of the robot, which are uniformly distributed around the end of the robot and used to measure the accurate distance values from each point to the surface of the workpiece.
[0042] S102, identifying feature points of a pre-assembly hole of the curved workpiece from the visual image, and determining pixel coordinates of a center of the pre-assembly hole based on the feature points.
[0043] Specifically, please continue to refer to Figure 2 The central part of the curved workpiece has a pre-assembly hole, which is a hole position processed on the workpiece before assembly and used as a hole for nail insertion / drilling and riveting operation. It should be noted that the specific shape of the pre-assembly hole is determined according to actual needs, which is not limited in the embodiment. For example, in an embodiment, the pre-assembly hole can be a circle or an ellipse.
[0044] Further, the feature points refer to pixel point information on the edge of the pre-assembly hole, and the feature points can be used for geometric fitting to help calculate the center position of the pre-assembly hole.
[0045] Further, the threshold segmentation or edge detection method can be used to separate the edge points of the pre-assembly hole from the image, which are located on the contour of the pre-assembly hole and form a closed curve. The feature points are sent to the nonlinear least squares fitting model to fit an optimal circular or elliptical contour.
[0046] Further, the center point coordinates (u, v) of the circle or ellipse are extracted from the fitting result, which is the pixel-level position of the pre-assembly hole in the current image.
[0047] S103, according to the distance data, determine the depth information of the robot end to the curved workpiece, and construct the first Jacobian matrix of the image feature according to the pixel coordinates and the depth information.
[0048] Specifically, the depth information refers to the spatial distance between the robot end and the workpiece surface where the pre-assembly hole is located in the Z direction (i.e. the optical axis direction). Combined with the pixel position information in the image, the image coordinates can be converted into spatial coordinates.
[0049] In specific implementation, the depth information can be calculated based on the following formula:
[0050] ;
[0051] wherein, is the depth information, is the reading of the i-th laser sensor, i is the label of the laser sensor, i is an integer from 1 to 4.
[0052] Further, the first Jacobian matrix is used to describe the mapping matrix of how the image feature error affects the robot end speed, in other words, the first Jacobian matrix is used to convert the data in the image space to the data in the control space.
[0053] In specific implementation, the first Jacobian matrix can be calculated based on the following formula:
[0054] ;
[0055] wherein, is the first Jacobian matrix, λ is the focal length of the camera lens, z is the height coordinate of the feature point along the z axis, and u and v are two coordinates in the pixel coordinates (u, v).
[0056] S104, for each laser sensor, according to the distance data obtained by the laser sensor, determine the local Jacobian matrix of the laser ranging point corresponding to the laser sensor.
[0057] Specifically, the laser ranging point refers to the spatial intersection point formed by the laser beam hitting the workpiece surface. Each laser sensor corresponds to a ranging point.
[0058] Further, the local Jacobian matrix refers to a linear mapping matrix describing the influence of the error of a laser ranging point on the robot end control quantity. The local Jacobian matrix describes the mapping relationship between the distance error of the laser point and the six-dimensional velocity of the robot end.
[0059] In specific implementation, the local Jacobian matrix can be calculated by the following formula:
[0060] ;
[0061] Wherein, is the local Jacobian matrix, is the focal length of the camera , and the ratio of the camera pixel size is is the distance data obtained by the i th laser sensor, is the coordinate of the laser ranging point corresponding to the i th laser sensor.
[0062] It should be noted that since the camera obtains the physical coordinate, the pixel coordinate calculated in the formula is calculated, and each pixel has an actual width and height on the sensor. When converting the physical millimeter coordinate into the pixel coordinate, the pixel size needs to be divided. Because the pixel of the camera used by us is square, the length and width are , so the following can be obtained by transformation .
[0063] S105, according to the local Jacobian matrix of all laser ranging points, determine the second Jacobian matrix of the distance feature, and according to the first Jacobian matrix and the second Jacobian matrix, determine the joint feature Jacobian matrix.
[0064] In specific implementation, the second Jacobian matrix can be calculated based on the following formula:
[0065] ;
[0066] Wherein, is the second Jacobian matrix, , , , respectively, the local Jacobian matrix corresponding to the four laser ranging points.
[0067] Further, the joint feature Jacobian matrix is a composite Jacobian matrix formed by integrating the first Jacobian matrix and the second Jacobian matrix according to certain rules, which is used to uniformly process the influence of image information and distance information on the robot.
[0068] In specific implementation, the joint feature Jacobian matrix can be calculated based on the following formula:
[0069] ;
[0070] wherein, is a joint feature Jacobian matrix, is a first Jacobian matrix, is a second Jacobian matrix.
[0071] S106, determining a feature point error matrix according to the pixel coordinates of the feature points in the visual image and pixel coordinates of feature points of the pre-assembly hole in the visual image of the curved workpiece collected under the target pose.
[0072] Specifically, the target pose can be a current pose: the pixel coordinates of the current pre-assembly hole center point obtained by image processing after the camera collects the image.
[0073] Further, the height and width corresponding to the visual image collected by the camera can be determined in combination with the parameters of the camera. For example, in an embodiment, the image captured by the camera is a 320x200 size image, and it is determined that the height of the camera is 320 and the width is 200 at this time.
[0074] In a specific implementation, the feature point error matrix can be calculated based on the following formula:
[0075] ;
[0076] wherein, is a feature point error matrix, , wherein, represents the distance between the pixel coordinates in the visual image and the pixel coordinates corresponding to the current pose, row is the height of the camera, and col is the width of the camera.
[0077] S107, determining a pose error matrix according to the current pose and the target pose, and determining an overall error matrix according to the pose error matrix and the feature point error matrix.
[0078] Specifically, the current pose and the target pose contain the translation and rotation angle of the robot at two time nodes.
[0079] In a specific implementation, the pose error matrix can be constructed according to the current pose and the target pose as shown below:
[0080] ;
[0081] wherein, is a pose error matrix, is the translation in three directions corresponding to the current pose, , is the translation in three directions corresponding to the target pose, , are rotation angles of three directions corresponding to the target pose. , , are rotation angles of three directions corresponding to the target pose.
[0082] Further, the overall error matrix contains both the spatial error of the position and attitude offset and the visual error of the image plane.
[0083] In specific implementation, the overall error matrix can be calculated based on the following formula:
[0084]
[0085] wherein, is the overall error matrix, is the feature point error matrix, is the pose error matrix.
[0086] S108, input the overall error matrix and the joint feature Jacobian matrix into the pre-trained decision model, determine the current control amount by the decision model, and control the robot according to the current control amount.
[0087] In specific implementation, the overall error matrix used to comprehensively describe the distance between the current robot and the target state and the joint feature Jacobian matrix used to describe the mapping relationship between the overall error and the robot action are taken as the state, and are input into the pre-trained deep reinforcement learning decision model. The model has learned how to make action decisions according to the state error through simulation training, and can output a current control amount, i.e., the direction and amplitude in which the robot end should move and rotate in the six-dimensional space.
[0088] Further, the robot adjusts the robot pose according to the current control amount determined by the decision model, gradually and accurately approaches the target position, and completes the automatic insertion or positioning task on the complex curved surface workpiece.
[0089] The robot hybrid visual servo positioning method provided by the embodiment, the first aspect: in the current pose of the robot, the image of the curved workpiece is collected in real time by the camera, and the distance information of the end to the surface of the workpiece is obtained by combining a plurality of laser ranging sensors, the two-dimensional visual feature and the three-dimensional distance perception ability are effectively fused, compared with the traditional method which only relies on image processing, the target detection and attitude recognition accuracy on the unstructured workpiece surface with significant curvature change is significantly improved through multi-source perception, the robustness and adaptability of the robot system in complex manufacturing scene are enhanced; the second aspect: through the image processing procedure, including image denoising, edge extraction and fitting, the geometric contour of the pre-assembled hole in the curved workpiece is accurately identified, and the image pixel coordinates of the center of the hole are obtained, which provides a key plane reference for subsequent three-dimensional positioning, so that the identification can be completed when there is certain reflection and light change on the surface of the workpiece, the stability of the hole recognition is significantly improved, and the problem of feature mismatch of the traditional visual servo on the complex curved surface is solved; the third aspect: combining the identified pixel coordinates of the center and the depth data provided by the plurality of laser range finders, the two-dimensional image feature can be effectively converted into three-dimensional space information, and an image feature Jacobian matrix (i.e. a first Jacobian matrix) is constructed, which can reflect the influence relationship of image feature error on the end speed of the robot, and the accurate modeling of the visual servo system to the end control is realized, thereby improving the responsiveness and stability of the robot attitude adjustment; the fourth aspect: for each laser sensor, the mapping relationship of the spatial coordinate change of the corresponding ranging point to the end speed can be calculated by combining its reading and installation pose, forming a local Jacobian matrix, the local Jacobian matrix provides a feedback basis for the normal error between the end and the surface of the workpiece from multiple angles, so that the robot can not only adjust the position, but also realize high-precision attitude adjustment, especially suitable for normal correction in the pre-assembled hole plug task; the fifth aspect: by integrating the first Jacobian matrix of the image feature and the second Jacobian matrix of the laser ranging, a joint feature Jacobian matrix (i.e. a joint feature Jacobian matrix) is constructed, which considers the control requirements of both position and plane normal dimensions, can improve the control integrity of the system for positioning on complex curved surfaces, solves the limitation of single sensor method in control dimension, so that the robot can maintain smooth motion while maintaining high precision; the sixth aspect: the overall error matrix and the joint feature Jacobian matrix are taken as inputs and sent into the pre-trained TD3 deep reinforcement learning decision model to dynamically generate the current control quantity and guide the robot to make real-time motion adjustment, avoiding the problems of manual parameter adjustment and fixed parameters in the traditional model, realizing high adaptability response to different workpiece attitudes, positions and curvatures, significantly improving the task success rate and generalization ability of the robot system, meeting the needs of high requirement assembly tasks such as automatic rivet insertion under complex curved surface, and in conclusion, the success rate and operation efficiency of the robot during assembly can be significantly improved, the workpiece can be accurately aligned with the hole vertically, and the pin insertion into the hole is realized.
[0090] Figure 3 The flowchart of the second embodiment of the robot hybrid visual servo positioning method provided in the present application is shown in FIG. 2. Please refer to Figure 3 In the above embodiment, the decision model is a reinforcement learning model based on a double Q network; the training of the decision model comprises the following steps:
[0091] S301, a double Q network is constructed, and the parameters of the double Q network are initialized; wherein the double Q network comprises a decision network for generating a control quantity and an evaluation network for evaluating the value of the control quantity.
[0092] Specifically, the double Q network is one of the core structures of the reinforcement learning algorithm TD3, which is composed of two groups of Q networks (i.e. value networks) for evaluating the value of the action under the current policy to alleviate the overestimation problem of the Q function.
[0093] Further, the decision network and the evaluation network work in parallel but have independent parameters.
[0094] It should be noted that the double Q network takes the smaller value in the update process to improve the estimation conservatism and thus enhance the stability of the policy.
[0095] Further, the neural network can be initialized before the training starts, including randomly setting the network weights and biases, so as to continuously optimize the policy through gradient descent in the subsequent training process.
[0096] S302, a plurality of sets of experience data are obtained by simulation experiments; wherein each set of experience data comprises a sequence composed of a plurality of state-action-reward value pairs and an overall reward value of the task corresponding to the sequence; the state data in each state-action-reward value pair comprises the visual image collected by the camera and the distance data collected by the laser sensor, the action data of the state-action-reward value pair is the control quantity, and the reward value of the state-action-reward value pair is the single-step reward value corresponding to the control quantity.
[0097] Specifically, the experience data refers to the training data collected by the robot through multiple interactions in the simulation task. These data are the "learning materials" of reinforcement learning and are stored in the "experience replay buffer".
[0098] In specific implementation, the experience data can include the state, action, reward and next state of the robot, etc.
[0099] Further, the state includes the image information collected by the camera and the distance information collected by the laser range finder; the action is the control quantity executed by the robot in the current state; and the reward is the immediate feedback (single-step reward value) obtained by the robot after executing the action, which is used to measure the goodness of the current action.
[0100] Further, a complete task process is a sequence, and the overall reward value is the cumulative sum of all single-step reward values in the complete task execution process, reflecting the total performance of the robot in completing the entire servo positioning task.
[0101] In implementation, for example, in an embodiment, a simulation environment including a six-degree-of-freedom robot, a monocular camera, and four laser range finders can be built in the CoppeliaSim simulation platform. A large amount of experience data sequences are collected by continuously running the robot plug task. In each simulation test, the robot starts from the initial pose and gradually adjusts the position and attitude. The system records the current state (including visual image and laser ranging data), the action performed (control quantity), and the single-step reward value after the action is performed at each step to form a state-action-reward triple. A plurality of such triples are concatenated to form a complete task sequence, and the overall reward value corresponding to the sequence is calculated to reflect the performance of the robot in the entire task process. All of these experience data are saved in the "experience replay buffer" for repeated sampling and training of the reinforcement learning model (such as TD3), so as to optimize the control strategy and enable the robot to more efficiently complete the servo positioning and plug operation of the complex curved surface structure in subsequent tasks.
[0102] A specific embodiment is given below to introduce in detail the process of obtaining experience data by simulation test:
[0103] (1) The initial pose of the robot is taken as the current pose, the state data corresponding to the current pose is collected, and the state data is input to the decision network to output the control quantity corresponding to the state data.
[0104] Specifically, the initial pose of the robot is initialized in the simulation environment as the current pose of the task. Then, the system collects the state data (including visual image and laser ranging value) under the pose and inputs it to the trained decision network to output the control quantity to be executed.
[0105] (2) The robot is controlled to act according to the control quantity, and after the robot acts, the updated state data corresponding to the new pose is collected, and the state type is determined according to the updated state data; the state type includes a task success state, a task failure state, and an intermediate state.
[0106] Specifically, the robot moves to a new pose according to the control quantity, and the system immediately collects the new state information and determines whether the current state is a task success state, a failure state, or an intermediate state according to the deviation from the target state.
[0107] (3) The single-step reward value corresponding to the control quantity is determined according to the state type.
[0108] Specifically, according to the state type, a single-step reward value corresponding to the action is calculated in combination with factors such as motion efficiency and pose deviation.
[0109] In a specific implementation, the single-step reward value can be as follows:
[0110] ;
[0111] wherein, is the single-step reward value, E is an energy consumption index representing energy consumption in a task process, is a speed smoothness, is a maximum step length of a task, and superscripts c and I represent is a current state, is an initial position.
[0112] A specific embodiment is given below to introduce in detail the step of determining the single-step reward value corresponding to the control amount according to the state type:
[0113] Step 1: when the state type is a task completion state, an expected path length of a robot from an initial pose to a target pose is calculated according to an initial joint angle of the robot at the initial pose and a target joint angle of the robot at the target pose.
[0114] Step 2: a joint angle under each action is determined according to a pose corresponding to each state action reward value in a current task.
[0115] Step 3: an actual path length of the robot from a starting point to an ending point is determined according to the joint angle corresponding to each state action reward value.
[0116] Step 4: an energy consumption index of current step control is determined according to the expected path length and the actual path length.
[0117] Step 5: an initial motion operability index of the robot under current step control is determined according to a joint angle of the robot at a current pose.
[0118] Step 6: an actual motion operability index of the robot under current step control is determined according to the initial motion operability index and all initial motion operability indexes of all step controls before the current step control.
[0119] Step 7: the single-step reward value corresponding to the control amount is determined according to the energy consumption index and the actual motion operability index.
[0120] In a specific implementation, the joint angle configuration of the robot at the beginning of the task and the target joint angle at the completion of the task are obtained, and the Euclidean distance between the two is calculated as the shortest path length under ideal conditions.
[0121] Further, the joint angle configuration of the robot at the moment is recorded every time a control action is executed; for each state-action pair in the entire task sequence, the corresponding actual joint angle is extracted to form a series of discrete points.
[0122] Further, the joint space distance between adjacent frames is calculated step by step for the joint angle sequence obtained in the previous step, and the obtained path is the real path length in the actual execution process of the robot, which is usually greater than the expected path.
[0123] Further, the energy consumption index can be calculated based on the following formula:
[0124] ;
[0125] where E is the energy consumption index, abs() represents the absolute value, q = [q1 q2 q3 q4 q5 q6] represents the joint values of the six-joint robot, represents the expected position of the robot, represents the initial position of the robot, is the k+1 step joint value, is the k step joint value, and the factor ω is the weight of the energy consumption of different joints, and its value is: ω = [4, 4, 2, 2, 1, 1].
[0126] Further, the motion operability is a comprehensive measure of the motion ability of the robot in each direction under a certain pose, and the motion operability provides an overall scalar description of the gain from the joint speed to the end effector speed .
[0127] Further, the actual motion operability index can be calculated based on the following formula:
[0128] ;
[0129] where, is the actual motion operability index, T is the motion efficiency, and E is the energy consumption index.
[0130] It can be understood that the initial motion operability index is taken as the starting point, and the actual motion operability index at this time can be determined by superimposing all the initial motion operability indexes before the current step, which represents the possibility of the current robot executing the task.
[0131] Further, the joint angle difference between two adjacent steps is weighted and accumulated to obtain an actual joint space path length of the robot from the start point to the end point.
[0132] Optionally, when the state type is the task failure state, a single-step reward value corresponding to the control amount is determined according to the feature point error matrix corresponding to the current pose and the feature point error matrix corresponding to the initial pose.
[0133] Specifically, the single-step reward value of the task failure can be calculated based on the following formula:
[0134] ;
[0135] wherein, is the single-step reward value of the task failure.
[0136] Optionally, when the state type is the intermediate state, a single-step reward value corresponding to the control amount is determined according to a preset maximum moving step number.
[0137] In specific implementation, the specific value of the maximum moving step number is set according to actual needs, which is not limited in the embodiment.
[0138] Further, the negative value of the reciprocal of the maximum moving step number can be directly determined as the single-step reward value corresponding to the control amount.
[0139] (4) The state data, the control amount, and the single-step reward value are spliced into a state-action-reward value pair.
[0140] In specific implementation, for example, in an embodiment, the state data A corresponds to the control amount a, the reward value of a is Ya, and the reward value pair of A-a-Ya is obtained by splicing the three quantities.
[0141] (5) When the state type is the task success or task failure state, an overall reward value of the current task is determined based on the single-step reward values in each state-action-reward value pair in the current task, and all state-action-reward value pairs in the current task and the overall reward value of the current task are saved as an experience data.
[0142] In specific implementation, the cumulative value of all single-step reward values in the current task represents the overall performance of the robot from the start point to the task termination (success or failure), and the experience data is an important indicator for measuring the pros and cons of the control strategy.
[0143] It should be noted that when the task is successful, all single-step reward values are directly accumulated as the overall reward value, and when the task fails, the single-step reward value at the failure moment can be directly taken as the overall reward value to more intuitively reflect the task execution.
[0144] (6) When the state type is an intermediate state, the new pose is taken as the current pose, and the process of collecting state data corresponding to the current pose is performed again.
[0145] Specifically, if the current state is still an intermediate state, the current new pose is taken as a new "current pose", and the next round of state collection and control decision cycle is entered, and the process is continuously executed until the task is completed or failed. This process is repeated continuously.
[0146] It can be understood that, by taking the initial pose of the robot as the starting point, the corresponding visual image and laser ranging data are acquired to form the current state, and the control amount is generated by inputting the pre-trained decision network, to drive the robot to complete an action. After the action is executed, the system collects the updated state data corresponding to the new pose, and judges the state type (success, failure or intermediate state) according to the deviation from the task target, so as to assign a single-step reward with a distinguishing degree to each control amount. Each state, action and reward is packaged into a state-action-reward value pair, when the task is completed (successfully or failed), the system accumulates all rewards in the entire task trajectory to calculate the overall reward, and saves the complete trajectory and total reward of the task as an experience data; if the task has not been completed, the current updated state is taken as a new starting point, and the data collection and control cycle is continued, so that an automated and efficient experience data generation process is realized, providing high-quality and diversified learning samples for the double Q network training based on the TD3 algorithm, thereby improving the strategy accuracy and robustness of the robot in the complex curved surface positioning scene.
[0147] S303, based on the overall reward value of the task corresponding to the sequence in each set of experience data, selecting part of the experience data from the multiple sets of experience data as training samples, and using the training samples to train the double Q network.
[0148] Specifically, the cumulative value of all single-step reward values in a sequence is obtained as the overall reward value, and the higher the overall reward value, the better the task is executed (faster, more accurate, lower energy consumption, closer to the target pose).
[0149] In specific implementation, each sequence contains the state, control action and corresponding single-step reward value of the robot during the execution of the alignment task, then the overall reward value of each sequence is calculated, which measures the overall performance of the task execution, based on these overall reward values, the sequences with better performance are preferentially selected from all experience data as training samples, and the state-action-reward pairs in these samples are input into the double Q network of the TD3 algorithm to guide the learning and updating of the action value of the two evaluation networks, thereby gradually improving the motion strategy of the robot in the complex curved surface servo positioning task.
[0150] A specific embodiment is given below to introduce in detail the process of selecting part of the experience data as training samples from multiple sets of experience data:
[0151] All experience data is sorted according to the task reward value corresponding to the experience data from high to low to obtain a sorting result;
[0152] According to the preset filtering ratio parameter and the number of experience data, the number of filtered samples is determined, and the first target sample corresponding to the number of filtered samples is filtered from the sorting result in descending order;
[0153] Experience data with an overall reward value in a preset range is filtered from the remaining samples as supplementary samples; wherein the upper limit of the preset range is the minimum overall reward value corresponding to the first target sample, and the lower limit of the preset range is the difference between the minimum overall reward value and a preset tolerance parameter;
[0154] The target sample and the supplementary sample are determined as training samples.
[0155] In specific implementation, based on simulation experiments, multiple complete experience data are stored, each data corresponding to a robot task execution process, including multiple state-action-reward value pairs and the overall reward value of the task, i.e. the total score of the task performance. First, all experience data is sorted according to the overall reward value from high to low to form an ordered experience sequence, wherein the experience at the front represents better performance of the robot in these tasks, shorter execution path, more stable control and smaller error.
[0156] Further, according to the preset filtering ratio parameter (for example, the top 10% or 20%) and the total number of experience data, the number of first target samples to be retained is calculated, and the corresponding number of high-quality experience is selected from the front according to the sorting order as the first batch of training samples. These data represent the positive demonstration of "good control behavior" in policy learning.
[0157] Further, part of the "suboptimal experience" is selected from the remaining unselected samples as supplementary samples. In specific implementation, for example, in an embodiment, a preset tolerance parameter (such as a reward fluctuation range ε) can be set, and the minimum overall reward value in the first target sample is taken as the upper limit, and the tolerance parameter is subtracted from the lower limit to form an allowed interval. The experience data in this interval is added as supplementary samples. In this way, through the allowed interval, samples that have not entered the first target sample but have certain value are retained. These supplementary samples can increase the adaptability of the policy to boundary states and non-optimal decisions.
[0158] Further, the first target sample and the supplementary sample are merged to form the training sample set of the current round of the reinforcement learning model, which is sent to the double Q network for policy update.
[0159] It can be understood that, by means of "optimizing and supplementing the boundary", the quality of the training data is ensured, and a certain degree of strategy diversity is reserved, and the learning speed, convergence stability and generalization ability of the robot in the complex curved surface positioning task are significantly improved.
[0160] The robot hybrid visual servo positioning method provided in the embodiment firstly solves the problem of overestimation of action value in the traditional single Q network by constructing a double Q network structure and setting two independent evaluation networks to estimate the value of the same action, thereby improving the accuracy of the model in judging the quality of the action and the stability of the control strategy; secondly, a large number of experiments are carried out in the simulation environment, and a plurality of groups of experience data containing image information and distance information are collected, each group of data completely covering the whole process of the robot from the initial pose to the target pose, thereby providing a rich and real perception-decision-feedback link for learning an accurate control strategy; finally, high-quality experience sequences are preferentially selected for training by evaluating the overall reward value of the task, thereby avoiding the interference of inefficient data on the training and accelerating the convergence speed of the strategy; in summary, by means of the double Q network architecture + high-quality simulation experience collection + sample screening strategy based on the overall reward, high-precision, adaptive and high-robust positioning control of the robot on the complex curved surface pre-assembly hole is realized, and the task completion efficiency and success rate are significantly improved.
[0161] A process of experimental verification is given below:
[0162] Specifically, a virtual environment is built on the CoppeliaSim simulation platform for training. Since the robot library of the CoppeliaSim software itself does not have an ABB IRB 120 robot, the URDF file of the selected robot is first imported, and the absolute positions of the robot DAE file and STL file in the file are simultaneously changed, so that collision and kinematics simulation can be realized subsequently. A vision camera and four laser range finders are imported at the end of the robot. The camera adopts an "eye on hand" layout, the camera resolution is "320x240" (the camera is directly installed at the end of the robot (usually near the end effector) and moves together with the arm.), and the depth information is ignored, so that the depth information provided by the laser range finder is used in the positioning process. The laser range finders are evenly distributed around the end effector of the robot arm and as close to the center as possible to ensure that as small an area as possible is covered, so that the normal of the target point plane can be calculated more accurately; at the same time, the workpiece is imported into the simulation environment, the workpiece has a thickness of 15 mm, and the surface is a single curvature surface, and a through hole is formed on the top for subsequent rivet insertion, and the direction of the hole is perpendicular to the curved surface.
[0163] Further, the algorithm runs on a Dell Precision-5820-Tower workstation with an Intel Xeon(R) w-2255 processor at a frequency of 3.7 GHz, 96 G memory, an NVIDIA RTX 3080 graphics card, an Ubuntu 20.04 operating system, an Anaconda python 3.6.15 environment configuration, and Pytorch 2.1.0. During the training process, the position of the workpiece is randomly refreshed each time to ensure the generalization ability and effectiveness of the training. The final trained model is saved as a.pth file, and the related data of the training is stored in the log.
[0164] Further, the CatalystRL framework used is a distributed framework for reproducible reinforcement learning research, which mainly includes three components: a sampler, a trainer, and a database. The sampler is responsible for obtaining states and observations from the environment, determining the next action through the current model, and applying these actions to the environment to obtain new states and rewards. These data are stored in a trajectory buffer. The trainer part adjusts the model parameters through various optimization algorithms to optimize the decision-making process of the model. The replay buffer and queue system are used to manage experience data, so that historical data can be efficiently used for learning. The database is used for persistent storage and query of trajectory data generated during the training process. In this study, the MongoDB system is used to efficiently manage and access data. Overall, this architecture realizes efficient training and parameter optimization of reinforcement learning models in continuous state space through these three interrelated components.
[0165] Further, the TD3 algorithm uses a double-layer “Actor-Critic” structure, with Actor and Critic each containing two hidden layers with 256 neurons each. Each Actor network corresponds to two Critic networks, with a learning rate of 0.0003 and a ReLU activation function. The batchsize is 16.
[0166] Further, after successful training, the data in the MongoDB system is read for analysis. It can be seen that the algorithm decreases rapidly and has small fluctuations, reaching convergence around 7000 steps, indicating that the robot has learned a reasonable strategy.
[0167] Further, the reward function during the training process almost never over-explores during the convergence process, and the overall growth is relatively stable. After 15000 steps, it can be stabilized above 1.6, and it can be considered that the pre-assembly hole insertion is basically completed, and the algorithm has good results.
[0168] Further, the training process is after the number of steps reaches 2500 steps, the robot motion efficiency is basically maintained at 2.8, indicating that the robot can perform fast task execution actions in the action space, reducing the motion time, and also speeding up the training process.
[0169] It can be understood from the above experimental results that the training time of the proposed algorithm is shorter, and the trained robot can quickly realize the re-pinning action for the designed pre-assembled hole, and still has a good success rate and efficiency when the workpiece position changes.
[0170] Another embodiment is given below for different experimental verification:
[0171] Further, the thickness of the workpiece is changed to 10mm and 20mm respectively, and is introduced into the simulation environment. First, load the trained model to perform 50 positioning experiments, and the position of the workpiece is still randomly generated during the process. Then, the algorithm is changed to PBVS and IBVS method to perform 50 experiments on the workpiece with a thickness of 15mm.
[0172] Further, the success rate is the highest when the workpiece thickness is the thinnest, which is 80%. Because the smaller the thickness, the closer the target point to the workpiece surface, the easier the action point at the bottom of the rivet to reach the target position point. Even if there is a slight deviation in the normal angle, the gap between the rivet and the hole diameter will make the positioning task complete. Although the positioning success rate decreases when the workpiece thickness increases, the model still continuously self-learns and trains, and the positioning task success rate of the 20mm thick workpiece can still reach 74% in 50 experiments. It reflects the correctness and generality.
[0173] Further, the BVS method has a low success rate because the workpiece position is randomly generated each time, and the feature points may disappear in the field of view during movement, resulting in positioning failure. The IBVS method is prone to cause the robot speed to exceed the limit due to the large distance, and the success rate is also low. The proposed method has obvious advantages.
[0174] Further, since the workpiece only changes in position during the training process, the hole is always located at the top, i.e. the orientation of the hole is always perpendicular to the ground. In order to verify the generalization of the algorithm, the position and orientation of the hole on the workpiece are redesigned under the premise that the workpiece thickness is 15mm and the direction of the hole is perpendicular to the curved surface, and 50 experiments are conducted in the simulation environment.
[0175] Further, the algorithm still basically meets the positioning requirements after the orientation of the hole changes. With the continuous self-training of the model during the positioning process, the success rate is high, indicating that the algorithm has a certain generalization and is still meaningful when facing different orientation holes or different curvature workpieces.
[0176] Corresponding to the foregoing embodiment of the robot hybrid visual servo positioning method, the application also provides an embodiment of a robot hybrid visual servo positioning device.
[0177] The embodiment of the robot hybrid visual servo positioning device of the application can be applied to a robot hybrid visual servo positioning apparatus. The device embodiment can be realized by software, or realized by hardware or a combination of software and hardware. Taking software realization as an example, as a device in a logical sense, it is formed by reading the corresponding computer program instructions in the non-volatile memory into the memory and running by the processor of the robot hybrid visual servo positioning apparatus where the device is located. From the hardware level, as shown in the figure, Figure 4 As shown in the figure, it is a hardware structure diagram of the robot hybrid visual servo positioning apparatus where the robot hybrid visual servo positioning device of the application is located. In addition to the processor, the memory, the network interface, and the non-volatile memory shown in the figure, the robot hybrid visual servo positioning apparatus where the device is located in the embodiment can also include other hardware according to the actual function of the robot hybrid visual servo positioning device, and details are not repeated. Figure 4 As shown in the figure, it is a hardware structure diagram of the robot hybrid visual servo positioning apparatus where the robot hybrid visual servo positioning device of the application is located. In addition to the processor, the memory, the network interface, and the non-volatile memory shown in the figure, the robot hybrid visual servo positioning apparatus where the device is located in the embodiment can also include other hardware according to the actual function of the robot hybrid visual servo positioning device, and details are not repeated.
[0178] Figure 5 The structure schematic diagram of the embodiment one of the robot hybrid visual servo positioning device provided by the application is shown in the figure. Please refer to Figure 5 The device provided by the embodiment includes a collection module 510, a determination module 520, and a control module 530. Wherein,
[0179] The collection module 510 is configured to collect a visual image of a curved workpiece based on the camera under the current pose, and collect distance data from the robot end to the curved workpiece based on the plurality of laser sensors;
[0180] The determination module 520 is configured to identify feature points of a pre-assembly hole of the curved workpiece from the visual image, and determine pixel coordinates of the center of the pre-assembly hole based on the feature points;
[0181] The determination module 520 is further configured to determine depth information of the robot end to the curved workpiece according to the distance data, and construct a first Jacobian matrix of the image feature according to the pixel coordinates and the depth information;
[0182] The determination module 520 is further configured to determine, for each laser sensor, a local Jacobian matrix of a laser ranging point corresponding to the laser sensor according to the distance data obtained by the laser sensor;
[0183] The determination module 520 is further configured to determine a second Jacobian matrix of the distance feature according to a local Jacobian matrix of all laser ranging points, and determine a joint feature Jacobian matrix according to the first Jacobian matrix and the second Jacobian matrix.
[0184] The determination module 520 is further configured to determine a feature point error matrix according to pixel coordinates of the feature points in the visual image and pixel coordinates of feature points of the pre-assembly hole in the visual image of the curved workpiece captured under the target pose.
[0185] The determination module 520 is further configured to determine a pose error matrix according to the current pose and the target pose, and determine an overall error matrix according to the pose error matrix and the feature point error matrix.
[0186] The control module 530 is configured to input the overall error matrix and the joint feature Jacobian matrix into a pre-trained decision model, determine a current control quantity by the decision model, and control the robot according to the current control quantity.
[0187] The device of the embodiment can be used to execute the method of the embodiment. Figure 1 The steps of the method embodiment are similar in implementation principle and process, and will not be described here.
[0188] Please continue to refer to Figure 4 The application further provides a robot hybrid visual servo positioning device, which comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the steps of any one of the methods provided in the first aspect of the application when executing the program.
[0189] The application further provides a computer readable storage medium, which stores a computer program, and the program is executed by the processor to implement the steps of any one of the methods provided in the application.
[0190] The implementation process of the functions and roles of each unit in the above device is specifically described in the implementation process of the corresponding steps in the above method, and will not be described here.
[0191] For the apparatus embodiment, since it basically corresponds to the method embodiment, the relevant part can be seen from the part of the method embodiment. The apparatus embodiment described above is only illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purpose of the application according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0192] The above only describes the preferred embodiments of the present application and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A robotic hybrid visual servoing positioning method, characterized in that, The method is applied to a robot, an end of the robot is loaded with a camera and a plurality of laser sensors; the method comprises: In a current pose, a visual image of a curved workpiece is captured based on the camera, and distance data of the robot end to the curved workpiece is captured based on the plurality of laser sensors; Feature points of a pre-assembly hole of the curved workpiece are identified from the visual image, and pixel coordinates of a center of the pre-assembly hole are determined based on the feature points; Depth information of the robot end to the curved workpiece is determined according to the distance data, and a first Jacobian matrix of an image feature is constructed according to the pixel coordinates and the depth information; For each laser sensor, a local Jacobian matrix of a laser ranging point corresponding to the laser sensor is determined according to the distance data obtained by the laser sensor; A second Jacobian matrix of a distance feature is determined according to the local Jacobian matrices of all laser ranging points, and a joint feature Jacobian matrix is determined according to the first Jacobian matrix and the second Jacobian matrix; A feature point error matrix is determined according to the pixel coordinates of the feature points in the visual image and pixel coordinates of feature points of the pre-assembly hole in a visual image of the curved workpiece captured in a target pose; A pose error matrix is determined according to the current pose and the target pose, and an overall error matrix is determined according to the pose error matrix and the feature point error matrix; The overall error matrix and the joint feature Jacobian matrix are input into a pre-trained decision model, a current control quantity is determined by the decision model, and the robot is controlled according to the current control quantity. The decision model is a reinforcement learning model based on a double Q network, and a training process of the decision model comprises: taking an initial pose of the robot as a current pose, capturing state data corresponding to the current pose, and determining a single-step reward value corresponding to a control quantity according to a state type; when the state type is a task completion state, calculating an expected path length of the robot from the initial pose to the target pose according to an initial joint angle of the robot in the initial pose and a target joint angle of the robot in the target pose; determining a joint angle in each action step according to a pose corresponding to each state action reward value; determining an actual path length of the robot from the starting point to the ending point according to the joint angle corresponding to each state action reward value; determining an energy consumption index of the current step control according to the expected path length and the actual path length; determining an actual motion operability index of the robot in the current step control according to the joint angle of the robot in the current pose; determining an initial motion operability index of the robot in the current step control according to the initial motion operability index and all initial motion operability indexes of all step controls before the current step control; and determining the single-step reward value corresponding to the control quantity according to the energy consumption index and the actual motion operability index.
2. The method of claim 1, wherein, The training process of the decision model comprises: constructing a double Q network and initializing parameters of the double Q network; wherein the double Q network comprises a decision network for generating a control quantity and an evaluation network for evaluating a control quantity value; obtaining a plurality of sets of experience data through simulation experiments; wherein each set of experience data comprises a sequence composed of a plurality of state-action-reward value pairs and an overall reward value of a task corresponding to the sequence; state data in each state-action-reward value pair comprises a visual image collected by a camera and distance data collected by a laser sensor, action data of the state-action-reward value pair is a control quantity, and a reward value of the state-action-reward value pair is a single-step reward value corresponding to the control quantity; selecting part of the experience data from the plurality of sets of experience data as training samples based on the overall reward value of the task corresponding to the sequence in each set of experience data, and using the training samples to train the double Q network.
3. The method of claim 2, wherein, The experience data is obtained through simulation experiments, comprising: inputting the state data into the decision network to output a control quantity corresponding to the state data from the decision network; controlling the robot to act according to the control quantity, collecting updated state data corresponding to a new pose after the robot acts, and determining a state type according to the updated state data; the state type comprises a task success state, a task failure state and an intermediate state; determining a single-step reward value corresponding to the control quantity according to the state type; concatenating the state data, the control quantity and the single-step reward value into a state-action-reward value pair; when the state type is the task success state or the task failure state, determining an overall reward value of a current task based on single-step reward values in each of the state-action-reward value pairs in the current task, and saving all state-action-reward value pairs in the current task and the overall reward value of the current task as an experience data; when the state type is the intermediate state, taking the new pose as a current pose and performing the process of collecting state data corresponding to the current pose again.
4. The method of claim 1, wherein, The single-step reward value corresponding to the control quantity is determined according to the state type, comprising: when the state type is the task failure state, determining the single-step reward value corresponding to the control quantity according to a feature point error matrix corresponding to the current pose and a feature point error matrix corresponding to an initial pose.
5. The method of claim 4, wherein, The single-step reward value corresponding to the control quantity is determined according to the state type, comprising: when the state type is the intermediate state, determining the single-step reward value corresponding to the control quantity according to a preset maximum number of movement steps.
6. The method of claim 3, wherein, The single-step reward value corresponding to the control quantity is determined according to the state type, comprising: when the state type is the intermediate state, determining the single-step reward value corresponding to the control quantity according to a preset maximum number of movement steps. The single-step reward value corresponding to the control quantity is determined according to the state type, comprising: sorting all experience data in descending order of overall reward values of tasks corresponding to the experience data to obtain a sorting result; determining a number of screening samples according to a preset screening proportion parameter and a number of the experience data, and screening first target samples corresponding to the number of screening samples from the sorting result in descending order. The experience data with the overall reward value in a preset range is screened out from the remaining samples as the supplementary samples; wherein, the upper limit of the preset range is the minimum overall reward value corresponding to the first target sample, and the lower limit of the preset range is the difference between the minimum overall reward value and a preset tolerance parameter; The target sample and the supplementary sample are determined as training samples.
7. A robotic hybrid visual servoing positioning apparatus, comprising: The device is applied to a robot, and an end of the robot is loaded with a camera and a plurality of laser sensors; the device comprises a collection module, a determination module and a control module; wherein, The collection module is configured to collect a visual image of a curved workpiece based on the camera and distance data of the robot end to the curved workpiece based on the plurality of laser sensors at a current pose; The determination module is configured to identify feature points of a pre-assembly hole of the curved workpiece from the visual image, and determine pixel coordinates of a center of the pre-assembly hole based on the feature points; The determination module is further configured to determine depth information of the robot end to the curved workpiece according to the distance data, and construct a first Jacobian matrix of an image feature according to the pixel coordinates and the depth information; The determination module is further configured to determine, for each laser sensor, a local Jacobian matrix of a laser ranging point corresponding to the laser sensor according to the distance data obtained by the laser sensor; The determination module is further configured to determine a second Jacobian matrix of distance features according to the local Jacobian matrices of all laser ranging points, and determine a joint feature Jacobian matrix according to the first Jacobian matrix and the second Jacobian matrix; The determination module is further configured to determine a feature point error matrix according to the pixel coordinates of the feature points in the visual image and pixel coordinates of feature points of the pre-assembly hole in a visual image collected at a target pose on the curved workpiece; The determination module is further configured to determine a pose error matrix according to the current pose and the target pose, and determine an overall error matrix according to the pose error matrix and the feature point error matrix; The control module is configured to input the overall error matrix and the joint feature Jacobian matrix into a pre-trained decision model, determine a current control amount by the decision model, and control the robot according to the current control amount. The decision model is a double Q network-based reinforcement learning model, and a training process of the decision model includes: taking an initial pose of the robot as a current pose, collecting state data corresponding to the current pose, and determining a single-step reward value corresponding to a control quantity according to a state type; when the state type is a task completion state, calculating an expected path length of the robot from the initial pose to the target pose according to an initial joint angle of the robot at the initial pose and a target joint angle of the robot at the target pose; determining a joint angle at each action under each state action reward value; determining an actual path length of the robot from the start point to the end point according to the joint angle corresponding to each state action reward value; determining an energy consumption index of the current step control according to the expected path length and the actual path length; determining an initial motion operability index of the robot at the current step control according to the joint angle of the robot at the current pose; determining an actual motion operability index of the robot at the current step control according to the initial motion operability index and all initial motion operability indexes of all step controls before the current step control; and determining the single-step reward value corresponding to the control quantity according to the energy consumption index and the actual motion operability index.
8. A robotic hybrid visual servoing positioning apparatus, comprising: Computer program product comprising a memory, a processor and a computer program stored on the memory and loadable into the processor, characterized in that the processor implements the steps of the method according to any one of claims 1 to 6 when executing the program.
Citation Information
Patent Citations
Mechanical arm control method based on deep reinforcement learning
CN116533249A
Mixed mode depth imaging
CN117043547A