Robot hybrid visual servo positioning method, device and equipment
By combining camera and laser sensor to identify feature points of curved workpieces, constructing Jacobian matrix and inputting it into the decision model, the accuracy and adaptability problems of traditional visual servoing methods on complex curved surfaces are solved, and efficient and stable robot assembly operations are realized.
Patent Information
- Application Number
- CN202511526520.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-24
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2045-10-24
AI Technical Summary
Traditional visual servoing methods suffer from large feature point extraction errors and low efficiency in deep reinforcement learning when dealing with complex curved surface structures, resulting in unstable robot servo control and difficulty in achieving high-precision and highly adaptable assembly.
A hybrid vision servo positioning method for robots is adopted, which combines a camera and multiple laser sensors. By identifying the feature points of pre-assembly holes on curved workpieces, Jacobian matrix and error matrix are constructed and input into a pre-trained decision model to generate control variables to guide robot movement.
It significantly improves the robot's target detection accuracy and posture recognition robustness on complex curved surfaces, enabling high-precision and stable posture adjustment and control, and improving assembly success rate and efficiency.
Smart Images

Figure CN121004614A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, and in particular to a method, apparatus and device for robot hybrid vision servo positioning. Background Technology
[0002] Robots are playing an increasingly important role in the field of high-end equipment manufacturing, and robotics technology has become a key force in improving production efficiency and ensuring assembly accuracy. In manufacturing, due to the large size, low rigidity, and high processing difficulty of structural components, how to quickly and accurately identify and position pre-assembly holes has become one of the key bottlenecks in improving manufacturing efficiency and quality.
[0003] Currently, a monocular vision-based servo positioning method has been explored. This method uses a camera to acquire images, extracts image features of pre-assembled holes on the workpiece surface, and combines image processing and servo control algorithms to drive the robot's end effector to complete pose adjustments and perform operations such as pin insertion. Simultaneously, laser ranging technology is gradually being applied to normal attitude correction. By measuring the distances between multiple laser points at the end effector and the workpiece surface, the surface normal is calculated, enabling fine-tuning of the end effector's attitude.
[0004] However, traditional visual servoing methods still have significant limitations when dealing with complex curved surface structures, pre-assembled holes with varying orientations, or unstable lighting environments. Firstly, surface reflections and irregular textures can easily cause errors in feature point extraction, leading to unstable image feature-driven control results. Secondly, existing deep reinforcement learning methods suffer from low learning efficiency and difficulty in parameter convergence when handling high-dimensional continuous motion planning for robots, affecting servo control performance. Therefore, there is an urgent need for a servo positioning method that integrates multi-sensor perception capabilities and advanced control strategies to meet the high-precision and highly adaptable requirements of automatic assembly of complex curved surface structures. Summary of the Invention
[0005] In view of this, this application provides a robot hybrid vision servo positioning method, apparatus and device to significantly improve the success rate and operational efficiency of the robot during assembly.
[0006] Specifically, this application is implemented through the following technical solution:
[0007] The first aspect of this application provides a hybrid visual servo positioning method for a robot, the method being applied to a robot whose end effector is equipped with a camera and multiple laser sensors; the method includes:
[0008] In the current pose, visual images of the curved workpiece are acquired based on the camera, and distance data from the robot end effector to the curved workpiece is acquired based on the multiple laser sensors.
[0009] Feature points of the pre-assembly hole of the curved workpiece are identified from the visual image, and the pixel coordinates of the center of the pre-assembly hole are determined based on the feature points.
[0010] Based on the distance data, the depth information from the robot end effector to the curved workpiece is determined, and a first Jacobian matrix of image features is constructed based on the pixel coordinates and the depth information;
[0011] For each laser sensor, the local Jacobian matrix of the laser ranging point corresponding to that laser sensor is determined based on the distance data acquired by that laser sensor.
[0012] Based on the local Jacobian matrices of all laser ranging points, determine the second Jacobian matrix of the distance features, and based on the first and second Jacobian matrices, determine the joint feature Jacobian matrix;
[0013] The feature point error matrix is determined based on the pixel coordinates of the feature points in the visual image and the pixel coordinates of the feature points of the pre-assembly holes in the visual image of the curved workpiece acquired under the target pose.
[0014] Based on the current pose and the target pose, a pose error matrix is determined, and based on the pose error matrix and the feature point error matrix, an overall error matrix is determined.
[0015] The overall error matrix and the joint feature Jacobian matrix are input into a pre-trained decision model, which determines the current control quantity and controls the robot according to the current control quantity.
[0016] A second aspect of this application provides a robot hybrid vision servo positioning device, the device being applied to a robot, the robot's end effector being equipped with a camera and multiple laser sensors; the device includes an acquisition module, a determination module, and a control module; wherein...
[0017] The acquisition module is used to acquire visual images of the curved workpiece based on the camera in the current pose, and to acquire distance data from the robot end effector to the curved workpiece based on the multiple laser sensors.
[0018] The determining module is used to identify feature points of the pre-assembly hole of the curved workpiece from the visual image, and determine the pixel coordinates of the center of the pre-assembly hole based on the feature points.
[0019] The determining module is further configured to determine the depth information from the robot end effector to the curved workpiece based on the distance data, and construct a first Jacobian matrix of image features based on the pixel coordinates and the depth information;
[0020] The determining module is further configured to, for each laser sensor, determine the local Jacobian matrix of the laser ranging point corresponding to that laser sensor based on the distance data acquired by that laser sensor;
[0021] The determining module is further configured to determine the second Jacobian matrix of the distance features based on the local Jacobian matrices of all laser ranging points, and to determine the joint feature Jacobian matrix based on the first Jacobian matrix and the second Jacobian matrix.
[0022] The determining module is further configured to determine the feature point error matrix based on the pixel coordinates of the feature points in the visual image and the pixel coordinates of the feature points of the pre-assembly holes in the visual image of the curved workpiece acquired under the target pose.
[0023] The determining module is further configured to determine a pose error matrix based on the current pose and the target pose, and to determine an overall error matrix based on the pose error matrix and the feature point error matrix.
[0024] The control module is used to input the overall error matrix and the joint feature Jacobian matrix into a pre-trained decision model, which determines the current control quantity and controls the robot according to the current control quantity.
[0025] A third aspect of this application provides a robot hybrid vision servo positioning device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the methods provided in the first aspect of this application.
[0026] The robot hybrid vision servo positioning method, apparatus, and device provided in this application first acquire visual images of the curved workpiece based on a camera in the current pose, and acquire distance data from the robot's end effector to the curved workpiece based on multiple laser sensors; secondly, feature points of pre-assembly holes on the curved workpiece are identified from the visual images, and the pixel coordinates of the center of the pre-assembly holes are determined based on the feature points; then, the depth information from the robot's end effector to the curved workpiece is determined based on the distance data, and a first Jacobian matrix of image features is constructed based on the pixel coordinates and depth information; then, for each laser sensor, the local Jacobian matrix of the laser ranging point corresponding to that laser sensor is determined based on the distance data acquired by that laser sensor; and then, based on all... The local Jacobian matrix of the laser ranging point is used to determine the second Jacobian matrix of the distance features. Based on the first and second Jacobian matrices, the joint feature Jacobian matrix is determined. Then, based on the pixel coordinates of the feature points in the visual image and the pixel coordinates of the feature points of the pre-assembled holes in the visual image of the curved workpiece acquired under the target pose, the feature point error matrix is determined. Subsequently, based on the current pose and the target pose, the pose error matrix is determined. Based on the pose error matrix and the feature point error matrix, the overall error matrix is determined. Finally, the overall error matrix and the joint feature Jacobian matrix are input into the pre-trained decision model, which determines the current control quantity and controls the robot according to the current control quantity.In this way, firstly, under the robot's current pose, real-time images of the curved workpiece are acquired using a camera, while multiple laser rangefinders are used to obtain multi-point distance information from the end effector to the workpiece surface. This effectively integrates two-dimensional visual features with three-dimensional distance perception capabilities. Compared to traditional methods that rely solely on image processing, multi-source perception significantly improves the accuracy of target detection and posture recognition on unstructured workpiece surfaces with significant curvature variations, enhancing the robustness and adaptability of the robot system in complex manufacturing scenarios. Secondly, through image processing procedures, including image denoising, edge extraction, and fitting, the geometric contours of pre-assembled holes in curved workpieces are accurately identified, and the image pixel coordinates of their centers are obtained, providing crucial information for subsequent three-dimensional positioning. The key uses a planar reference, allowing for recognition even when the workpiece surface has some reflection or changes in lighting, significantly improving the stability of circular hole recognition and solving the feature mismatch problem of traditional visual servoing on complex curved surfaces. Thirdly, by combining the recognized center pixel coordinates with depth data provided by multiple laser rangefinders, two-dimensional image features can be effectively converted into three-dimensional spatial information, and an image feature Jacobian matrix (i.e., the first Jacobian matrix) can be constructed. This matrix reflects the influence of image feature errors on the robot's end-effector speed, achieving precise modeling of the end-effector control by the visual servoing system, thereby improving the responsiveness and stability of robot posture adjustment. Fourthly, for each laser sensor, its readings and installation pose are combined... The mapping relationship between the spatial coordinate changes of the corresponding ranging point and the end-effector velocity can be calculated, forming a local Jacobian matrix. The local Jacobian matrix provides a feedback basis for the normal error between the end-effector and the workpiece surface from multiple perspectives, enabling the robot not only to adjust its position but also to achieve high-precision attitude adjustment, which is particularly suitable for normal correction in pre-assembly hole insertion tasks. Fifthly, by integrating the first Jacobian matrix of image features and the second Jacobian matrix of laser ranging, a joint feature Jacobian matrix (i.e., joint feature Jacobian matrix) is constructed. This matrix considers the control requirements of both position and planar normal dimensions, which can improve the control integrity of the system for positioning complex curved surfaces and solve the limitations of single sensor methods in the control dimension. The limitations of the robot allow it to maintain high precision while ensuring smooth movement. Sixthly, the overall error matrix and joint feature Jacobian matrix are input into a pre-trained TD3 deep reinforcement learning decision model to dynamically generate the current control quantity, guiding the robot to make real-time motion adjustments. This avoids the problems of manual parameter tuning and parameter fixing in traditional models, achieving a highly adaptive response to different workpiece postures, positions, and curvatures. This significantly improves the robot system's task success rate and generalization ability, meeting the needs of high-requirement assembly tasks such as automatic rivet insertion on complex curved surfaces. In summary, this significantly improves the robot's success rate and operational efficiency during assembly, ensuring that the workpiece can be accurately aligned perpendicularly with the insertion hole, thus enabling pin insertion. Attached Figure Description
[0027] Figure 1 A flowchart of an embodiment of the robot hybrid vision servo positioning method provided in this application;
[0028] Figure 2 A schematic diagram of a robot shown in an exemplary embodiment of this application;
[0029] Figure 3 A flowchart of Embodiment 2 of the robot hybrid vision servo positioning method provided in this application;
[0030] Figure 4 This is a hardware structure diagram of the robot hybrid vision servo positioning device, which is the robot hybrid vision servo positioning device of this application.
[0031] Figure 5 This is a schematic diagram of the structure of a first embodiment of the robot hybrid vision servo positioning device provided in this application. Detailed Implementation
[0032] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application.
[0033] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used herein are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0034] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0035] The following specific embodiments are given to illustrate the technical solution of this application in detail.
[0036] Figure 1 This is a flowchart of an embodiment of the robot hybrid vision servo positioning method provided in this application. Please refer to... Figure 1The method provided in this embodiment is applied to a robot, the end effector of which is equipped with a camera and multiple laser sensors; the method includes:
[0037] S101. In the current pose, acquire visual images of the curved workpiece based on the camera, and acquire distance data from the robot end effector to the curved workpiece based on the multiple laser sensors.
[0038] It should be noted that the method described in this embodiment can be used for the automated assembly of large-sized thin-walled components in aircraft manufacturing. For example, a six-degree-of-freedom industrial robot can be used to perform the positioning task in the pin positioning stage before the drilling and riveting process. Since the workpieces are structures such as aircraft tail fins, fuselage sections, and reinforcing ribs, these workpieces are often made of sheet metal and profiles, and have characteristics such as large size, poor rigidity, and complex curvature. To ensure the subsequent riveting steps, the robot's pin alignment is particularly important.
[0039] Specifically, the current pose refers to the real-time position and orientation of the robot's end effector during task execution.
[0040] Figure 2 This is a schematic diagram of a robot illustrating an exemplary embodiment of this application. Please refer to... Figure 2 The camera refers to a monocular camera mounted at the end of a robot's robotic arm, used to acquire image information of the workpiece. In practice, the camera can be configured as an "eye on the hand," meaning it is fixedly mounted at the end of the robot and moves with it.
[0041] For further details, please refer to [link / reference]. Figure 2 The laser sensors consist of multiple laser rangefinders installed at the end of the robot. In a specific implementation, there may be four laser rangefinders installed at the end of the robot, evenly distributed around the perimeter, used to measure the precise distance from each point to the workpiece surface.
[0042] S102. Identify the feature points of the pre-assembly hole of the curved workpiece from the visual image, and determine the pixel coordinates of the center of the pre-assembly hole based on the feature points.
[0043] For details, please continue to refer to... Figure 2 The curved workpiece has a pre-assembly hole in its center. This pre-assembly hole is a pre-machined hole on the workpiece before assembly, used for inserting rivets / drilling / riveting operations. It should be noted that the specific shape of the pre-assembly hole is determined according to actual needs, and this embodiment does not limit it. For example, in one embodiment, the pre-assembly hole can be circular or elliptical.
[0044] Furthermore, feature points refer to the pixel information on the edge of the pre-assembled hole. Geometric fitting can be performed using feature points to help calculate the center position of the pre-assembled hole.
[0045] Furthermore, threshold segmentation or edge detection methods can be used to separate the edge points of the pre-assembled holes from the image. These points are located on the contour of the pre-assembled holes and form closed curves. The above feature points are fed into a nonlinear least squares fitting model to fit an optimal circular or elliptical contour.
[0046] Furthermore, the center point coordinates (u, v) of the circle or ellipse are extracted from the fitting results, which are the pixel-level positions of the pre-assembled holes in the current image.
[0047] S103. Based on the distance data, determine the depth information from the robot end to the curved workpiece, and construct a first Jacobian matrix of image features based on the pixel coordinates and the depth information.
[0048] Specifically, depth information refers to the spatial distance in the Z direction (i.e., the optical axis direction) between the robot end effector and the workpiece surface where the pre-assembly hole is located. Combined with the pixel position information in the image, the image coordinates can be converted into spatial coordinates.
[0049] In practice, depth information can be calculated based on the following formula:
[0050] ;
[0051] in, For depth information, This represents the reading of the i-th laser sensor, where i is the label of the laser sensor and can be an integer from 1 to 4.
[0052] Furthermore, the first Jacobian matrix is used to describe the mapping matrix of how image feature errors affect the robot's end-effector velocity; in other words, the first Jacobian matrix is used to convert data in the image space into data in the control space.
[0053] In practice, the first Jacobian matrix can be calculated based on the following formula:
[0054] ;
[0055] in, Let λ be the first Jacobian matrix, λ be the focal length of the camera lens, z be the height coordinate of the feature point along the z-axis, and u and v be two coordinates in the pixel coordinate (u, v).
[0056] S104. For each laser sensor, determine the local Jacobian matrix of the laser ranging point corresponding to the laser sensor based on the distance data acquired by the laser sensor.
[0057] Specifically, a laser ranging point refers to the spatial intersection point formed when the laser beam strikes the workpiece surface. Each laser sensor corresponds to one ranging point.
[0058] Furthermore, the local Jacobian matrix refers to a linear mapping matrix that describes the influence of the error at a certain laser ranging point on the robot's end-effector control. The local Jacobian matrix characterizes the mapping relationship between the distance error at that laser point and the six-dimensional velocity of the robot's end-effector.
[0059] In practice, the local Jacobian matrix can be calculated using the following formula:
[0060] ;
[0061] in, It is a local Jacobian matrix. It is the camera's focal length. The ratio to the size of camera pixels, For the distance data acquired by the i-th laser sensor, Let be the coordinates of the laser ranging point corresponding to the i-th laser sensor.
[0062] It should be noted that since the camera obtains physical coordinates, the formula calculates pixel coordinates. Each pixel has an actual width and height on the sensor. When converting physical millimeter coordinates to pixel coordinates, it is necessary to divide by the pixel size. Furthermore, because the pixels of the camera we use are square, with both length and width... Therefore, it can be converted into .
[0063] S105. Based on the local Jacobian matrices of all laser ranging points, determine the second Jacobian matrix of the distance features, and based on the first Jacobian matrix and the second Jacobian matrix, determine the joint feature Jacobian matrix.
[0064] In practice, the second Jacobian matrix can be calculated based on the following formula:
[0065] ;
[0066] in, This is the second Jacobian matrix. , , , These are the local Jacobian matrices corresponding to the four laser ranging points.
[0067] Furthermore, the joint feature Jacobian matrix is a composite Jacobian matrix formed by integrating the first and second Jacobian matrices according to certain rules, and is used to uniformly process the influence of image information and distance information on the robot.
[0068] In practice, the joint characteristic Jacobian matrix can be calculated based on the following formula:
[0069] ;
[0070] in, For the joint characteristic Jacobian matrix, This is the first Jacobian matrix. This is the second Jacobian matrix.
[0071] S106. Determine the feature point error matrix based on the pixel coordinates of the feature points in the visual image and the pixel coordinates of the feature points of the pre-assembly holes in the visual image of the curved workpiece acquired under the target pose.
[0072] Specifically, the target pose can be the current pose: the pixel coordinates of the current pre-assembled hole center point obtained by image processing after the camera acquires the image.
[0073] Furthermore, the height and width of the visual image captured by the camera can be determined by combining the camera's parameters. For example, in one embodiment, if the image captured by the camera is 320×200 pixels, the camera height is determined to be 320 pixels and the width to be 200 pixels.
[0074] In practice, the feature point error matrix can be calculated based on the following formula:
[0075] ;
[0076] in, The feature point error matrix, This represents the distance between pixel coordinates in the visual image and the pixel coordinates corresponding to the current pose, where row is the camera height and col is the camera width.
[0077] S107. Determine the pose error matrix based on the current pose and the target pose, and determine the overall error matrix based on the pose error matrix and the feature point error matrix.
[0078] Specifically, the current pose and the target pose include the robot's translation and rotation angle at two time points.
[0079] In practice, a pose error matrix can be constructed based on the current pose and the target pose, as shown below:
[0080] ;
[0081] in, Here is the pose error matrix. These represent the translation amounts in the three directions corresponding to the current pose. , These represent the translation amounts in the three directions corresponding to the target pose. , These are the rotation angles in the three directions corresponding to the current pose. , , The rotation angles in the three directions corresponding to the target pose.
[0082] Furthermore, the overall error matrix includes both spatial errors related to position and orientation offsets and visual errors related to the image plane.
[0083] In practice, the overall error matrix can be calculated based on the following formula:
[0084]
[0085] in, The overall error matrix, The feature point error matrix, Let be the pose error matrix.
[0086] S108. Input the overall error matrix and the joint feature Jacobian matrix into the pre-trained decision model, and let the decision model determine the current control quantity, and control the robot according to the current control quantity.
[0087] In practice, the overall error matrix, which comprehensively describes the distance between the current robot and the target state, and the joint feature Jacobian matrix, which describes the mapping relationship between the overall error and the robot's actions, are used as the state. The input is fed into a pre-trained deep reinforcement learning decision model. This model has learned how to make action decisions based on the state error through simulation training and can output a current control variable, that is, the direction and magnitude of the robot end effector's movement and rotation in six-dimensional space.
[0088] Furthermore, the robot adjusts its posture according to the current control variables determined by the decision-making model, gradually and precisely approaching the target position to complete the automatic pin insertion or positioning task on the complex curved surface workpiece.
[0089] The robot hybrid vision servo positioning method provided in this embodiment has two aspects. First, under the robot's current pose, it uses a camera to acquire images of the curved workpiece in real time, and simultaneously combines multiple laser rangefinders to obtain multi-point distance information from the end effector to the workpiece surface. This effectively integrates two-dimensional visual features with three-dimensional distance perception capabilities. Compared to traditional methods that rely solely on image processing, this multi-source perception significantly improves the accuracy of target detection and posture recognition on unstructured workpiece surfaces with significant curvature changes, enhancing the robustness and adaptability of the robot system in complex manufacturing scenarios. Second, through an image processing workflow, including image denoising, edge extraction, and fitting, it accurately identifies the geometric contours of pre-assembled holes in curved workpieces and obtains the image pixel coordinates of their centers. The first aspect is that the identification of circular holes provides a crucial planar reference for subsequent 3D positioning. This allows for identification even when the workpiece surface has certain reflections or changes in lighting, significantly improving the stability of circular hole identification and solving the problem of feature mismatch on complex curved surfaces in traditional visual servoing. Secondly, by combining the identified center pixel coordinates with depth data provided by multiple laser rangefinders, the 2D image features can be effectively converted into 3D spatial information, and an image feature Jacobian matrix (i.e., the first Jacobian matrix) can be constructed. This matrix reflects the influence of image feature errors on the robot's end effector speed, achieving precise modeling of the end effector control by the visual servoing system, thereby improving the responsiveness and stability of the robot's posture adjustment. Thirdly, for each laser sensor, the system... By combining the readings and installation posture, the mapping relationship between the spatial coordinate changes of the corresponding ranging point and the end effector velocity can be calculated, forming a local Jacobian matrix. This local Jacobian matrix provides feedback on the normal error between the end effector and the workpiece surface from multiple perspectives, enabling the robot not only to adjust its position but also to achieve high-precision posture adjustment, particularly suitable for normal correction in pre-assembly hole insertion tasks. Fifthly, by integrating the first Jacobian matrix of image features with the second Jacobian matrix of laser ranging, a joint feature Jacobian matrix (i.e., the joint feature Jacobian matrix) is constructed. This matrix simultaneously considers the control requirements of both position and planar normal dimensions, improving the system's control integrity for complex curved surface positioning and addressing the limitations of single-sensor methods in control. The limitations in dimensionality allow the robot to maintain stable movement while maintaining high precision. Sixthly, the overall error matrix and joint feature Jacobian matrix are input into a pre-trained TD3 deep reinforcement learning decision model to dynamically generate the current control quantity, guiding the robot to make real-time motion adjustments. This avoids the problems of manual parameter tuning and parameter fixing in traditional models, achieving a highly adaptive response to different workpiece postures, positions, and curvatures. This significantly improves the robot system's task success rate and generalization ability, meeting the needs of high-requirement assembly tasks such as automatic rivet insertion on complex curved surfaces. In summary, this significantly improves the robot's success rate and operational efficiency during assembly, ensuring that the workpiece can be accurately aligned perpendicularly with the insertion hole, achieving pin insertion.
[0090] Figure 3 This is a flowchart of Embodiment 2 of the robot hybrid vision servo positioning method provided in this application. Please refer to... Figure 3 Based on the above embodiments, the decision model is a reinforcement learning model based on a double Q network; the training steps of the decision model include:
[0091] S301. Construct a dual-Q network and initialize the parameters of the dual-Q network; wherein the dual-Q network includes a decision network for generating control variables and an evaluation network for evaluating the value of control variables.
[0092] Specifically, the dual Q network is one of the core structures of the reinforcement learning algorithm TD3. It consists of two sets of Q networks (i.e., value networks) used to evaluate the value of actions under the current policy, thereby alleviating the problem of Q function overestimation.
[0093] Furthermore, the decision network and the evaluation network operate in parallel, but with independent parameters.
[0094] It should be noted that the dual-Q network takes a smaller value during the update process to increase the conservatism of the valuation, thereby enhancing the stability of the strategy.
[0095] Furthermore, before training begins, the neural network can be initialized, including randomly setting network weights and biases, so that the strategy can be continuously optimized through gradient descent during subsequent training.
[0096] S302. Obtain multiple sets of experience data using simulation experiments; wherein, each set of experience data includes a sequence consisting of multiple state-action reward value pairs, and the overall reward value of the task corresponding to the sequence; the state data in each state-action reward value pair includes visual images acquired by the camera and distance data acquired by the laser sensor, the action data of the state-action reward value pair is the control variable, and the reward value of the state-action reward value pair is the single-step reward value corresponding to the control variable.
[0097] Specifically, experiential data refers to the training data collected by the robot through multiple interactions during simulation tasks. This data serves as the "learning material" for reinforcement learning and is stored in the "experience replay buffer."
[0098] In practice, experience data can include the robot's state, actions, rewards, and next state.
[0099] Furthermore, the state includes image information captured by the camera and distance information captured by the laser rangefinder; the action is the control quantity executed by the robot in the current state; the reward is the immediate feedback (single-step reward value) obtained by the robot after executing the action, which is used to measure the quality of the current action.
[0100] Furthermore, a complete task process is a sequence, and the overall reward value refers to the cumulative sum of the reward values of all individual steps during the complete execution of a task, reflecting the overall performance of the robot in completing the entire servo positioning task.
[0101] In specific implementation, for example, in one embodiment, a simulation environment containing a six-DOF robot, a monocular camera, and four laser rangefinders can be built in the CoppeliaSim simulation platform. By continuously running the robot's pin-insertion task, a large number of experience data sequences are collected. In each simulation experiment, the robot gradually adjusts its position and attitude from the initial pose. At each step, the system records: the current state (including visual image and laser ranging data), the action (control variable) executed, and the single-step reward value after the action is executed, forming a state-action-reward triplet. Multiple such triplets are connected in series to form a complete task sequence, and the overall reward value corresponding to the sequence is calculated to reflect the robot's performance throughout the task. All these experience data are stored in an "experience replay buffer" for reinforcement learning models (such as TD3) to repeatedly sample and train, thereby optimizing the control strategy and enabling the robot to more efficiently complete the servo positioning and pin-insertion operation of complex curved surface structures in subsequent tasks.
[0102] The following is a specific example illustrating the process of obtaining empirical data using simulation experiments:
[0103] (1) Take the robot’s initial pose as the current pose, collect the state data corresponding to the current pose, and input the state data into the decision network. The decision network outputs the control quantity corresponding to the state data.
[0104] Specifically, the robot's initial pose is initialized in the simulation environment as the current pose of the task. Then, the system collects the state data (including visual images and laser range values) under this pose and inputs it into the pre-trained decision network, which outputs the control quantity that should be executed at the moment.
[0105] (2) Control the robot to act according to the control quantity, and after the robot acts, collect the updated state data corresponding to the new pose, and determine the state type according to the updated state data; the state type includes task success state, task failure state and intermediate state.
[0106] Specifically, after the robot moves according to the control quantity and enters a new posture, the system then collects new state information and determines whether the current state is a successful state, a failed state, or an intermediate state based on the deviation from the target state.
[0107] (3) Determine the single-step reward value corresponding to the control quantity based on the state type.
[0108] Specifically, the single-step reward value corresponding to the action is calculated based on the state type, combined with factors such as motion efficiency and posture deviation.
[0109] In practice, the single-step reward value can be as follows:
[0110] ;
[0111] in, E represents the single-step reward value, and E represents the energy consumption index, which characterizes the energy consumption during the task. For speed smoothness, The maximum step size for the heavy task is given by ep, where the superscripts c and I represent respectively. This is the current state. This is the initial position.
[0112] The following is a specific embodiment that details the steps for determining the single-step reward value corresponding to the control quantity based on the state type:
[0113] Step 1: When the state type is task completed, calculate the expected path length of the robot from the initial pose to the target pose based on the initial joint angle of the robot in the initial pose and the target joint angle of the robot in the target pose.
[0114] Step 2: Determine the joint angle for each action based on the pose corresponding to the reward value of each state action in the current task.
[0115] Step 3: Determine the actual path length of the robot from the starting point to the ending point based on the joint rotation angles corresponding to the reward values of each state action.
[0116] Step 4: Determine the energy consumption index controlled in the current step based on the desired path length and the actual path length.
[0117] Step 5: Based on the joint rotation angles of the robot in the current pose, determine the initial motion operability index of the robot in the current step.
[0118] Step Six: Based on the initial motion operability index and all the initial motion operability indices of the robot in all steps before the current step, determine the actual motion operability index of the robot in the current step.
[0119] Step 7: Determine the single-step reward value corresponding to the control quantity based on the energy consumption index and the actual movement operability index.
[0120] In practice, the robot's joint angle configuration at the start of the task and the target joint angle at the end of the task are obtained, and the Euclidean distance between the two is calculated as the shortest path length under ideal conditions.
[0121] Furthermore, for each control action executed, the system records the current joint angle configuration of the robot; for each state-action pair in the entire task sequence, the corresponding actual joint angles are extracted to form a series of discrete points.
[0122] Furthermore, the joint space distance between adjacent frames is calculated step by step based on the joint angle sequence obtained in the previous step. The resulting path is the actual path length during the robot's actual execution process, which is usually greater than the expected path.
[0123] Furthermore, the energy consumption index can be calculated based on the following formula:
[0124] ;
[0125] Where E is the energy consumption index, abs() represents the absolute value, and q=[q1 q2 q3 q4 q5 q6] represents the joint values of a six-joint robot. Represents the robot's desired position. Represents the robot's initial position. This is the joint value at step k+1. Let ω be the joint value at step k. The factor ω is used as the weight of energy consumption for different joints, and its value is: ω=[4,4,2,2,1,1].
[0126] Furthermore, the operability of the movement It is a comprehensive measure of a robot's ability to move in all directions at a given pose. Motion maneuverability provides information from joint velocities. Speed to end effector The overall scalar description of the gain.
[0127] Furthermore, the actual motion operability index can be calculated based on the following formula:
[0128] ;
[0129] in, T represents the actual operability index of the exercise, E represents the exercise efficiency index, and E represents the energy consumption index.
[0130] Understandably, by using the initial motion operability index as a starting point and superimposing all the initial motion operability indices before the current step, the actual motion operability index at this point can be determined, representing the likelihood of the robot performing the task at this moment.
[0131] Furthermore, the joint angle differences between two adjacent steps are weighted and accumulated to obtain the actual joint space path length of the robot from the starting point to the ending point.
[0132] Optionally, when the state type is a task failure state, the single-step reward value corresponding to the control quantity is determined based on the feature point error matrix corresponding to the current pose and the feature point error matrix corresponding to the initial pose.
[0133] Specifically, the single-step reward value for task failure can be calculated based on the following formula:
[0134] ;
[0135] in, This is the single-step reward value for a failed task.
[0136] Optionally, when the state type is an intermediate state, the single-step reward value corresponding to the control quantity is determined according to the preset maximum number of movement steps.
[0137] In practice, the specific value of the maximum number of steps is set according to actual needs, and this embodiment does not limit it.
[0138] Furthermore, the negative value of the reciprocal of the maximum number of steps can be directly determined as the single-step reward value corresponding to the control quantity.
[0139] (4) Combine the state data, the control variable, and the single-step reward value into a state-action reward value pair.
[0140] In a specific implementation, for example, in one embodiment, state data A corresponds to control quantity a, the reward value of a is calculated as Ya, and the three quantities are concatenated to obtain the reward value pair Aa-Ya.
[0141] (5) When the status type is task success or task failure, the overall reward value of the current task is determined based on the single-step reward value of each state action reward value pair in the current task, and all state action reward value pairs in the current task and the overall reward value of the current task are saved as a piece of experience data.
[0142] In practice, the cumulative value of all single-step reward values in the current task represents the robot's overall performance from the start point to the end of the task (success or failure). Empirical data is an important indicator for measuring the quality of control strategies.
[0143] It should be noted that when a task is successful, all individual reward values are directly added together as the overall reward value. When a task fails, the individual reward values at the moment of failure can be directly used as the overall reward value to more intuitively reflect the task execution status.
[0144] (6) When the state type is intermediate state, the new pose is used as the current pose, and the process of collecting the state data corresponding to the current pose is executed again.
[0145] Specifically, if the current state is still intermediate, the new pose is taken as the new "current pose," and the process enters the next state acquisition and control decision-making loop, continuing until the task is completed or fails. This process is repeated continuously.
[0146] Understandably, starting from the robot's initial pose, the system acquires corresponding visual images and laser ranging data to form the current state, which is then input into a pre-trained decision network to generate control variables, driving the robot to complete one action. After the action is executed, the system collects updated state data corresponding to the new pose and determines the state type (success, failure, or intermediate state) based on the deviation from the task objective, thereby assigning a discriminative single-step reward to each control variable. Each state, action, and reward is packaged into a state-action-reward value pair. When the task ends (success or failure), the system accumulates all rewards throughout the entire task trajectory to calculate the overall reward and saves the complete trajectory and total reward of the task as a piece of experience data. If the task is not yet completed, the current updated state is used as a new starting point to continue the data acquisition and control loop. This achieves an automated and highly efficient experience data generation process, providing high-quality and diverse learning samples for the training of the dual-Q network based on the TD3 algorithm, thereby improving the robot's policy accuracy and robustness in complex curved surface localization scenarios.
[0147] S303. Based on the overall reward value of the task corresponding to the sequence in each set of experience data, select a portion of the experience data from the multiple sets of experience data as training samples, and use the training samples to train the dual-Q network.
[0148] Specifically, the overall reward value is obtained by summing up the reward values of all individual steps in a sequence. The higher the overall reward, the better the task is performed (which can be faster, more accurate, lower energy consumption, or closer to the target posture).
[0149] In practice, each sequence contains the robot's state, control actions, and corresponding single-step reward value during the alignment task. Then, the overall reward value is calculated for each sequence, which measures the overall performance of the task. Based on these overall reward values, the better-performing sequences are selected from all empirical data as training samples. The state-action-reward pairs in these samples are input into the dual-Q network of the TD3 algorithm to guide the two evaluation networks in learning and updating the action value, thereby gradually improving the robot's motion strategy in complex curved surface servo localization tasks.
[0150] The following is a specific example illustrating the process of selecting a subset of empirical data from multiple sets of empirical data as training samples:
[0151] Sort all experience data from highest to lowest according to the overall reward value of the task corresponding to that experience data, and obtain the sorting result;
[0152] Based on the preset screening ratio parameter and the amount of the empirical data, the number of screening samples is determined, and the first target sample corresponding to the number of screening samples is selected from the sorting results in descending order.
[0153] From the remaining samples, empirical data whose overall reward value falls within a preset range are selected as supplementary samples; wherein, the upper limit of the preset range is the minimum overall reward value corresponding to the first target sample, and the lower limit of the preset range is the difference between the minimum overall reward value and the preset tolerance parameter;
[0154] The target sample and the supplementary sample are determined as training samples.
[0155] In practice, based on simulation experiments, multiple complete sets of experience data are stored. Each set of data corresponds to a robot task execution process, including multiple state-action-reward value pairs and the overall reward value of the task, i.e. the total score of task performance. First, all experience data are sorted from high to low according to the overall reward value to form an ordered experience sequence. The experience data at the top indicates that the robot performs better in these tasks, with shorter execution paths, more stable control, and smaller errors.
[0156] Furthermore, based on the preset screening ratio parameters (e.g., taking the top 10% or 20%) and the total amount of experience data, the number of the first target samples to be retained is calculated, and a corresponding number of high-quality experiences are selected from front to back according to the sorting order as the first batch of training samples. These data represent positive examples of "good control behavior" in strategy learning.
[0157] Furthermore, from the remaining unselected samples, some "suboptimal experiences" are further screened as supplementary samples. In specific implementation, for example, in one embodiment, a preset tolerance parameter (such as the reward fluctuation range ε) can be set. The minimum overall reward value in the first target sample is used as the upper limit, and the tolerance parameter is subtracted from it as the lower limit to form an allowable interval. The experience data within this interval is added as supplementary samples. In this way, samples that have certain value but did not enter the first target sample are retained through the allowable interval. These supplementary samples can increase the adaptability of the strategy to boundary states and non-optimal decisions.
[0158] Furthermore, the first target sample and the supplementary sample are merged to form the training sample set for the current round of the reinforcement learning model, which is then fed into the double-Q network for policy updates.
[0159] Understandably, by using the method of "selecting the best from the best and supplementing the boundaries", the quality of training data is ensured while retaining a certain degree of strategy diversity, which significantly improves the robot's learning speed, convergence stability and generalization ability in complex surface localization tasks.
[0160] The robot hybrid vision servo localization method provided in this embodiment firstly constructs a dual-Q network structure, setting up two independent evaluation networks to estimate the value of the same action, thus solving the problem of overestimation of action value in traditional single-Q networks, thereby improving the accuracy of the model's judgment on action quality and the stability of the control strategy. Secondly, through extensive experiments in a simulation environment, multiple sets of empirical data containing image and distance information are collected. Each set of data completely covers the entire process of the robot from the initial pose to the target pose, providing a rich and realistic perception-decision-feedback link for learning precise control strategies. Finally, by evaluating the overall reward value of the task, high-quality empirical sequences are prioritized for training, avoiding interference from inefficient data and accelerating the convergence speed of the strategy. In summary, through the dual-Q network architecture + high-quality simulation experience collection + sample selection strategy based on overall reward, high-precision, adaptive, and robust positioning control of the robot for pre-assembled holes on complex curved surfaces is achieved, significantly improving task completion efficiency and success rate.
[0161] The experimental verification process is given below:
[0162] Specifically, a virtual environment was built on the CoppeliaSim simulation platform for training. Since the CoppeliaSim software's built-in robot library did not include the ABB IRB 120 robot, the URDF file of the selected robot was first imported. Simultaneously, the absolute positions of the robot's DAE and STL files were modified to ensure subsequent collision and kinematic simulations could be performed. A vision camera and four laser rangefinders were imported into the robot's end effector. The cameras were arranged in an "eye-on-hand" layout with a resolution of 320×240 (the cameras were directly mounted on the robot's arm end effector (usually near the end effector) and moved with the arm). Ignoring depth information was selected to ensure that the depth information provided by the laser rangefinders was used during positioning. The laser rangefinders were evenly distributed around the end effector of the robotic arm, as close to the center as possible to ensure coverage of the smallest possible area, thereby more accurately calculating the normal to the plane containing the target point. Simultaneously, the workpiece was imported into the simulation environment. The workpiece was 15mm thick, with a single-curvature surface and a through hole at the top for subsequent rivet insertion; the hole's direction was perpendicular to the surface.
[0163] Furthermore, the algorithm was run on a Dell Precision-5820-Tower workstation with an Intel Xeon(R)w-2255 processor (3.7 GHz), 96 GB of RAM, an NVIDIA RTX 3080 graphics card, Ubuntu 20.04 operating system, and Anaconda Python 3.6.15 and PyTorch 2.1.0. During training, the location of the artifacts was randomly refreshed each time to ensure the generalization ability and effectiveness of the training. The finally trained model was saved as a .pth file, and the relevant training data was also stored in the log.
[0164] Furthermore, the CatalystRL framework employed is a distributed framework for reproducible reinforcement learning research, primarily consisting of three components: a sampler, a trainer, and a database. The sampler is responsible for acquiring states and observations from the environment, determining the next action based on the current model, and then applying these actions to the environment to obtain new states and rewards. This data is stored in a trajectory buffer. The trainer uses various optimization algorithms to adjust model parameters, optimize the model's decision-making process, and manages experience data using a replay buffer and a queue system, thereby efficiently utilizing historical data for learning. The database is used for persistent storage and querying of trajectory data generated during training; this study uses MongoDB for efficient data management and access. Overall, this architecture, through these three cooperating components, achieves efficient training and parameter optimization of reinforcement learning models in a continuous state space.
[0165] Furthermore, the TD3 algorithm adopts a two-layer "Actor-Critic" structure. Both the Actor and Critic contain two hidden layers, each with 256 neurons. Each Actor network corresponds to two Critic networks. The learning rate is set to 0.0003, the activation function is ReLU, and the batch size is 16.
[0166] Furthermore, after successful training, data from the MongoDB system was analyzed. It can be seen that the algorithm descents relatively quickly with minimal fluctuations, reaching convergence in approximately 7000 steps, indicating that the robot has learned a relatively reasonable strategy.
[0167] Furthermore, the reward function during training showed almost no overexploration during convergence, and its growth was relatively stable. After 15,000 steps, it could be considered to be basically complete in terms of pre-assembly hole insertion, indicating that the algorithm had good performance.
[0168] Furthermore, after the number of steps reached 2500, the robot's motion efficiency remained around 2.8, indicating that the robot could perform rapid task execution actions within the action space, reducing motion time and accelerating the training process.
[0169] It is understandable that the above experimental results show that the proposed algorithm has a short training time, and the trained robot can quickly perform the nailing action for the designed pre-assembled holes, and still has a good success rate and efficiency when the workpiece position changes.
[0170] Another embodiment is given below for different experimental verification:
[0171] Furthermore, the thickness of the workpiece was changed to 10mm and 20mm respectively, and imported into the simulation environment. First, the pre-trained model was loaded and 50 positioning experiments were conducted. During the process, the position of the workpiece was still randomly generated. Then, the algorithm was changed to PBVS and IBVS methods, and 50 experiments were also conducted on a workpiece with a thickness of 15mm.
[0172] Furthermore, the model achieves the highest success rate of 80% when the workpiece thickness is at its thinnest. This is because the thinner the workpiece, the closer the target point is to the workpiece surface, making it easier for the rivet's base to reach the target location. Even with slight deviations in the normal angle, the gap between the rivet and the hole ensures successful positioning. Although the success rate decreases with increasing workpiece thickness, the model continues to learn and train, maintaining a 74% success rate for positioning a 20mm thick workpiece in 50 experiments. This demonstrates the model's accuracy and versatility.
[0173] Furthermore, the BVS method has a low success rate because the workpiece position is randomly generated each time, and the feature points may disappear from the field of view during the movement, leading to positioning failure. The IBVS method is also prone to robot speed exceeding the limit due to large distances, resulting in a low success rate. The proposed method has obvious advantages.
[0174] Furthermore, since the workpiece only changed position during the training process, and the hole was always at the top, meaning the hole was always perpendicular to the ground, in order to verify the generalizability of the algorithm, the position and orientation of the hole on the workpiece were redesigned while keeping the workpiece thickness at 15mm and the hole orientation perpendicular to the curved surface. The results were then imported into the simulation environment and 50 experiments were conducted.
[0175] Furthermore, the algorithm still basically meets the positioning requirements even after the orientation of the hole changes. As the model continuously trains itself during the positioning process, the success rate is relatively high, indicating that the algorithm has a certain degree of generalization and is still meaningful when facing holes with different orientations or even workpieces with different curvatures.
[0176] Corresponding to the aforementioned embodiment of a robot hybrid vision servo positioning method, this application also provides an embodiment of a robot hybrid vision servo positioning device.
[0177] An embodiment of the robot hybrid vision servo positioning device disclosed in this application can be applied to a robot hybrid vision servo positioning device. The device embodiment can be implemented through software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of the robot hybrid vision servo positioning device loading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 4 The diagram shown is a hardware structure diagram of a robot hybrid vision servo positioning device, which is the subject of this application. (Except for...) Figure 4 In addition to the processor, memory, network interface, and non-volatile memory shown, the robot hybrid vision servo positioning device in the embodiment may also include other hardware depending on the actual function of the robot hybrid vision servo positioning device, which will not be described in detail here.
[0178] Figure 5 This is a schematic diagram of the structure of an embodiment of the robot hybrid vision servo positioning device provided in this application. Please refer to... Figure 5 The device provided in this embodiment includes a data acquisition module 510, a determination module 520, and a control module 530; wherein,
[0179] The acquisition module 510 is used to acquire visual images of the curved workpiece based on the camera in the current pose, and to acquire distance data from the robot end to the curved workpiece based on the multiple laser sensors.
[0180] The determining module 520 is used to identify feature points of the pre-assembly hole of the curved workpiece from the visual image, and determine the pixel coordinates of the center of the pre-assembly hole based on the feature points.
[0181] The determining module 520 is further configured to determine the depth information from the robot end to the curved workpiece based on the distance data, and construct a first Jacobian matrix of image features based on the pixel coordinates and the depth information;
[0182] The determining module 520 is further configured to determine, for each laser sensor, the local Jacobian matrix of the laser ranging point corresponding to that laser sensor based on the distance data acquired by that laser sensor;
[0183] The determining module 520 is further configured to determine a second Jacobian matrix of distance features based on the local Jacobian matrices of all laser ranging points, and to determine a joint feature Jacobian matrix based on the first Jacobian matrix and the second Jacobian matrix.
[0184] The determining module 520 is further configured to determine a feature point error matrix based on the pixel coordinates of the feature points in the visual image and the pixel coordinates of the feature points of the pre-assembly holes in the visual image of the curved workpiece acquired under the target pose.
[0185] The determining module 520 is further configured to determine a pose error matrix based on the current pose and the target pose, and to determine an overall error matrix based on the pose error matrix and the feature point error matrix.
[0186] The control module 530 is used to input the overall error matrix and the joint feature Jacobian matrix into a pre-trained decision model, and the decision model determines the current control quantity and controls the robot according to the current control quantity.
[0187] The apparatus of this embodiment can be used to perform... Figure 1 The steps of the method embodiment shown are similar in principle and process, and will not be repeated here.
[0188] Please continue to refer to Figure 4 This application also provides a robot hybrid vision servo positioning device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the methods provided in the first aspect of this application.
[0189] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods provided in this application.
[0190] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0191] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0192] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A robot hybrid vision servo localization method, characterized in that, The method is applied to a robot, the end effector of which is equipped with a camera and multiple laser sensors; the method includes: In the current pose, visual images of the curved workpiece are acquired based on the camera, and distance data from the robot end effector to the curved workpiece is acquired based on the multiple laser sensors. Feature points of the pre-assembly hole of the curved workpiece are identified from the visual image, and the pixel coordinates of the center of the pre-assembly hole are determined based on the feature points. Based on the distance data, the depth information from the robot end effector to the curved workpiece is determined, and a first Jacobian matrix of image features is constructed based on the pixel coordinates and the depth information; For each laser sensor, the local Jacobian matrix of the laser ranging point corresponding to that laser sensor is determined based on the distance data acquired by that laser sensor. Based on the local Jacobian matrices of all laser ranging points, determine the second Jacobian matrix of the distance features, and based on the first and second Jacobian matrices, determine the joint feature Jacobian matrix; The feature point error matrix is determined based on the pixel coordinates of the feature points in the visual image and the pixel coordinates of the feature points of the pre-assembly holes in the visual image of the curved workpiece acquired under the target pose. Based on the current pose and the target pose, a pose error matrix is determined, and based on the pose error matrix and the feature point error matrix, an overall error matrix is determined. The overall error matrix and the joint feature Jacobian matrix are input into a pre-trained decision model, which determines the current control quantity and controls the robot according to the current control quantity.
2. The method according to claim 1, characterized in that, The decision model is a reinforcement learning model based on a double-Q network; the training process of the decision model includes: Construct a dual-Q network and initialize the parameters of the dual-Q network; wherein the dual-Q network includes a decision network for generating control variables and an evaluation network for evaluating the value of control variables; Multiple sets of empirical data are obtained through simulation experiments. Each set of empirical data includes a sequence of multiple state-action reward value pairs and the overall reward value of the task corresponding to the sequence. The state data in each state-action reward value pair includes visual images acquired by the camera and distance data acquired by the laser sensor. The action data of the state-action reward value pair is the control variable, and the reward value of the state-action reward value pair is the single-step reward value corresponding to the control variable. Based on the overall reward value of the task corresponding to the sequence in each set of experience data, a portion of experience data is selected from the multiple sets of experience data as training samples, and the training samples are used to train the dual-Q network.
3. The method according to claim 2, characterized in that, The acquisition of empirical data through simulation experiments includes: The robot's initial pose is taken as the current pose. The state data corresponding to the current pose is collected and input into the decision network. The decision network outputs the control quantity corresponding to the state data. The robot is controlled to move according to the control quantity, and after the robot moves, the updated state data corresponding to the new pose is collected, and the state type is determined according to the updated state data; the state type includes task success state, task failure state and intermediate state; Determine the single-step reward value corresponding to the control quantity based on the state type; The state data, the control variable, and the single-step reward value are concatenated into a state-action-reward value pair; When the status type is task success or task failure, the overall reward value of the current task is determined based on the single-step reward value of each state action reward value pair in the current task, and all state action reward value pairs in the current task, as well as the overall reward value of the current task, are saved as a piece of empirical data. When the state type is intermediate, the new pose is used as the current pose, and the process of collecting the state data corresponding to the current pose is executed again.
4. The method according to claim 3, characterized in that, Determining the single-step reward value corresponding to the control quantity based on the state type includes: When the state type is task completed, the expected path length of the robot from the initial pose to the target pose is calculated based on the initial joint angle of the robot in the initial pose and the target joint angle of the robot in the target pose. Based on the pose corresponding to the action reward value of each state in the current task, determine the joint angle under each action step; Based on the joint rotation angles corresponding to the reward values of each state action, determine the actual path length of the robot from the starting point to the ending point; The energy consumption index for the current step is determined based on the expected path length and the actual path length. Based on the joint rotation angles of the robot in its current pose, determine the initial motion operability index of the robot in the current step. Based on the initial motion operability index and all the initial motion operability indices of all steps controlled before the current step, the actual motion operability index of the robot in the current step is determined. The single-step reward value corresponding to the control quantity is determined based on the energy consumption index and the actual motion operability index.
5. The method according to claim 4, characterized in that, Determining the single-step reward value corresponding to the control quantity based on the state type includes: When the state type is a task failure state, the single-step reward value corresponding to the control quantity is determined based on the feature point error matrix corresponding to the current pose and the feature point error matrix corresponding to the initial pose.
6. The method according to claim 5, characterized in that, Determining the single-step reward value corresponding to the control quantity based on the state type includes: When the state type is an intermediate state, the single-step reward value corresponding to the control quantity is determined according to the preset maximum number of movement steps.
7. The method according to claim 3, characterized in that, Based on the overall reward value of the task corresponding to the sequence in each set of experience data, a portion of experience data is selected as training samples from the multiple sets of experience data, including: Sort all experience data from highest to lowest according to the overall reward value of the task corresponding to that experience data, and obtain the sorting result; Based on the preset screening ratio parameter and the amount of the empirical data, the number of screening samples is determined, and the first target sample corresponding to the number of screening samples is selected from the sorting results in descending order. From the remaining samples, empirical data whose overall reward value falls within a preset range are selected as supplementary samples; wherein, the upper limit of the preset range is the minimum overall reward value corresponding to the first target sample, and the lower limit of the preset range is the difference between the minimum overall reward value and the preset tolerance parameter; The target sample and the supplementary sample are determined as training samples.
8. A robot hybrid vision servo positioning device, characterized in that, The device is applied to a robot, the robot's end effector being equipped with a camera and multiple laser sensors; the device includes a data acquisition module, a determination module, and a control module; wherein, The acquisition module is used to acquire visual images of the curved workpiece based on the camera in the current pose, and to acquire distance data from the robot end effector to the curved workpiece based on the multiple laser sensors. The determining module is used to identify feature points of the pre-assembly hole of the curved workpiece from the visual image, and determine the pixel coordinates of the center of the pre-assembly hole based on the feature points. The determining module is further configured to determine the depth information from the robot end effector to the curved workpiece based on the distance data, and construct a first Jacobian matrix of image features based on the pixel coordinates and the depth information; The determining module is further configured to, for each laser sensor, determine the local Jacobian matrix of the laser ranging point corresponding to that laser sensor based on the distance data acquired by that laser sensor; The determining module is further configured to determine the second Jacobian matrix of the distance features based on the local Jacobian matrices of all laser ranging points, and to determine the joint feature Jacobian matrix based on the first Jacobian matrix and the second Jacobian matrix. The determining module is further configured to determine the feature point error matrix based on the pixel coordinates of the feature points in the visual image and the pixel coordinates of the feature points of the pre-assembly holes in the visual image of the curved workpiece acquired under the target pose. The determining module is further configured to determine a pose error matrix based on the current pose and the target pose, and to determine an overall error matrix based on the pose error matrix and the feature point error matrix. The control module is used to input the overall error matrix and the joint feature Jacobian matrix into a pre-trained decision model, which determines the current control quantity and controls the robot according to the current control quantity.
9. A robot hybrid vision servo positioning device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the program, implements the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Mechanical arm control method based on deep reinforcement learning
CN116533249A
Mixed mode depth imaging
CN117043547A
Method and system for positioning angle of strip head of inner ring of steel coil
CN117329991A
Robot 3D visual servo method based on composite learning and homography matrix
CN117733868A
Self-calibration system of rigid connecting rod-flexible continuum hybrid mechanical arm and implementation method of self-calibration system
CN118752483A
Cited By
Robot measurement method and system with offline path planning and online visual correction
CN122617984A