Space pose correction method and device for automatic locking dexterous hand end effector
By using online vision correction methods and devices, the problems of grasping error and high hardware cost of dexterous hands in industrial automatic locking systems have been solved, realizing low-cost, easy-to-change flexible locking that can meet the needs of multi-variety, small-batch production.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LINGXIN QIAOSHOU (BEIJING) TECH CO LTD
- Filing Date
- 2026-06-16
- Publication Date
- 2026-07-14
AI Technical Summary
Existing industrial automatic locking systems cannot simultaneously achieve compatibility with general tools, gripping posture accuracy, and changeover efficiency. They suffer from poor flexibility, high costs, or complex control, making it difficult to adapt to the flexible production needs of multiple varieties and small batches.
The screwdriver is placed in a preset calibration position by a robotic arm-driven dexterous hand. Multiple frames of screwdriver bit images are captured by a camera device, and pixel-level center coordinates are extracted. Based on a preset model, deviation compensation is performed and mapped to the physical displacement in the robotic arm's base coordinate system. This achieves the coincidence of the screwdriver bit center with the theoretical working point, completes the pose correction, and performs the screw fastening operation.
It achieves precision compensation for dexterous hand gripping of general-purpose tools, reduces hardware costs, simplifies control logic, improves changeover efficiency, and supports flexible configuration according to workpiece precision level, balancing efficiency and quality.
Smart Images

Figure CN122378764A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and industrial automation technology, and more specifically to an online visual correction method and apparatus for the spatial pose of a dexterous hand end effector used in automatic locking. Background Technology
[0002] Currently, the mainstream industrial automatic screw fastening systems fall into two main technical categories: dedicated tightening modules and collaborative robots paired with customized end effectors. Dedicated tightening modules integrate servo motors, torque sensors, and precision guide rails, relying on a rigid structure to ensure fastening positioning accuracy. However, they are only compatible with screws and workpieces of fixed specifications, resulting in limited versatility and high equipment costs. Collaborative robots equipped with customized gripping actuators can perform screw gripping and fastening operations, but when changing product models, the entire end effector structure needs to be replaced, leading to lengthy debugging times.
[0003] Existing fastening devices are all custom-developed for specific working conditions and cannot be adapted to general-purpose electric screwdrivers, resulting in high hardware costs. Dexterous hands exhibit random posture deviations when grasping general-purpose tools, hindering their flexibility in precision assembly. Visual servo tracking control logic is cumbersome, and multi-degree-of-freedom collaborative control can lead to insufficient system stability. Currently, there is no mature application model combining dexterous hands with general-purpose tools and supporting precision compensation, making it difficult to balance end effector costs and changeover efficiency. Existing technologies are ill-suited to the requirements of flexible production operations with diverse product types and small batches.
[0004] The information in the background section is merely intended to illustrate the general background of the invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art. Summary of the Invention
[0005] To address at least some of the technical problems in the prior art, the present invention provides a method and apparatus for online visual correction of the spatial pose of a dexterous hand end effector for automatic locking. Specifically, the present invention includes the following:
[0006] A first aspect of the present invention provides an online visual correction method for the spatial pose of a dexterous hand end effector for automatic locking, the method comprising:
[0007] A screwdriver is placed at a preset calibration station by a robotic arm that drives a dexterous hand; wherein the preset calibration station is located at a preset position relative to the camera device, the robotic arm is a multi-degree-of-freedom robotic arm, and the dexterous hand is a multi-finger dexterous hand;
[0008] The camera device captures multiple frames of screwdriver bit images;
[0009] The multi-frame bit images are input into a preset model to extract the pixel-level center coordinates of the multi-frame bit images;
[0010] The pixel deviation to be compensated is obtained based on the pixel-level center coordinates of the multi-frame header and the preset reference center;
[0011] Based on preset rules, the pixel deviation that needs to be compensated is mapped to the physical displacement in the robot arm's base coordinate system;
[0012] The robotic arm is controlled to perform a small translational motion in the XY plane according to the physical displacement, so as to drive the dexterous hand and the screwdriver it grasps to move synchronously until the center of the screwdriver bit coincides with the theoretical working point, thus completing the posture correction;
[0013] Maintain the corrected position and perform the screw fastening operation.
[0014] Optionally, acquiring multiple frames of screwdriver bit images via the camera device includes: using the dexterous hand to rotate the screwdriver at a uniform angular velocity of... The screwdriver is rotated, and multiple frames of bit images of the screwdriver are captured by the camera device located above the screwdriver. The unit is rad / s;
[0015] Inputting the multi-frame head images into a preset model to extract the pixel-level center coordinates of the multi-frame head includes: the camera frame rate of the camera device is... ,in, That is, the number of frames collected per second, used to determine the sampling time interval. and total number of sampled frames ,in The rotation period;
[0016] No. Frame acquisition time The corresponding rotation angle is ;in, The initial rotation phase angle, in rad, is determined by the random initial angle of the bit at the first sampling moment. After Gaussian filtering and contrast enhancement through contrast-limited adaptive histogram equalization, the multi-frame bit images are input into the preset model to obtain the detection box. Then, the pixel-level bit center coordinates of the multi-frame bit images are calculated using the gray-scale centroid method. Unit: pixel; wherein, the detection box is used to define the position of the batch head in the image; In the horizontal direction, It is in the vertical direction.
[0017] Optionally, the preset reference center includes the ROI box reference center, and the pixel deviation to be compensated is obtained based on the multi-frame header pixel-level center coordinates and the preset reference center, including:
[0018] Calculate the observed pixel deviation between the pixel-level center coordinates of the multi-frame batch header and the reference center of the ROI box;
[0019] The observed pixel deviation is estimated by least squares based on the eccentricity error model to eliminate the periodic error introduced by the eccentricity of the bit manufacturing, and the true system deviation is obtained. The true system deviation is used as the pixel deviation that needs to be compensated.
[0020] Optionally, the observed pixel deviation is estimated using a least-squares method based on the eccentricity error model to eliminate the periodic error introduced by the eccentricity in the bit manufacturing process, thus obtaining the true system deviation, which includes:
[0021] Let the eccentricity vector between the geometric center and the rotation center of the bit be... Unit: pixel; Module length is Unit: pixel; Direction angle: Unit: rad, number Frame observation pixel deviation Includes true system bias Eccentricity error and random noise :
[0022]
[0023]
[0024] in, The first Frame in Pixel deviation in direction, unit: pixel; : True systematic deviation in direction, unit: pixel; The first Frame in Random measurement noise in direction, unit: pixel;
[0025] make , ,in, For the eccentric vector in Projected components of the axis, unit: pixel. For the eccentric vector in Projected components of the axis, unit: pixel; construct a linear least squares model. ,
[0026] in:
[0027]
[0028] Represents the observed pixel vector, unit: pixel, by Frame pixel deviations are formed by stacking them in frame order. 3D column vector;
[0029]
[0030] in, Design matrix, dimensions Due to the rotation angle Construct a known coefficient matrix;
[0031] : The vector of states to be estimated, dimension It includes true system bias and eccentricity error;
[0032]
[0033] Noise vector, dimension Assuming each component is independent and identically distributed, with variance , ;
[0034] The optimal estimate is:
[0035]
[0036] : The least squares estimate, dimension ;
[0037] Extracting the true system bias estimate , and eccentricity parameters ,in , For estimating vectors The first and second components.
[0038] The above Eccentricity estimate, unit: pixel, derived from the estimated projection components. The calculated eccentric vector magnitude.
[0039] Optionally, data is collected during the rotation process. Frame image ( ), No. The deviation of the observed pixels obtained from frame detection is By using least squares estimation, the estimated covariance matrix is output while eliminating eccentricity errors. , used to assess the reliability of the estimate; where, Random measurement noise variance, unit: pixel²; when the standard deviation and One of them exceeds the threshold When this happens, the number of sampling frames is automatically increased and the estimation is recalculated; among them, Estimate the covariance matrix, dimension This reflects the uncertainty of each parameter to be estimated; : The estimated standard deviation of the true systematic bias in direction, in pixels, reflects... Reliability of the estimates; : The estimated standard deviation of the true systematic bias in direction, in pixels, reflects... Reliability of the estimates; Convergence criterion threshold (convergence criterion threshold, unit: pixel, determined by back-calculation based on the desired physical accuracy, such as...) (pixel corresponds to a physical precision of approximately ±0.15mm).
[0040] Optionally, this method also includes: a convergence criterion, which determines that the estimate is reliable and compensation can be performed when the following conditions are met simultaneously:
[0041]
[0042] If the requirements are not met, increase the number of sampling frames. ;in: This refers to the number of sampling frames added each time, typically 5 to 10 frames. The larger the value, the smaller the covariance of the least squares estimate, and the more reliable the estimate.
[0043] The above steps, based on preset rules, map the pixel deviations that need to be compensated into physical displacements in the robot arm's base coordinate system, including:
[0044] Let the camera intrinsic parameter matrix of the imaging device be:
[0045]
[0046] in, The camera is Focal length in a direction, unit: pixel; : Camera principal point coordinates, unit: pixel, i.e., the coordinates of the intersection of the camera optical axis and the image plane;
[0047] The reference center point of the ROI in the pixel coordinate system is The estimated true system deviation is ;
[0048] Considering lens radial distortion, define the normalized plane radius:
[0049]
[0050] in Average focal length, unit: pixel; Normalized plane radius, dimensionless, represents the normalized distance of a pixel relative to the principal point, used for distortion compensation calculation;
[0051] Pixel deviation after distortion compensation:
[0052]
[0053]
[0054] The basic scale is obtained through calibration. ,in Unit: mm / pixel , This represents the radial distortion coefficient of the lens; it takes into account the actual distance between the camera and the bit. Distance from calibration Differences, adjust the scale:
[0055]
[0056] in, and The unit is mm;
[0057] The compensated displacement in the base coordinate system of the robotic arm is:
[0058] .
[0059] Optionally, the method further includes:
[0060]
[0061] At that time, it is determined to be "aligned";
[0062] in, Pixel error tolerance, unit: pixel, set according to workpiece accuracy requirements;
[0063] High-precision mode: pixel pixel corresponds to a physical precision of approximately ±0.1~0.2mm;
[0064] Standard mode: pixel pixel corresponds to a physical precision of approximately ±0.3~0.5mm;
[0065] Quick Mode: pixel Pixel priority is given to ensuring the cycle time of the operation.
[0066] Optionally, the method further includes:
[0067] The displacement is received by the robotic arm. , );
[0068] Combining the repeatability of the robotic arm Total compensation error budget:
[0069]
[0070] Total compensation error, in mm, is a composite of visual estimation error, scale mapping error, and robotic arm execution error. They are respectively Physical error of orientation visual estimation error mapping, unit: mm.
[0071] Optionally, controlling the robotic arm to perform micro-translational motion in the XY plane according to the physical displacement includes:
[0072] This correction mechanism only applies to the XY plane, and the rotation angle... Keep the preset value unchanged:
[0073]
[0074] Compensation transformation matrix, dimension , used to describe pure translation compensation transformation in the base coordinate system of the robotic arm; : Identity matrix.
[0075] Optionally, the preset model is trained through the following steps:
[0076] Multiple screwdriver bit images were collected as training samples;
[0077] The training samples are input into the YOLOv11 object detection network for training, so that it learns the visual features of the bit in the actual scene;
[0078] Perform intrinsic parameter calibration and pixel and physical scale calibration on the camera used to acquire images;
[0079] The screwdriver bit image is input into the visual calibration and algorithm preprocessing module to obtain the preset model and calibration parameters.
[0080] A second aspect of the present invention provides an online visual correction device for the spatial pose of a dexterous hand end effector for automatic locking, the device comprising:
[0081] A placement module is used to place a screwdriver at a preset calibration station by driving a dexterous hand through a robotic arm; wherein the preset calibration station is located at a preset position relative to the camera device, the robotic arm is a multi-degree-of-freedom robotic arm, and the dexterous hand is a multi-finger dexterous hand;
[0082] The acquisition module is used to acquire multiple frames of screwdriver bit images through the camera device;
[0083] The extraction module is used to input the multi-frame head images into a preset model to extract the pixel-level center coordinates of the multi-frame head;
[0084] The first processing module is used to obtain the pixel deviation to be compensated based on the pixel-level center coordinates of the multi-frame header and the preset reference center.
[0085] The second processing module is used to map the pixel deviation to be compensated into a physical displacement in the base coordinate system of the robotic arm based on a preset rule.
[0086] The third processing module is used to control the robotic arm to perform a micro-translational motion in the XY plane according to the physical displacement, so as to drive the dexterous hand and the screwdriver it grasps to move synchronously until the center of the screwdriver bit coincides with the theoretical working point, thus completing the posture correction.
[0087] The screw fastening module is used to maintain the corrected position and perform screw fastening operations.
[0088] A third aspect of the present invention provides an electronic device comprising:
[0089] Processor; and
[0090] Stored program memory,
[0091] The program includes instructions that, when executed by the processor, cause the processor to perform the method described in any of the first aspect embodiments described above.
[0092] A fourth aspect of the present invention provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are configured to cause a computer to perform the method described in any one of the embodiments of the first aspect above.
[0093] In a fifth aspect, the present invention provides a computer program product comprising instructions which, when executed, cause a computer to perform the method described in any one of the embodiments of the first aspect.
[0094] This invention provides an online visual correction method for the spatial pose of a dexterous hand end effector used in automatic screw fastening. The method includes: using a robotic arm to drive the dexterous hand to place a screwdriver at a preset correction station; wherein the preset correction station is located at a preset position relative to a camera device, the robotic arm is a multi-degree-of-freedom robotic arm, and the dexterous hand is a multi-finger dexterous hand; acquiring multiple frames of screwdriver bit images using the camera device; inputting the multiple frames of screwdriver bit images into a preset model to extract the pixel-level center coordinates of the screwdriver bit; obtaining the pixel deviation to be compensated based on the pixel-level center coordinates of the screwdriver bit and a preset reference center; mapping the pixel deviation to be compensated into a physical displacement in the robotic arm's base coordinate system based on preset rules; controlling the robotic arm to perform a micro-translational motion in the XY plane according to the physical displacement, so as to drive the dexterous hand and the screwdriver it grasps to move synchronously until the center of the screwdriver bit coincides with the theoretical working point, thus completing the pose correction; maintaining the corrected pose and performing the screw fastening operation. This solution addresses the problems of existing industrial automated clamping systems, which cannot balance compatibility with general-purpose tools, gripping posture accuracy, and changeover efficiency, resulting in poor flexibility, high cost, or complex control. By employing an "online correction after gripping" mechanism and a YOLOv11 vision inspection closed loop, the random posture error of the dexterous hand gripping a general-purpose tool is transformed into a compensable physical displacement, overcoming the limitations of repeatability accuracy of the dexterous hand. At the same time, software-defined accuracy replaces hardware stacking, achieving low-cost, easy-to-change, flexible clamping, and supporting flexible configuration of tolerance according to the workpiece accuracy level, balancing efficiency and quality.
[0095] It should be understood that the above general description and the following specific embodiments are merely exemplary and illustrative, and do not limit the scope of the invention. Attached Figure Description
[0096] The accompanying drawings, which are part of the specification of this invention, illustrate exemplary embodiments of the invention. The drawings, together with the description in the specification, serve to illustrate the principles of the invention.
[0097] Figure 1 The flowchart illustrates an online visual correction method for the spatial pose of a dexterous hand end effector used for automatic locking, provided as a specific embodiment of the present invention.
[0098] Figure 2 This is a schematic diagram of an online visual correction device for the spatial pose of a dexterous hand end effector used for automatic locking, provided as a specific embodiment of the present invention.
[0099] Figure 3 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0100] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the spirit of the contents disclosed in the present invention will be clearly explained below with reference to the accompanying drawings and detailed description. After understanding the embodiments of the present invention, any person skilled in the art can make changes and modifications based on the technology taught in the present invention without departing from the spirit and scope of the present invention.
[0101] The illustrative embodiments and descriptions of the present invention are used to explain the invention, but are not intended to limit the invention. Furthermore, elements / components using the same or similar reference numerals in the drawings and embodiments are used to represent the same or similar parts.
[0102] The terms "first," "second," etc., used in this document are not intended to specifically refer to order or sequence, nor are they intended to limit the invention. They are merely used to distinguish elements or operations described using the same technical terms.
[0103] The directional terms used in this article, such as up, down, left, right, front, or back, are for reference only when referring to the accompanying drawings. Therefore, the use of directional terms is for illustrative purposes and not to limit this work.
[0104] The terms “include,” “including,” “have,” “contain,” etc., used in this article are all open-ended terms, meaning that they include but are not limited to.
[0105] The term "and / or" as used herein includes any or all of the things mentioned.
[0106] The term "multiple" in this article includes "two" and "more than two"; the term "multiple groups" in this article includes "two groups" and "more than two groups".
[0107] The terms "approximately," "about," etc., used herein are intended to modify any quantity or error that may vary slightly, but these slight variations or errors do not change the essence of the quantity or error. Generally, the range of slight variations or errors modified by such terms may be 20% in some embodiments, 10% in others, 5% in still others, or other values. Those skilled in the art should understand that the aforementioned values can be adjusted according to actual needs and are not limited thereto.
[0108] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0109] When expressions such as "at least one of A, B, and C" are used, they should generally be interpreted in accordance with the meaning commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, systems having A alone, having B alone, having C alone, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). When expressions such as "at least one of A, B, or C" are used, they should generally be interpreted in accordance with the meaning commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, or C" should include, but is not limited to, systems having A alone, having B alone, having C alone, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). A person skilled in the art should also understand that any conjunction and / or phrase that substantially arbitrarily indicates two or more optional items, whether in the specification, claims, or drawings, should be understood to indicate the possibility of including one of these items, either of these items, or both items. For example, the phrase “A or B” should be understood as including the possibility of “A” or “B”, or “A and B”.
[0110] Currently, there are two main core technology routes for industrial automatic locking systems. Both solutions can achieve precision locking operations, but they have significant shortcomings in terms of flexibility, cost control, and engineering implementation. They cannot adapt to the flexible production needs of multiple varieties and small batches, and they also greatly limit the large-scale application of dexterous hands in industrial precision locking assembly scenarios.
[0111] The first type is a dedicated tightening module solution. This solution uses a dedicated tightening device integrating a servo motor, torque sensor, and precision guide rail. Relying on a rigid connection structure, it ensures the coaxiality of the bit and screw, achieving a positioning accuracy of ±0.1mm, which meets the requirements of high-precision fastening operations. However, this solution has extremely poor adaptability. The device is designed only for specific screw specifications and fixed workpiece types, resulting in limited adaptability and very low flexibility. Furthermore, the high degree of customization of dedicated equipment leads to high overall hardware investment costs and extremely poor adaptability to production changes.
[0112] The second category is the collaborative robot + customized end effector solution. This solution uses a custom-designed screwdriver gripping mechanism mounted on the end effector of the collaborative robot, along with mechanical positioning pins or elastic clamping mechanisms to offset some of the tool gripping errors, thereby completing the locking and fastening operation. However, this solution also has its implementation challenges. When changing equipment models, the entire customized end effector needs to be replaced, resulting in a cumbersome changeover process, long debugging cycles, low production adaptation efficiency, and difficulty in quickly responding to production needs during product category switching.
[0113] Overall, the two existing mainstream automatic screw fastening solutions share common core defects and several unresolved technical pain points. First, dedicated tightening modules and customized end effectors are specifically designed for specific applications and are incompatible with general-purpose electric screwdrivers, requiring continuous investment in customized hardware, resulting in high overall equipment investment costs and low reusability. Second, existing technologies fail to address the random posture error problem when dexterous hands grasp general-purpose tools, severely limiting the full utilization of the dexterous hand's high flexibility. In actual operation, when a dexterous hand grasps a general-purpose electric screwdriver, the randomness of the finger envelope posture causes the screwdriver bit to exhibit a random deviation of ±1~3mm in the XY plane. This deviation is directly transmitted to the fastening point, frequently causing operational failures such as screw alignment failures and thread damage. Furthermore, the insufficient repeatability of the dexterous hand has become a core bottleneck for its application in precision assembly scenarios.
[0114] To address tool alignment deviation issues, existing technologies have developed visual servo optimization solutions. These solutions involve mounting a camera on the hand or tool end to track the target position in real time and provide feedback to control the robotic arm's movement, thereby improving alignment accuracy. However, this optimization solution has significant technical limitations and engineering implementation challenges: First, the solution requires extremely high control frequency and must be coupled in real time with the multi-degree-of-freedom motion of the dexterous hand, resulting in complex control logic. Second, the camera's field of view changes in real time with the hand's movements, causing dynamic fluctuations in the equipment calibration relationship, which significantly increases the overall system complexity. Third, this solution is difficult to adapt to precision alignment scenarios with close-range, narrow-field-of-view operations, and its accuracy cannot meet the requirements for precision locking.
[0115] In addition, existing visual servoing solutions generally adopt a real-time servo control strategy of "grasping and adjusting at the same time". The control link needs to process two core algorithms simultaneously: the closed-loop control of the dexterous hand grasping force and the visual servoing motion planning. The high coupling of multiple algorithm modules greatly increases the difficulty of algorithm development and engineering implementation, and also leads to a decrease in system stability.
[0116] In summary, the current field of industrial automatic locking has not yet formed a mature technical paradigm of "dexterous hand + general tool + precision compensation". Existing technologies cannot meet the three core requirements of low cost, rapid changeover and high precision operation of end effectors in flexible automated production lines, and are difficult to adapt to the modern flexible production mode of multi-variety and small batch. This is also the core improvement direction of this invention against existing technologies.
[0117] The following technical problems exist in the existing technology:
[0118] 1) The dexterous hand's gripping error has not been effectively eliminated. Existing solutions either ignore this error, resulting in insufficient locking accuracy; or they employ complex vision servo real-time compensation, leading to poor system stability.
[0119] 2) Dedicated hardware is expensive. Dedicated tightening modules or customized end effectors need to be specially designed for specific tools and cannot reuse general-purpose electric screwdrivers, resulting in high hardware investment and maintenance costs;
[0120] 3) Insufficient flexibility in model changeover. Existing solutions often require replacement of physical hardware or recalibration when changing models, resulting in long debugging cycles and difficulty in adapting to multi-variety, small-batch production modes;
[0121] 4) High complexity of control algorithms. Visual servoing solutions need to handle dynamic calibration and multi-degree-of-freedom coordinated control in real time, making them difficult to implement in engineering.
[0122] This invention provides an online visual correction method for the spatial pose of a dexterous hand end effector used in automatic locking, such as... Figure 1 As shown, the method includes the following steps:
[0123] Step S101: The robotic arm drives the dexterous hand to place the screwdriver at a preset calibration station. This preset calibration station is located at a preset position relative to the camera device. The robotic arm is a multi-degree-of-freedom robotic arm, and the dexterous hand is a multi-finger dexterous hand. In an optional embodiment, the multi-degree-of-freedom robotic arm drives the multi-finger dexterous hand to move to the screwdriver storage position. The dexterous hand performs a grasping action to fully grip the electric screwdriver. Subsequently, the robotic arm raises the screwdriver to a fixed height directly above the camera, which is the preset calibration station. This step completes tool acquisition and enters the calibration preparation state.
[0124] Step S102: Acquire multiple frames of screwdriver bit images using a camera device.
[0125] Step S103: Input multiple frames of head images into a preset model to extract the pixel-level center coordinates of the head images; in an optional embodiment, the preset model may be a YOLOv11 object detection network.
[0126] Step S104: Obtain the pixel deviation to be compensated based on the pixel-level center coordinates of the multi-frame batch head and the preset reference center.
[0127] Step S105: Based on preset rules, the pixel deviation to be compensated is mapped to a physical displacement in the robot arm's base coordinate system. Regarding steps S104-S105 above, in an optional embodiment, multiple screwdriver bit images are acquired as training samples; these training samples are input into the YOLOv11 object detection network for training, enabling it to learn the visual features of the screwdriver bit in a real-world scene; intrinsic parameter calibration and pixel-to-physical scale calibration are performed on the camera used to acquire the images; the screwdriver bit images are input into the visual calibration and algorithm preprocessing module to obtain the preset model and calibration parameters. Specifically, hundreds of image data of screwdriver bits positioned directly above the camera are acquired, and the YOLOv11 object detection network is trained to learn the visual features of the screwdriver bit in a real-world scene; camera intrinsic parameter calibration and pixel-to-physical scale calibration are completed. The main execution body for this step is the visual calibration and algorithm preprocessing module, with the screwdriver bit image dataset as input and the trained YOLOv11 detection model and calibration parameters as output. In the dataset construction phase, the dataset can be expanded to include multiple categories (e.g., the Benchmark for 6D Object Pose Estimation, or BOP) datasets for different industrial part types, or it can incorporate perceptual inputs such as red-green-blue color images + depth images (RGB-D) and point cloud fusion to adapt to more industrial operation scenarios. "Upward-facing camera" refers to a camera arrangement where the optical axis is perpendicular to the horizontal work surface and shooting upwards, which is fundamentally different from the hand-eye camera at the end effector of the robotic arm.
[0128] The YOLOv11 object detection network described above is only an example. In the target region extraction stage, YOLO can be replaced with other trainable high-precision instance segmentation networks (such as Mask R-CNN, RT-DETR-Seg, etc.). As long as its output mask can be used as a priori for the target region of SAM-6D, it still belongs to the alternative implementation of this invention.
[0129] Step S106: Control the robotic arm to perform a small translational motion in the XY plane according to the physical displacement, so as to drive the dexterous hand and the screwdriver it grasps to move synchronously until the center of the screwdriver bit coincides with the theoretical working point, thus completing the posture correction.
[0130] Step S107: Maintain the corrected pose and perform the screw tightening operation. Specifically, the robotic arm drives the corrected screwdriver to move to the target screw position and performs the tightening operation. Since the corrected bit positioning accuracy can reach ±0.3mm, it meets the precision tightening requirements for M2.5 and larger screws. "±0.3mm accuracy" refers to the positioning accuracy of the bit center in the XY plane, which can be adjusted in the algorithm through the pixel error tolerance parameter.
[0131] The overall process of this optional embodiment is as follows: a dexterous hand picks up the electric screwdriver from the material rack and lifts it to the calibration station → the camera captures the screwdriver bit image → the YOLOv11 model detects the center of the screwdriver bit → pixel deviation is calculated and mapped to physical displacement → the robotic arm performs a small-amplitude translation in the XY plane to complete the calibration → it moves to the work point to perform fastening. "YOLOv11" refers to the YOLO series target detection model released by Ultralytics in 2024, and its segmentation version (YOLOv11-seg) can be used to achieve more accurate screwdriver bit contour extraction.
[0132] The working mechanism of this invention is as follows: the random pose error introduced by the dexterous hand grasping is transformed into a statically observable visual feedback quantity through a specific time window of "after grasping and before operation"; pixel-level center positioning is achieved using deep learning target detection, and then mapped to controllable physical motion quantities of the robotic arm through a calibration scale, forming a closed-loop correction link of "detection-solution-compensation". Compared with real-time visual servoing, the "post-correction" strategy of this optional embodiment decouples visual detection from mechanical motion, avoiding coupling with the complex control of the dexterous hand, and the system architecture is simple and reliable; compared with dedicated hardware solutions, this optional embodiment improves the operation accuracy of general-purpose tools with software algorithms, significantly reducing hardware costs. This solution addresses the problems of existing industrial automated clamping systems, which cannot balance compatibility with general-purpose tools, gripping posture accuracy, and changeover efficiency, resulting in poor flexibility, high cost, or complex control. By employing an "online correction after gripping" mechanism and a YOLOv11 vision inspection closed loop, the random posture error of the dexterous hand gripping a general-purpose tool is transformed into a compensable physical displacement, overcoming the limitations of repeatability accuracy of the dexterous hand. At the same time, software-defined accuracy replaces hardware stacking, achieving low-cost, easy-to-change, flexible clamping, and supporting flexible configuration of tolerance according to the workpiece accuracy level, balancing efficiency and quality.
[0133] Step S102 above involves acquiring multiple frames of screwdriver bit images using the camera device. In an optional embodiment, the dexterous hand causes the screwdriver to rotate at a uniform angular velocity of... To rotate, among which, The unit is rad / s, determined by the constant rotational speed of the screwdriver controlled by the dexter's buttons at low speed. Multiple frames of screwdriver bit images are captured by a camera located above the screwdriver. These multiple frames are then input into the aforementioned preset model to extract the pixel-level center coordinates of the screwdriver bit. The implementation method is as follows: the camera frame rate of the camera device is... ,in, That is, the number of frames collected per second, used to determine the sampling time interval. and total number of sampled frames ,in The rotation period is _____.
[0134] Sampling strategy alternatives: Fixed frame averaging can be replaced by adaptive sampling (such as dynamically adjusting the number of sampling points according to the rotation period), or median filtering / weighted averaging can be used instead of arithmetic averaging. As long as multi-frame acquisition and fusion are completed during the screwdriver rotation process to eliminate eccentricity error, they all fall within the protection scope of this invention.
[0135] No. Frame acquisition time (Unit: s) , The corresponding rotation angle is ;in, The initial rotation phase angle, in rad, is determined by the random initial angle of the bit at the first sampling moment. After Gaussian filtering for noise reduction and Contrast Limited Adaptive Histogram Equalization (CLAHE) for contrast enhancement, the multi-frame bit images are input into the preset model to obtain the detection box. Then, the pixel-level bit center coordinates of the multi-frame bit images are calculated using the gray-scale centroid method. The unit is pixels; the detection box is used to define the position of the batch head in the image. In the horizontal direction, It is in the vertical direction.
[0136] In an optional embodiment, the aforementioned preset reference center includes the reference center of the Region of Interest (ROI) bounding box. Step 104 involves obtaining the pixel deviation to be compensated based on the multi-frame screwdriver pixel-level center coordinates and the preset reference center. In an optional embodiment, the observed pixel deviation between the multi-frame screwdriver pixel-level center coordinates and the ROI bounding box reference center is calculated respectively. The observed pixel deviation is estimated by least squares based on the eccentricity error model to eliminate the periodic error introduced by the screwdriver manufacturing eccentricity, thus obtaining the true system deviation. This true system deviation is used as the pixel deviation to be compensated. Specifically, the dexterous hand holds down the screwdriver button to make it rotate at a low speed, during which an upward-facing industrial camera repeatedly captures screwdriver images. Each frame of image is input into the trained YOLOv11 model to extract the screwdriver pixel-level center coordinates, and the pixel deviation between each frame and the ROI bounding box reference center is calculated. The deviation of the multi-frames is estimated by least squares based on the eccentricity error model to eliminate the periodic error introduced by the screwdriver manufacturing eccentricity, thus obtaining the true system deviation. The pixel deviation is mapped to the physical displacement in the robot arm's base coordinate system according to the calibration scale. This rotational sampling mechanism can eliminate the eccentricity error of screwdriver bits introduced by manufacturing or assembly. Here, "pixel deviation" refers to the difference in pixel coordinates between the center of the screwdriver bit and the reference center of the ROI detected in a single frame image, which includes true system deviation, eccentricity error, and random noise. "True system deviation" refers to the core pose deviation caused only by the randomness of the dexterous hand's grip after eliminating eccentricity error through multi-frame estimation, i.e., the amount that needs to be compensated. "Eccentricity error" refers to the periodic positional deviation of the screwdriver bit caused by the misalignment of the rotation center and geometric center due to manufacturing tolerances or assembly clearances. This error can be effectively suppressed by the rotational sampling and averaging mechanism.
[0137] In a specific optional embodiment, the observed pixel deviation is estimated using least squares based on the eccentricity error model to eliminate the periodic error introduced by the eccentricity of the bit manufacturing, and the true system deviation is obtained, including: let the eccentricity vector between the geometric center and the rotation center of the bit be... (Unit: pixel, a two-dimensional vector in pixel coordinates, representing the direction and magnitude of the deviation of the bit's geometric center from the rotation center); Module length is... Unit: pixel (eccentricity, i.e., the length of the eccentric vector); direction angle: Unit: rad, angle between the eccentric vector and the horizontal axis of the image coordinate system), the... Frame observation pixel deviation Includes true system bias Eccentricity error and random noise :
[0138]
[0139]
[0140] in, The first Frame in Pixel deviation in direction; unit: pixel; : The true system deviation of the orientation, i.e. the core pose deviation introduced by the grasping to be estimated, in pixels; The first Frame in Random measurement noise in direction, unit: pixel;
[0141] make , ,in, For the eccentric vector in Projected components of the axis, unit: pixel. For the eccentric vector in Projected components of the axis, unit: pixel; construct a linear least squares model. ,
[0142] in:
[0143]
[0144] Represents the observed pixel vector, unit: pixel, by Frame pixel deviations are formed by stacking them in frame order. 3D column vector;
[0145]
[0146] in, Design matrix, dimensions Due to the rotation angle Construct a known coefficient matrix;
[0147] : The vector of states to be estimated, dimension It includes true system bias and eccentricity error;
[0148]
[0149] Noise vector, dimension Assuming each component is independent and identically distributed, with variance , ;
[0150] The optimal estimate is:
[0151]
[0152] : The least squares estimate, dimension ;
[0153] Extracting the true system bias estimate , and eccentricity parameters ,in , For estimating vectors The first and second components.
[0154] The above Eccentricity estimate, unit: pixel, derived from the estimated projection components. The calculated eccentric vector magnitude.
[0155] In the pose refinement stage, the standard Iterative Closest Point (ICP) algorithm can be replaced by point-to-surface ICP, weighted ICP, ICP with normal constraints, or other iterative optimization methods based on rigid body registration. As long as the object of the algorithm is the initial pose result of SAM-6D and it is used to further optimize the final 6D pose, it can be regarded as an equivalent alternative of the present invention.
[0156] In an optional embodiment, the data is collected during the rotation process. Frame image ( ), No. The deviation of the observed pixels obtained from frame detection is By using least squares estimation, the estimated covariance matrix is output while eliminating eccentricity errors. , used to assess the reliability of the estimate; where, Random measurement noise variance, unit: pixel²; when the standard deviation and One of them exceeds the threshold When this happens, the number of sampling frames is automatically increased and the estimation is recalculated; among them, Estimate the covariance matrix, dimension This reflects the uncertainty of each parameter to be estimated; : The estimated standard deviation of the true systematic bias in direction, in pixels, reflects... Reliability of the estimates; : The estimated standard deviation of the true systematic bias in direction, in pixels, reflects... Reliability of the estimates; Convergence criterion threshold (convergence criterion threshold, unit: pixel, determined by back-calculation based on the desired physical accuracy, such as...) (pixel corresponds to a physical precision of approximately ±0.15mm).
[0157] In an optional embodiment, regarding the convergence criterion, the estimation is deemed reliable and compensation can be performed when the following conditions are met simultaneously:
[0158]
[0159] If the requirements are not met, increase the number of sampling frames. ;in: This refers to the number of sampling frames added each time, typically 5 to 10 frames. The larger the value, the smaller the covariance of the least squares estimate, and the more reliable the estimate.
[0160] Step S105 above involves mapping the pixel deviation to be compensated into a physical displacement in the robot arm's base coordinate system based on a preset rule. In an optional embodiment, the camera intrinsic parameter matrix of the imaging device is set as follows:
[0161]
[0162] in, The camera is Focal length in a direction, unit: pixel; : Camera principal point coordinates, unit: pixel, i.e., the coordinates of the intersection of the camera optical axis and the image plane;
[0163] The reference center point of the ROI in the pixel coordinate system is (Unit: pixel, pre-calibrated reference center of the region of interest, corresponding to the ideal position of the bit), the estimated true systematic bias is: (Unit: pixel);
[0164] Considering lens radial distortion, define the normalized plane radius:
[0165]
[0166] in Average focal length, unit: pixel; Normalized plane radius, dimensionless, represents the normalized distance of a pixel relative to the principal point, used for distortion compensation calculation;
[0167] Pixel deviation after distortion compensation:
[0168]
[0169]
[0170] The basic scale is obtained through calibration. ,in Unit: mm / pixel , Indicates the radial distortion coefficient of the lens. The first-order radial distortion coefficient dominates barrel / pincushion distortion, and is related to... Proportional The radial distortion coefficient of a second-order lens, and Proportional; taking into account the actual distance between the camera and the bit. (Unit: mm, actual vertical distance from the camera's optical center to the bit's end face at the current acquisition moment) and calibration distance (Unit: mm, standard vertical distance from the optical center of the camera to the end face of the bit during scale calibration) Differences in scale, corrected scale:
[0171]
[0172] ( Basic scale, unit: mm / pixel, at calibration distance The following was obtained through precision displacement experiments; : Distance-corrected scale, unit: mm / pixel, used to convert pixel deviation into physical displacement.
[0173] The compensated displacement in the base coordinate system of the robotic arm is:
[0174] . ( : Robotic arm base coordinate system The directional compensation displacement, in mm, is directly used to drive the robotic arm to perform micro-translation.
[0175] In one alternative embodiment:
[0176] If the alignment is correct, it is considered "aligned"; otherwise, a compensation movement is performed.
[0177] in, Pixel error tolerance, unit: pixel, is set according to the workpiece accuracy requirements. The pixel error tolerance is a software-configurable parameter that can be adjusted in the algorithm according to different workpiece accuracy requirements, taking into account both detection speed and positioning accuracy.
[0178] High-precision mode: pixel pixel corresponds to a physical precision of approximately ±0.1~0.2mm;
[0179] Standard mode: pixel pixel corresponds to a physical precision of approximately ±0.3~0.5mm;
[0180] Quick Mode: pixel Pixel priority is given to ensuring the cycle time of the operation.
[0181] In one alternative embodiment:
[0182] The robotic arm receives the displacement ( , The robot performs a small-amplitude translational motion in the XY plane to align the center of the bit with the theoretical working point, thus completing the pose correction. This corrected pose serves as the reference for subsequent fastening operations. The repeatability accuracy of the robotic arm is considered. (Unit: mm, repeatability error of the robotic arm body in the XY plane, given in the robotic arm specification sheet), total compensation error budget:
[0183]
[0184] Total compensation error, in mm, is a composite of visual estimation error, scale mapping error, and robotic arm execution error. They are respectively Physical error of orientation visual estimation error mapping, unit: mm.
[0185] Step S106 above involves controlling the robotic arm to perform a small translational motion in the XY plane according to the physical displacement. In an optional embodiment, the correction mechanism only acts on the XY plane, and the rotation angle... Keep the preset value unchanged:
[0186]
[0187] Compensation transformation matrix, dimension , used to describe pure translation compensation transformation in the base coordinate system of the robotic arm; : Identity matrix.
[0188] The key timing feature of this optional embodiment is "post-grip correction":
[0189] This timing sequence decouples the complex control of the dexterous hand (T1) from visual inspection (T2-T3) and mechanical compensation (T4) in time, with each stage controlling a single object; at the same time, through eccentricity error modeling and distortion correction, the system error sources are eliminated one by one, improving the final compensation accuracy.
[0190] This optional embodiment proposes an "online visual correction after gripping" method and a flexible automatic locking system. Its core highlight is: the first "online correction after gripping" mechanism, which transforms the ±1~3mm repeatability error introduced by the randomness of the finger envelope posture when the dexterous hand grips a general tool into an observable and quantifiable visual feedback quantity. Error compensation is achieved by the slight translation of the robotic arm's base coordinate system, thereby breaking through the limitation of the dexterous hand's repeatability accuracy on the system performance. In practical implementation, after the dexterous hand grasps the screwdriver and raises it to a fixed height, an upward-facing industrial camera deployed on the worktable repeatedly captures images of the screwdriver bit (sampling is also performed during the screwdriver's low-speed rotation to eliminate eccentricity errors introduced by the bit's manufacturing or assembly). Each frame is input into a trained YOLOv11 deep learning model to achieve pixel-level center detection of the screwdriver bit. The pixel deviation between each frame and the reference center of the ROI (Region of Interest) is calculated. Based on the eccentricity error model, the deviations of multiple frames are estimated using least squares to eliminate periodic errors and obtain the true system deviation. Then, the pixel deviation is mapped to the physical displacement in the XY direction under the robot arm's base coordinate system according to a calibrated scale, driving the robot arm to perform micro-translation, forming a closed-loop correction link of "visual detection - deviation calculation - mechanical compensation". This solution transforms the dexterous hand's grasping error from a system bottleneck into an eliminable quantity, enabling general-purpose electric screwdrivers to achieve specialized-level precision operation capabilities, achieving ±0.3mm sub-millimeter precision fastening, meeting the assembly requirements of M2.5 and larger screws.
[0191] This optional embodiment proposes a low-cost combination solution of a dexterous hand and a universal electric screwdriver. It replaces hardware-based precision with software-defined precision, significantly reducing the cost and changeover threshold of end effectors in flexible automated production lines. This solves the problems of existing dedicated tightening modules and customized end effectors requiring specific design for particular tools, incompatibility with universal tools, and high hardware investment costs. Simultaneously, through a software-configurable design with pixel error tolerance, the same hardware platform can flexibly adapt to the fastening requirements of workpieces with different precision levels, balancing efficiency and quality. This optional embodiment aims to address the shortcomings of existing technologies, such as the difficulty in eliminating random pose errors when a dexterous hand grasps a universal tool, the complexity of the control chain in vision servoing solutions requiring real-time coupling with the dexterous hand's multiple degrees of freedom, the high coupling degree of the "grasp and adjust simultaneously" strategy algorithm leading to poor system stability, and the lack of a mature "dexterous hand + universal tool + precision compensation" technology paradigm. The optional embodiments of this invention innovatively propose: (1) a specific operation sequence of "online correction after gripping" and its application in pose compensation of the end effector of the dexterous hand; (2) a method for center detection of the end effector and micro-translation compensation of the robot arm base coordinate system based on a fixed-table vision unit; (3) a coupling mechanism between a deep learning target detection model and robot arm motion control, including a pixel-physical space mapping scale calibration method; (4) a software-configurable design of pixel error tolerance and a parameter adaptive method under different precision requirement scenarios; (5) a flexible automatic fastening system architecture of dexterous hand + general electric screwdriver + online vision correction. The expected effects of this invention also include providing a reusable technical paradigm for the application of dexterous hands in industrial precision operation scenarios, effectively balancing the cost, changeover efficiency and fastening accuracy of flexible automated production lines.
[0192] This embodiment also provides an online visual correction device for the spatial pose of a dexterous hand end effector for automatic locking. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0193] This embodiment provides an online visual correction device for the spatial pose of a dexterous hand end effector used in automatic locking, such as... Figure 2 As shown, the device includes:
[0194] The placement module 21 is used to place a screwdriver at a preset calibration station by driving a dexterous hand through a robotic arm; wherein the preset calibration station is located at a preset position relative to the camera device, the robotic arm is a multi-degree-of-freedom robotic arm, and the dexterous hand is a multi-finger dexterous hand;
[0195] Acquisition module 22 is used to acquire multiple frames of screwdriver bit images through the camera device;
[0196] Extraction module 23 is used to input the multi-frame head images into a preset model to extract the pixel-level center coordinates of the multi-frame head;
[0197] The first processing module 24 is used to obtain the pixel deviation to be compensated based on the pixel-level center coordinates of the multi-frame header and the preset reference center.
[0198] The second processing module 25 is used to map the pixel deviation to be compensated into a physical displacement in the base coordinate system of the robotic arm based on a preset rule.
[0199] The third processing module 26 is used to control the robotic arm to perform a micro-translational motion in the XY plane according to the physical displacement, so as to drive the dexterous hand and the screwdriver it grasps to move synchronously until the center of the screwdriver bit coincides with the theoretical working point, thus completing the posture correction.
[0200] The screw fastening module 27 is used to maintain the corrected position and perform screw fastening operations.
[0201] Further functional descriptions of the above modules are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0202] An exemplary embodiment of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a method according to an embodiment of the present invention.
[0203] An exemplary embodiment of the present invention also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of the present invention.
[0204] An exemplary embodiment of the present invention also provides a computer program product, including a computer program, wherein, when executed by a computer's processor, the computer program is used to cause the computer to perform a method according to an embodiment of the present invention.
[0205] refer to Figure 3The present invention will now describe a structural block diagram of an electronic device that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0206] like Figure 3 As shown, the electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. The RAM 303 may also store various programs and data required for the operation of the device 300. The computing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0207] Multiple components in electronic device 300 are connected to I / O interface 305, including: input unit 306, output unit 307, storage unit 308, and communication unit 309. Input unit 306 can be any type of device capable of inputting information to electronic device 300. Input unit 306 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 307 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 308 may include, but is not limited to, disk and optical disk. Communication unit 309 allows electronic device 300 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0208] The computing unit 301 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above. For example, the methods of the above embodiments can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 300 via ROM 302 and / or communication unit 309. In some embodiments, the computing unit 301 can be configured to perform the methods according to embodiments of the present invention by any other suitable means (e.g., by means of firmware).
[0209] The program code used to implement the methods of the present invention can be written in any combination of one or more programming languages. This program code can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0210] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0211] As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0212] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0213] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0214] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0215] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A method for online visual correction of the spatial pose of a dexterous hand end effector for automatic locking, characterized in that, The method includes: A screwdriver is placed at a preset calibration station by a robotic arm that drives a dexterous hand; wherein the preset calibration station is located at a preset position relative to the camera device, the robotic arm is a multi-degree-of-freedom robotic arm, and the dexterous hand is a multi-finger dexterous hand; The camera device captures multiple frames of screwdriver bit images; The multi-frame bit images are input into a preset model to extract the pixel-level center coordinates of the multi-frame bit images; The pixel deviation to be compensated is obtained based on the pixel-level center coordinates of the multi-frame header and the preset reference center; Based on preset rules, the pixel deviation that needs to be compensated is mapped to the physical displacement in the robot arm's base coordinate system; The robotic arm is controlled to perform a small translational motion in the XY plane according to the physical displacement, so as to drive the dexterous hand and the screwdriver it grasps to move synchronously until the center of the screwdriver bit coincides with the theoretical working point, thus completing the posture correction; Maintain the corrected position and perform the screw fastening operation.
2. The method for online visual correction of the spatial pose of a dexterous hand end effector for automatic locking as described in claim 1, characterized in that, Acquiring multiple frames of screwdriver bit images via the camera device includes: using the dexterous hand to make the screwdriver rotate at a uniform angular velocity of... The screwdriver is rotated, and multiple frames of bit images of the screwdriver are captured by the camera device located above the screwdriver. The unit is rad / s; Inputting the multi-frame head images into a preset model to extract the pixel-level center coordinates of the multi-frame head includes: the camera frame rate of the camera device is... ,in, That is, the number of frames collected per second, used to determine the sampling time interval. and total number of sampled frames ,in The rotation period; No. Frame acquisition time The corresponding rotation angle is ;in, The initial rotation phase angle, in rad, is determined by the random initial angle of the bit at the first sampling moment. After Gaussian filtering and contrast enhancement through contrast-limited adaptive histogram equalization, the multi-frame bit images are input into the preset model to obtain the detection box. Then, the pixel-level bit center coordinates of the multi-frame bit images are calculated using the gray-scale centroid method. Unit: pixel; wherein, the detection box is used to define the position of the batch head in the image; In the horizontal direction, It is in the vertical direction.
3. The method for online visual correction of the spatial pose of a dexterous hand end effector for automatic locking as described in claim 2, characterized in that, The preset reference center includes the ROI box reference center. The pixel deviation to be compensated, based on the multi-frame header pixel-level center coordinates and the preset reference center, includes: Calculate the observed pixel deviation between the pixel-level center coordinates of the multi-frame batch header and the reference center of the ROI box; The observed pixel deviation is estimated by least squares based on the eccentricity error model to eliminate the periodic error introduced by the eccentricity of the bit manufacturing, and the true system deviation is obtained. The true system deviation is used as the pixel deviation that needs to be compensated.
4. The online visual correction method for the spatial pose of a dexterous hand end effector for automatic locking as described in claim 3, characterized in that, Based on the eccentricity error model, the observed pixel deviation is estimated using least squares to eliminate the periodic error introduced by the eccentricity in the bit manufacturing process, resulting in the true system deviation, which includes: Let the eccentricity vector between the geometric center and the rotation center of the bit be... Unit: pixel; Module length is Unit: pixel; Direction angle: Unit: rad, number Frame observation pixel deviation Includes true system bias Eccentricity error and random noise : in, The first Frame in Pixel deviation in direction, unit: pixel; : True systematic deviation in direction, unit: pixel; The first Frame in Random measurement noise in direction, unit: pixel; make , ,in, For the eccentric vector in Projected components of the axis, unit: pixel. For the eccentric vector in Projected components of the axis, unit: pixel; construct a linear least squares model. , in: Represents the observed pixel vector, unit: pixel, by Frame pixel deviations are formed by stacking them in frame order. 3D column vector; in, Design matrix, dimensions Due to the rotation angle Construct a known coefficient matrix; : The vector of states to be estimated, dimension It includes true system bias and eccentricity error; Noise vector, dimension Assuming each component is independent and identically distributed, with variance , ; The optimal estimate is: : The least squares estimate, dimension ; Extracting the true system bias estimate , and eccentricity parameters ,in , For estimating vectors The first and second components.
5. The method for online visual correction of the spatial pose of a dexterous hand end effector for automatic locking according to claim 4, characterized in that, The method further includes: Assume data is collected during the rotation process. Frame image ( ), No. The deviation of the observed pixels obtained from frame detection is By using least squares estimation, the estimated covariance matrix is output while eliminating eccentricity errors. , used to assess the reliability of the estimate; where, Randomly measured noise variance, unit: pixel²; When standard deviation and One of them exceeds the threshold When needed, automatically increase the number of sampling frames and re-estimate; in, Estimate the covariance matrix, dimension This reflects the uncertainty of each parameter to be estimated; : The estimated standard deviation of the true systematic bias in direction, in pixels, reflects... Reliability of the estimates; : The estimated standard deviation of the true systematic bias in direction, in pixels, reflects... Reliability of the estimates; Convergence threshold, unit: pixel, determined by working backward from the desired physical accuracy.
6. The method for online visual correction of the spatial pose of a dexterous hand end effector for automatic locking according to claim 5, characterized in that, The method further includes: Convergence Criterion: An estimate is considered reliable and compensation can be performed when the following conditions are met simultaneously: If the requirement is not met, increase the number of sampling frames. ;in: This represents the number of sampling frames added each time.
7. The method for online visual correction of the spatial pose of a dexterous hand end effector for automatic locking according to claim 6, characterized in that, Based on preset rules, the pixel deviations that need to be compensated are mapped to physical displacements in the robotic arm's base coordinate system, including: Let the camera intrinsic parameter matrix of the imaging device be: in, The camera is Focal length in a direction, unit: pixel; : Camera principal point coordinates, unit: pixel, i.e., the coordinates of the intersection of the camera optical axis and the image plane; The reference center point of the ROI in the pixel coordinate system is The estimated true system deviation is ; Considering lens radial distortion, define the normalized plane radius: in Average focal length, unit: pixel; Normalized plane radius, dimensionless, represents the normalized distance of a pixel relative to the principal point, used for distortion compensation calculation; Pixel deviation after distortion compensation: The basic scale is obtained through calibration. ,in Unit: mm / pixel , This represents the radial distortion coefficient of the lens; taking into account the actual distance between the camera and the bit. Distance from calibration Differences, correct scale: in, and The unit is mm; The compensated displacement in the base coordinate system of the robotic arm is: 。 8. The method for online visual correction of the spatial pose of a dexterous hand end effector for automatic locking according to claim 7, characterized in that, The method further includes: At that time, it is determined to be "aligned"; in, Pixel error tolerance, unit: pixel, set according to workpiece accuracy requirements; High-precision mode: pixel pixel corresponds to a physical precision of approximately ±0.1~0.2mm; Standard mode: pixel pixel corresponds to a physical precision of approximately ±0.3~0.5mm; Quick Mode: pixel Pixel priority is given to ensuring the cycle time of the operation.
9. The method for online visual correction of the spatial pose of a dexterous hand end effector for automatic locking according to claim 7, characterized in that, The method further includes: The displacement is received by the robotic arm. , ); Combining the repeatability of the robotic arm Total compensation error budget: Total compensation error, in mm, is a composite of visual estimation error, scale mapping error, and robotic arm execution error. They are respectively Physical error of orientation visual estimation error mapping, unit: mm.
10. The method for online visual correction of the spatial pose of a dexterous hand end effector for automatic locking according to claim 9, characterized in that, Controlling the robotic arm to perform micro-translational motion in the XY plane according to the physical displacement includes: This correction mechanism only applies to the XY plane, and the rotation angle... Keep the preset value unchanged: Compensation transformation matrix, dimension , used to describe pure translation compensation transformation in the base coordinate system of the robotic arm; : Identity matrix.
11. The method for online visual correction of the spatial pose of a dexterous hand end effector for automatic locking according to any one of claims 1 to 9, characterized in that, The preset model is trained through the following steps: Multiple screwdriver bit images were collected as training samples; The training samples are input into the YOLOv11 object detection network for training, so that it learns the visual features of the bit in the actual scene; Perform intrinsic parameter calibration and pixel and physical scale calibration on the camera used to acquire images; The screwdriver bit image is input into the visual calibration and algorithm preprocessing module to obtain the preset model and calibration parameters.
12. An online visual correction device for the spatial pose of a dexterous hand end effector used in automatic locking, characterized in that, The device includes: A placement module is used to place a screwdriver at a preset calibration station by driving a dexterous hand through a robotic arm; wherein the preset calibration station is located at a preset position relative to the camera device, the robotic arm is a multi-degree-of-freedom robotic arm, and the dexterous hand is a multi-finger dexterous hand; The acquisition module is used to acquire multiple frames of screwdriver bit images through the camera device; The extraction module is used to input the multi-frame head images into a preset model to extract the pixel-level center coordinates of the multi-frame head; The first processing module is used to obtain the pixel deviation to be compensated based on the pixel-level center coordinates of the multi-frame header and the preset reference center. The second processing module is used to map the pixel deviation to be compensated into a physical displacement in the base coordinate system of the robotic arm based on a preset rule. The third processing module is used to control the robotic arm to perform a micro-translational motion in the XY plane according to the physical displacement, so as to drive the dexterous hand and the screwdriver it grasps to move synchronously until the center of the screwdriver bit coincides with the theoretical working point, thus completing the posture correction. The screw fastening module is used to maintain the corrected position and perform screw fastening operations.
13. An electronic device, comprising: processor; as well as Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-11.
14. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.
15. A computer program product, characterized in that, The computer program product includes instructions that, when executed, cause a computer to perform the method of any one of claims 1-11.