An unknown hand internal posture anti-sticking shaft hole assembly method based on visual tactile perception

By combining a dual-finger heterogeneous visual-tactile sensor system with admittance control, the initial posture deviation and jamming problems of the inner shaft of the robot hand are solved, realizing efficient and robust shaft-hole assembly, which is suitable for robot assembly tasks in complex environments.

CN122353256APending Publication Date: 2026-07-10SOUTHEAST UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHEAST UNIV
Filing Date
2026-04-30
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively perceive and correct initial posture deviations of the robot hand's internal axis in unstructured environments, especially prone to jamming on complex surface textures. Furthermore, existing methods rely on cumbersome physical calibration or data-driven training, resulting in weak transferability.

Method used

A dual-finger heterogeneous visual-tactile sensor system is adopted, which combines anchorless and anchored sensors. Principal component analysis and dual masking algorithms are used for attitude pre-alignment and low-delay force feedback. Combined with admittance control and Archimedes spiral motion for hole finding, a normal escape mechanism is designed to remove the jamming.

Benefits of technology

It achieves highly robust attitude calculation and low-delay force feedback without wrist force sensors, improves the assembly success rate in complex environments, and avoids the phase lag and jamming problems of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122353256A_ABST
    Figure CN122353256A_ABST
Patent Text Reader

Abstract

This invention discloses a method for anti-jamming shaft hole assembly based on visual-tactile perception for unknown hand posture. One visual-tactile sensor with and without anchor points are installed on the robot gripper. The orientation of the main axis of the gripped shaft is calculated using the visual-tactile image without anchor points, and the gripper posture is adjusted to achieve pre-alignment before insertion. To correct insertion failures caused by shaft posture estimation errors, the visual-tactile image with anchor points is used to extract low-latency relative force feedback in real time through a dual-mask spatial averaging algorithm. This guides a helical search based on force-potential hybrid admittance control to establish a stable contact state, achieving tactile-guided hole finding. Finally, for shafts with complex surface textures that experience jamming due to abrupt friction changes, a bidirectional escape mechanism involving inward and outward contraction along the helical normal direction is employed to release contact constraints, achieving compliant insertion guided by tactile perception. This invention achieves visual-tactile insertion operation in situations where there are no wrist force sensors and the initial posture of the shaft within the gripper is unknown, significantly improving the assembly success rate and robustness in complex working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot compliant control and visual-tactile perception technology, specifically to a method for anti-jamming shaft hole assembly based on visual-tactile perception of unknown hand posture. Background Technology

[0002] Shaft and hole assembly plays a crucial role in modern automated manufacturing systems. With the increasing demand for higher assembly quality and efficiency, robot-assisted assembly has become widespread. However, in real-world unstructured environments, workpiece manufacturing tolerances, fixture positioning errors, and the kinematic uncertainties of robot systems make assembly strategies that rely solely on high-precision position control difficult to succeed.

[0003] To achieve robotic shaft and hole assembly, existing research has introduced multiple modalities of perception feedback. A patent search revealed that Guo Lei et al. applied for an invention patent, "A Vision-Guided Screw / Threaded Sleeve Installation Method" (application number CN202512028132.6). This method utilizes symmetrically set acquisition poses to obtain point clouds of the threaded holes, and then uses point cloud segmentation to obtain the point cloud of the inner wall of the thread, thereby accurately obtaining the axial direction and center position of the threaded hole in the current workpiece, guiding the robot to successfully complete automatic assembly. However, in precision shaft and hole assembly scenarios, vision systems often face severe occlusion problems, making it difficult to obtain high-quality point clouds of screws and threads. Leng Jun et al. applied for an invention patent entitled "Robot Control System, Method and Robot Based on Vision and Force Fusion" (application number: CN202512007000.5). This robot control system and method based on vision and force fusion uses an artificial intelligence processing unit to analyze visual features in real time and predict mechanical parameters. Combined with an adaptive impedance control unit, it dynamically fuses position commands and force feedback, achieving precise adaptation to the characteristics of the target object and compliant control of the interaction process. It possesses real-time perception and dynamic adjustment capabilities, effectively avoiding component damage caused by mechanical mismatch and improving the reliability and efficiency of precision assembly. However, this method cannot directly obtain local geometric information of the contact surface and is difficult to perceive the initial posture deviation of the shaft of screw-like objects within the gripper. Gao Xiao et al. applied for an invention patent entitled "A Robot Shaft Hole Assembly System and Method Based on Teaching Learning" (application number: CN201811275792.8). This method uses a six-dimensional force / torque sensor and a passive flexible RCC device, and relies on manual teaching to collect data. It learns the mapping relationship between force and velocity through a Gaussian mixture model (GMM) to complete the assembly. This method requires complex preliminary manual teaching and data collection, and is highly dependent on expensive external six-dimensional force sensors and external compliant mechanisms. Furthermore, it cannot detect and correct hand posture deviations. Huang Panfeng et al. applied for an invention patent, "A Fine Perception Method for Axis-Hole Operation Based on Visual-Tactile Perception" (application number: CN202411223338.3), which uses a two-finger visual-tactile sensor to acquire optical flow images. By calculating the optical flow resultant vector, it estimates the motion trend of the clamped object after contact, and then infers the relative position of the shaft-hole contact point to achieve alignment. This method relies on the dynamic calculation of optical flow after the object and hole physically contact and deform. It not only suffers from phase lag and computational delay problems caused by the optical flow algorithm, but also fails to address the initial pose deviation before contact and lacks a recovery mechanism for jamming phenomena under complex working conditions.Li Mingfu et al. applied for an invention patent "A Robot Shaft Hole Assembly Method Based on Deep Reinforcement Learning and Admittance Control" (application number: CN202211369853.3). This method explores and learns the hole-searching and insertion actions in shaft hole assembly through deep reinforcement learning (DQN / DDPG network) and traditional force sensors, and outputs admittance control parameters. This method not only requires a long time of simulation or actual training to converge the network, but also has weak migration ability for workpieces of different shapes and materials. Moreover, when facing shaft assembly with complex surface textures (such as threads and discontinuous protrusions), once it gets stuck at the edge of the hole, deep reinforcement learning often cannot quickly provide an effective physical escape strategy.

[0004] In contrast, visual-tactile sensors, with their high-resolution local contact perception capabilities, can continuously and simultaneously sense the posture of objects within the hand and the contact force between the hand and the object, something that single vision or force sensors cannot achieve. Visual-tactile sensors are widely used in robotic skill manipulation. Achu Wilson et al. proposed combining tactile-guided low-level motion control with high-level vision-based task parsing to achieve tasks such as plugging (see "Cable Routing and Assembly using Tactile-driven Motion Primitives", ICRA 2023), demonstrating the crucial role of visual-tactile sensors when visual signals are obstructed. Zicai Peng et al. proposed using visual and tactile perception based on particle filtering to estimate the object's posture (see "High-Precision Object Pose Estimation Using Visual-Tactile Information for Dynamic Interactions in Robotic Grasping", ICRA 2025), using visual information to track the posture of the touched object in real time and using displacement data obtained from tactile sensors to estimate the posture changes of the grasped object. These methods are insufficiently concerned with the hand posture of objects with uncertain initial poses, or require multiple contacts to determine the hand posture, making them unsuitable for industrial assembly scenarios. Furthermore, Michele Mirto et al. proposed a novel method combining robotic arms and electromechanical tools to automate connector assembly (see "Towards Automated Connector Assembly: Wire Insertion Combining Tactile and Vision Sensors", AIM 2025). This method uses tactile sensors to estimate the local shape and position of the wires, while a camera determines the pin axis for alignment. However, such force-sensing methods based on tactile vision rely on cumbersome physical calibration or data-driven training, making it difficult to adapt to different scenarios and sensors. Moreover, existing assembly strategies lack physical recovery mechanisms to address the issue of complex surface axes getting stuck near holes. Therefore, how to robustly assemble objects with large initial hand poses using tactile vision sensors is a significant challenge that needs to be addressed. Summary of the Invention

[0005] The purpose of this invention is to propose an anti-jamming shaft hole assembly method based on visual-tactile perception of unknown hand posture, thereby achieving highly robust posture pre-alignment, low-delay force feedback, and jamming self-recovery assembly.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] This invention provides a method for anti-jamming shaft hole assembly based on visual-tactile perception of unknown hand posture, comprising the following steps:

[0008] Step 1: Construct a two-finger heterogeneous sensing system. Use two parallel grippers to grasp the shaft to be assembled. The first finger of the gripper is equipped with a visual-tactile sensor without anchor points, and the second finger of the gripper is equipped with a visual-tactile sensor with anchor points.

[0009] Step 2: Use the first fingertip to obtain a tactile depth image without anchor interference, extract the maximum deformation area caused by the contact between the shaft and the first fingertip visual tactile sensor as the effective pixel set, calculate the covariance matrix of the point set by the principal component analysis algorithm and perform eigenvalue decomposition, take the eigenvector direction corresponding to the maximum eigenvalue as the projection principal axis direction of the shaft in the plane of the first fingertip visual tactile sensor, and then calculate the deflection angle of the shaft relative to the axis of the first fingertip visual tactile sensor and the translation correction vector in the image space. Combine the kinematic transformation matrix of the robotic arm to complete the pre-alignment of the shaft hole posture before assembly.

[0010] Step 3: Control the end of the robotic arm to move vertically downward along the aligned posture direction towards the worktable with the fixed hole. Use the second fingertip to obtain the anchor point displacement field. Use the noise mask and the core area mask to perform a dual mask spatial averaging calculation on the anchor point displacement field to obtain the resultant force estimate value parallel to the contact surface of the second fingertip visual-tactile sensor with low delay. When the resultant force estimate value is greater than the set contact threshold, it is determined that the shaft and the worktable surface have established contact.

[0011] Step 4: Perform Archimedes spiral motion to find the hole with the contact point as the center. During the spiral hole finding process, the estimated value of the local resultant force of the contact surface extracted in real time by the second fingertip visual tactile sensor is directly used as the external force feedback input of the admittance controller to drive the admittance controller to update the position of the end of the robot arm in real time, so that the shaft applies a constant contact force to the worktable. When the estimated value of the resultant force is detected to drop sharply and remain below the set hole entry threshold, it is determined that the shaft has aligned with the hole position and initially slid into the hole. At this time, the Archimedes spiral hole finding motion ends, and the end of the robot arm is controlled to move vertically downward along the axis of the hole to perform the shaft insertion operation.

[0012] Step 5: During the helical hole finding and insertion stage, if the resultant force estimate extracted by the anchor point displacement field of the second fingertip visual tactile sensor drops sharply and continues to exceed the given threshold, it is determined to be stuck. The robotic arm end is controlled to move inward or outward along the normal direction of the current helical position to try to release the constraint. If the movement distance reaches the limit and the constraint is not released, the hole finding state is re-entered until the extracted resultant force estimate meets the preset hole entry conditions and the shaft reaches the preset depth to complete the assembly.

[0013] Furthermore, in step 2, the deflection angle of the axis relative to the axis of the first fingertip visual-tactile sensor and the translation correction vector in the image space are calculated (i.e., the tactile depth image is processed), specifically as follows:

[0014] Set percentile parameters Calculate the corresponding percentiles of the entire map depth value set. As an adaptive threshold generation binary mask Extract the set of coordinates of all valid pixels within the mask area;

[0015] Calculate the mean values ​​of the centroid's horizontal and vertical coordinates of the contact area, construct the covariance matrix of this point set, perform eigenvalue decomposition on the covariance matrix to obtain the largest eigenvector, and use this eigenvector to calculate the deflection angle of the axis relative to the axis of the first fingertip visual-tactile sensor. ;

[0016] The principal axis of the shaft is modeled as a straight line passing through the centroid and with a known slope. The x-coordinate of the intersection point of this line and the left edge of the tactile image is calculated to obtain the translation correction vector in the image space. ;

[0017] Combined with the pixel-to-physical-size ratio known from the first fingertip visual-tactile sensor It is possible to construct a homogeneous transformation matrix from the visual-tactile sensor coordinate system of the first fingertip to the straight line modeled by the screw spindle. :

[0018] ;

[0019] in, Indicates the rotation angle is The rotation matrix, This refers to the translation correction vector in the image space calculated in step 2. The pixel-to-physical-size ratio of the first fingertip visual-tactile sensor. This represents the translation matrix.

[0020] Substituting the transformation matrix from the robot arm base to the flange, we can obtain the target pose of the robot arm end effector in the base coordinate system, which aligns the screw to the hole shaft. :

[0021] ;

[0022] in, The preset ideal calibration reference pose (in the base coordinate system) is used. This is the inverse matrix of the homogeneous transformation matrix from the coordinate system of the first fingertip visual-tactile sensor to the screw spindle, calculated in step two above. It is the inverse of the transformation matrix from the end flange to the gripper of the robotic arm.

[0023] Furthermore, in step 3, the displacement field of the anchor point is calculated using a dual-mask spatial averaging method, employing both a noise mask and a core region mask. Specifically:

[0024] The displacement field of the anchor point of the second fingertip visual tactile sensor is obtained, with the anchor point position without contact force after grasping as a reference;

[0025] A noise mask is constructed by setting a fixed threshold for the displacement field of the anchor point. :

[0026] ;

[0027] in, This represents the set of magnitudes of all anchor point displacement vectors in the current anchor point displacement field. This represents the maximum displacement modulus generated by all anchor points in the current displacement field. A fixed scaling factor is used to filter out deformation interference in non-contact areas outside the maximum deformation region. A preset fixed threshold, defaulting to 1.1 pixels, is used to filter out minute displacements caused by environmental disturbances; a fixed scaling factor is set to construct the core area mask. :

[0028] ;

[0029] in This is a fixed scaling factor, with a default value of 0.5, used to filter out deformation interference in non-contact areas outside the maximum deformation region;

[0030] The effective displacement set is obtained by multiplying the anchor point displacement field with the noise mask and the core region mask respectively and taking the intersection. ;

[0031] Summing the displacement vectors of all anchor points within the effective displacement set, dividing by the number of anchor point displacements in the effective displacement set, and then multiplying by the calibration coefficients for anchor point displacements and actual forces, yields the real-time estimate of the resultant force. :

[0032] ;

[0033] in, This represents the number of effective displacement sets after double masking.

[0034] Furthermore, in step 4, the constant contact force applied by the shaft to the worktable surface is achieved using an admittance control strategy, specifically as follows:

[0035] Build including virtual quality Virtual damping and virtual stiffness coefficient Admittance control model:

[0036] ;

[0037] in The preset desired contact force, and Right now .

[0038] The difference between the estimated target contact force and the estimated resultant force is used as the force input, and the target acceleration is obtained by solving the admittance control differential equation. The speed is obtained by one integration. In discrete control systems, the position correction is updated using the Euler integral. ;

[0039] In discrete control systems, position is updated using the Euler integral. :

[0040] ;

[0041] in Given the current height of the robotic arm, a compliant search with constant contact force is achieved. In the first The position correction amount calculated by the admittance controller in each discrete control cycle.

[0042] Furthermore, in step 5, the specific judgment and execution logic for normal escape processing is as follows:

[0043] During the spiral hole-finding process, the angle between the resultant force direction estimated by the second fingertip visual-tactile sensor and the direction of the main axis of the shaft is determined. If the included angle remains greater than the set threshold (Default is 0.174rad) If the axis is determined to encounter a discontinuous protrusion, the robotic arm is controlled to retract inward along the spiral normal to avoid the obstacle.

[0044] During the insertion process, if the estimated resultant force drops sharply due to the complex texture of the shaft end surface and fails to reach the target insertion depth, the machine enters a stuck state. At this time, admittance control is maintained to control the end of the robotic arm to move outward along the normal of the current position of the spiral until the resultant force drops sharply again and the insertion state is re-entered.

[0045] When attempting to release the stuck state, once the robotic arm moves to the threshold, it lifts the shaft along the hole axis (vertically upward) to leave the hole position and re-enters the hole-finding state.

[0046] Compared with existing technologies, this invention proposes an anti-jamming shaft hole assembly method based on visual-tactile perception of unknown hand posture. Compared with a large number of force-controlled insertion hole technologies based on wrist force sensors, this invention fundamentally changes the force control algorithm based on six-dimensional wrist force / torque feedback to a force control algorithm based on contact force within the gripper. Thus, robot insertion hole operation is achieved using only visual-tactile sensors within the gripper without wrist force sensors.

[0047] This invention achieves hole-axis alignment by estimating the axis posture based on visual-tactile perception and adjusting the robot posture. Therefore, it can be used when the initial posture of the axis inside the gripper is unknown, such as when the initial posture of the screw is not fixed by the gripper or when the screw is handed to the robot by a person.

[0048] To correct for insertion failures caused by shaft attitude estimation errors, this invention utilizes anchored visual-tactile images for helical search based on force-potential hybrid admittance control, achieving tactile-guided hole finding. Furthermore, to overcome jamming during shaft insertion, a bidirectional escape mechanism involving inward and outward contraction along the helical normal direction is proposed to release contact constraints, achieving compliant insertion guided by tactile feedback and ensuring a high success rate for the final insertion task.

[0049] Compared with the prior art, the beneficial effects of the present invention are:

[0050] A dual-finger heterogeneous sensing architecture is adopted to decouple geometric sensing (without anchor points) and mechanical sensing (with anchor points) at the physical level, achieving highly robust attitude calculation without the need for complex data-driven networks. A simplified force estimation algorithm based on dual masks is proposed. Compared with traditional methods based on complex physical models and sliding window filters, this algorithm avoids phase lag while filtering out high-frequency noise, achieving low-latency and high signal-to-noise ratio force tracking. The designed normal escape jamming handling mechanism breaks the limitation of traditional single spiral search that is prone to jamming at hole edges or discontinuous protrusions. It uses active bidirectional normal motion to physically restore contact constraints, greatly improving the assembly success rate of complex textured shafts in unstructured environments. Attached Figure Description

[0051] Figure 1 This is a flowchart illustrating the anti-jamming shaft hole assembly method for unknown hand posture based on visual-tactile perception according to the present invention.

[0052] Figure 2 This is a schematic diagram illustrating the calculation of the shaft's attitude and offset within the gripper in step two of this invention;

[0053] Figure 3 This is a schematic diagram of the jamming scenario in the unknown hand posture anti-jamming shaft hole assembly method based on visual-tactile perception in step five of the present invention.

[0054] Figure 4This is a schematic diagram illustrating the formation of the stuck state and the escape direction in step five of the present invention. Detailed Implementation

[0055] like Figure 1 As shown, a method for anti-jamming shaft hole assembly based on visual-tactile perception of unknown hand posture includes the following steps:

[0056] Step 1: Construct a two-finger heterogeneous sensing system. Use two parallel grippers to grasp the shaft to be assembled. The first finger of the gripper is equipped with a visual-tactile sensor without anchor points, and the second finger of the gripper is equipped with a visual-tactile sensor with anchor points.

[0057] Step 2, as follows Figure 2 As shown, a tactile depth image without anchor interference is obtained using the first fingertip. The maximum deformation area caused by the contact between the shaft and the first fingertip visual tactile sensor is extracted as the effective pixel set. The covariance matrix of the point set is calculated by the principal component analysis algorithm and eigenvalue decomposition is performed. The direction of the eigenvector corresponding to the maximum eigenvalue is taken as the projection principal axis direction of the shaft in the plane of the first fingertip visual tactile sensor. Then, the deflection angle of the shaft relative to the axis of the first fingertip visual tactile sensor and the translation correction vector in the image space are calculated. Combined with the kinematic transformation matrix of the robotic arm, the shaft hole attitude pre-alignment before assembly is completed.

[0058] Step 3: Control the end of the robotic arm to move vertically downward along the aligned posture direction towards the worktable with the fixed hole. Use the second fingertip to obtain the anchor point displacement field. Use the noise mask and the core area mask to perform a dual mask spatial averaging calculation on the anchor point displacement field to obtain the resultant force estimate value parallel to the contact surface of the second fingertip visual-tactile sensor with low delay. When the resultant force estimate value is greater than the set contact threshold, it is determined that the shaft and the worktable surface have established contact.

[0059] Step 4: Perform Archimedes spiral motion to find the hole with the contact point as the center. During the spiral hole finding process, the admittance control strategy is used to apply a constant contact force to the shaft on the worktable surface. The estimated resultant force is used as an external feedback input to drive the admittance controller to update the position of the end of the robot arm. When the estimated resultant force is detected to drop sharply and remain below the set hole entry threshold, it is determined that the shaft has aligned with the hole position and has initially slid into the hole. At this time, the Archimedes spiral hole finding motion ends, and the end of the robot arm is controlled to move vertically downward along the axis of the hole to perform the shaft insertion operation.

[0060] Step 5, as follows Figure 3 As shown, during the spiral hole finding and insertion stage, if the estimated resultant force extracted from the anchor point displacement field of the second fingertip visual-tactile sensor experiences a non-monotonic sudden drop and continuously exceeds a given threshold, it is determined to be stuck in a state of stagnation. Figure 4The control robot arm end moves inward or outward along the normal direction of the current spiral position to try to release the constraint. If the movement distance reaches the limit and the constraint is not released, it will re-enter the hole-finding state until the extracted resultant force estimate value meets the preset hole-entry conditions and the shaft reaches the preset depth to complete the assembly.

[0061] In step 2, when processing the tactile depth image, the percentile parameter is first set. And calculate the corresponding percentiles of the entire map depth value set. As an adaptive threshold generation binary mask Then, the coordinate set of all effective pixels within the mask area is extracted, the mean of the centroid's horizontal and vertical coordinates of the contact area is calculated, and the covariance matrix of this point set is constructed. The covariance matrix is ​​then decomposed into eigenvalues ​​to obtain the largest eigenvector. This eigenvector is used to calculate the deflection angle of the axis relative to the axis of the first fingertip visual-tactile sensor. Therefore, the principal axis of the shaft is modeled as a straight line passing through the centroid with a known slope. The x-coordinate of the intersection point of this line and the left edge of the tactile image is calculated to obtain the translation correction vector in the image space. .

[0062] After modeling the object's pose along the hand's main axis, the pixel-to-physical-size ratio coefficients known from the first fingertip visual-tactile sensor are combined. It is possible to construct a homogeneous transformation matrix from the coordinate system of the first fingertip visual-tactile sensor to the straight line modeled by the screw spindle. :

[0063] ;

[0064] Substituting the transformation matrix from the robot arm base to the flange, we can obtain the target pose of the robot arm end effector in the base coordinate system, which aligns the screw to the hole shaft. :

[0065] ;

[0066] in, The preset ideal calibration reference pose (in the base coordinate system).

[0067] In step 3, the displacement field of the anchor point at the second fingertip is obtained. The anchor point position without contact force after grasping is used as a reference, and a fixed threshold is set to construct a noise mask. :

[0068] ;

[0069] in A preset fixed threshold, defaulting to 1.1 pixels, is used to filter out minute displacements caused by environmental disturbances; a fixed scaling factor is set to construct the core area mask. :

[0070] ;

[0071] in A fixed scaling factor, defaulting to 0.5, is used to filter out deformation interference in non-contact areas outside the maximum deformation region. The anchor point displacement field is then multiplied by the noise mask and the core region mask, and the intersection is taken to obtain the effective displacement set. By summing the displacement vectors of all anchor points within the effective displacement set, dividing by the number of anchor point displacements in the effective displacement set, and then multiplying by the calibration coefficients of the anchor point displacements and actual forces, the real-time resultant force estimate is obtained. :

[0072] ;

[0073] in, This represents the number of effective displacement sets after double masking.

[0074] In step 4, construct a system containing virtual mass. Virtual damping Virtual stiffness coefficient Admittance control model:

[0075] ;

[0076] The difference between the estimated target contact force and the estimated resultant force is used as the force input, and the target acceleration is obtained by solving the admittance control differential equation. The speed is obtained by one integration. In discrete control systems, the position correction is updated using the Euler integral. Specifically:

[0077] ;

[0078] ;

[0079] ;

[0080] in, Indicates the first The target acceleration is calculated from the admittance model for each cycle. and They represent the first The speed of the current cycle and the previous cycle and They represent the first The position correction amount between the current cycle and the previous cycle The sampling period of the discrete control system;

[0081] In discrete control systems, position is updated using the Euler integral. :

[0082] ;

[0083] in This is the current height of the robotic arm, enabling a compliant search with constant contact force.

[0084] In step 5, the specific judgment and execution logic for the normal escape process is as follows: During the spiral hole-finding process, the angle between the resultant force direction estimated by the second fingertip visual-tactile sensor and the direction of the main axis of the shaft is determined. If the included angle remains greater than the set threshold (Default is 0.174rad) If the axis is determined to encounter a discontinuous protrusion, the robotic arm is controlled to retract inward along the spiral normal to avoid the obstacle.

[0085] During the insertion process, if the estimated resultant force drops sharply due to the complex texture of the shaft end surface and fails to reach the target insertion depth, the machine enters a stuck state. At this time, admittance control is maintained to control the end of the robotic arm to move outward along the normal of the current position of the spiral until the resultant force drops sharply again and the insertion state is re-entered.

[0086] When attempting to release the stuck state, once the robotic arm moves to the threshold, it lifts the shaft along the hole axis (vertically upward) to leave the hole position and re-enters the hole-finding state.

[0087] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any modifications or equivalent changes made based on the technical essence of the present invention shall still fall within the scope of protection claimed by the present invention.

Claims

1. A method for anti-jamming shaft hole assembly based on visual-tactile perception of unknown hand posture, characterized in that, Includes the following steps: Step 1: Construct a two-finger heterogeneous sensing system. Use two parallel grippers to grasp the shaft to be assembled. The first finger of the gripper is equipped with a visual-tactile sensor without anchor points, and the second finger of the gripper is equipped with a visual-tactile sensor with anchor points. Step 2: Use the first fingertip to obtain a tactile depth image without anchor interference, extract the maximum deformation area caused by the contact between the shaft and the first fingertip visual tactile sensor as the effective pixel set, calculate the covariance matrix of the point set by the principal component analysis algorithm and perform eigenvalue decomposition, take the eigenvector direction corresponding to the maximum eigenvalue as the projection principal axis direction of the shaft in the plane of the first fingertip visual tactile sensor, and then calculate the deflection angle of the shaft relative to the axis of the first fingertip visual tactile sensor and the translation correction vector in the image space. Combine the kinematic transformation matrix of the robotic arm to complete the pre-alignment of the shaft hole posture before assembly. Step 3: Control the end of the robotic arm to move vertically downward along the aligned posture direction towards the worktable with the fixed hole. Use the second fingertip to obtain the anchor point displacement field. Use the noise mask and the core area mask to perform a dual mask spatial averaging calculation on the anchor point displacement field to obtain the resultant force estimate value parallel to the contact surface of the second fingertip visual-tactile sensor with low delay. When the resultant force estimate value is greater than the set contact threshold, it is determined that the shaft and the worktable surface have established contact. Step 4: Perform Archimedes spiral motion to find the hole with the contact point as the center. During the spiral hole finding process, the estimated value of the local resultant force of the contact surface extracted in real time by the second fingertip visual tactile sensor is directly used as the external force feedback input of the admittance controller to drive the admittance controller to update the position of the end of the robot arm in real time, so that the shaft applies a constant contact force to the worktable. When the estimated value of the resultant force is detected to drop sharply and remain below the set hole entry threshold, it is determined that the shaft has aligned with the hole position and initially slid into the hole. At this time, the Archimedes spiral hole finding motion ends, and the end of the robot arm is controlled to move vertically downward along the axis of the hole to perform the shaft insertion operation. Step 5: During the helical hole finding and insertion stage, if the resultant force estimate extracted by the anchor point displacement field of the second fingertip visual tactile sensor drops sharply and continues to exceed the given threshold, it is determined to be stuck. The robotic arm end is controlled to move inward or outward along the normal direction of the current helical position to try to release the constraint. If the movement distance reaches the limit and the constraint is not released, the hole finding state is re-entered until the extracted resultant force estimate meets the preset hole entry conditions and the shaft reaches the preset depth to complete the assembly.

2. The method for anti-jamming shaft hole assembly based on visual-tactile perception of unknown hand posture according to claim 1, characterized in that, In step 2, the deflection angle of the axis relative to the axis of the first fingertip visual-tactile sensor and the translation correction vector in the image space are calculated, specifically including the following steps: Set percentile parameters, calculate the corresponding percentile points of the entire image depth value set as adaptive thresholds to generate a binary mask, and extract the coordinate set of all effective pixels within the mask area; Calculate the mean values ​​of the centroid's horizontal and vertical coordinates of the contact area, construct the covariance matrix of the point set, perform eigenvalue decomposition on the covariance matrix to obtain the largest eigenvector, and use this eigenvector to calculate the deflection angle of the shaft relative to the axis of the first fingertip visual-tactile sensor. Model the main axis of the shaft as a straight line passing through the centroid and with a known slope. Calculate the x-coordinate of the intersection point of this straight line and the left edge of the tactile image to obtain the translation correction vector in the image space. By combining the known pixel-to-physical-size scaling factor of the first fingertip visual-tactile sensor, a homogeneous transformation matrix can be constructed from the coordinate system of the first fingertip visual-tactile sensor to the straight line modeled by the screw spindle. Substituting this into the transformation matrix from the robot arm base to the flange, the target pose of the robot arm end in the base coordinate system can be obtained, which aligns the screw to the hole shaft.

3. The method for anti-jamming shaft hole assembly based on visual-tactile perception of unknown hand posture according to claim 2, characterized in that, In step 3, the displacement field of the anchor point is calculated using a dual-mask spatial averaging method with both a noise mask and a core region mask. This specifically includes the following steps: The displacement field of the anchor point of the second fingertip visual tactile sensor is obtained, and the anchor point position without contact force after grasping is used as a reference. A noise mask is constructed by setting a fixed threshold for the displacement field of the anchor point. : ; in, Represents the first in the displacement field of the anchor point Displacement vector of each anchor point; Represents the first in the displacement field of the anchor point The length of the anchor point; A preset fixed threshold is used to filter out minute displacements caused by environmental disturbances; Indicates dissatisfaction Other situations; A core area mask is constructed by setting a fixed scaling factor for the noise mask. : ; in, This represents the set of magnitudes of all anchor point displacement vectors in the current anchor point displacement field. This represents the maximum displacement modulus generated by all anchor points in the current displacement field. A fixed scaling factor is used to filter out deformation interference in non-contact areas outside the maximum deformation region; The effective displacement set is obtained by multiplying the anchor point displacement field with the noise mask and the core region mask respectively and taking the intersection. ; Summing the displacement vectors of all anchor points within the effective displacement set, dividing by the number of anchor point displacements in the effective displacement set, and then multiplying by the calibration coefficients for anchor point displacements and actual forces, yields the real-time estimate of the resultant force. : ; in, This represents the number of effective displacement sets after double masking. Indicates in Set of effective displacements at any given time The Middle The displacement vector of each anchor point.

4. The method for anti-jamming shaft hole assembly based on visual-tactile perception of unknown hand posture according to claim 3, characterized in that, In step 4, the constant contact force applied by the shaft to the worktable surface is achieved using an admittance control strategy, specifically as follows: An admittance control model incorporating virtual mass, virtual damping, and virtual stiffness coefficients is constructed. The difference between the estimated target contact force and the resultant force is calculated as the force input. The target acceleration is obtained by solving the admittance control differential equation. The velocity is obtained by one integration. The position correction is then updated in the discrete control system using Euler integration. Finally, the position is updated in the discrete control system using Euler integration.

5. The method for anti-jamming shaft hole assembly based on visual-tactile perception of unknown hand posture according to claim 4, characterized in that, In step 5, the specific judgment and execution logic for normal escape processing is as follows: During the spiral hole-finding process, the angle between the resultant force direction estimated by the second fingertip visual tactile sensor and the main axis direction of the shaft is determined. If the angle is continuously greater than the set threshold, it is determined that the shaft encounters a discontinuous protrusion obstacle, and the robotic arm is controlled to retract inward along the spiral normal to avoid the obstacle. During the insertion process, if the estimated resultant force drops sharply due to the complex texture of the shaft end surface and fails to reach the target insertion depth, the machine enters a stuck state. At this time, admittance control is maintained to control the end of the robotic arm to move outward along the normal of the current position of the spiral until the resultant force drops sharply again and the insertion state is re-entered. When attempting to release the stuck state, once the robotic arm moves to the threshold, it lifts the shaft vertically upward along the hole axis, leaving the hole position and re-entering the hole-finding state.

6. The method for anti-jamming shaft hole assembly based on visual-tactile perception of unknown hand posture according to claim 5, characterized in that, Using the pixel-to-physical-size scaling factor known from the first fingertip visual-tactile sensor, a homogeneous transformation matrix is ​​constructed from the coordinate system of the first fingertip visual-tactile sensor to the straight line modeled by the screw spindle. : ; in, Indicates the rotation angle is The rotation matrix, This refers to the translation correction vector in the image space calculated in step 2. The pixel-to-physical-size ratio of the first fingertip visual-tactile sensor. This represents the translation matrix.

7. The method for anti-jamming shaft hole assembly based on visual-tactile perception of unknown hand posture according to claim 2, characterized in that, The formula for calculating the corresponding percentile of the total map depth value set in step 2 is as follows: ; in, For the set adaptive depth threshold, This means finding the i-th element in the set. Percentile function The pixel coordinates in the tactile depth image are Depth value at that location, This represents the coordinates of all valid pixels in the tactile depth image.

Citation Information

Patent Citations

  • A robot shaft hole assembly system and method based on teaching-learning

    CN109382828B

  • Robot shaft hole assembling method based on deep reinforcement learning and admittance control

    CN115674204A

  • Axle hole operation fine perception method based on visual tactile perception

    CN119098986A

  • Robot control system and method based on visual sense and force sense fusion and robot

    CN121403414A

  • Screw / thread bushing mounting method based on visual guidance

    CN121491965A