Visual positioning method and system for precise connector assembly

Through multimodal image fusion and deep learning technology, combined with the collaborative processing of vision and force control systems, high-precision, real-time positioning and assembly of precision connectors are achieved, solving the accuracy and real-time problems of traditional methods in complex environments and improving assembly efficiency and stability.

CN120707802AInactive Publication Date: 2025-09-26DONGGUAN HAIHONG INTELLIGENT TECH CO LTD
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510815958.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve high-precision, real-time positioning and assembly of precision connectors in complex environments, especially when lighting conditions change, backgrounds are complex, and connector structures are diverse. Traditional methods are unable to accurately identify and correct tiny deviations in real time.

Method used

Multimodal HDR fusion and distortion correction processing are adopted, combined with a composite convolutional neural network for coarse positioning of feature areas. Through precise positioning of geometric features and hand-eye calibration, the position and pose instructions of the robot end effector are generated, and vision-force control hybrid assembly processing is performed.

Benefits of technology

It improves the speed and accuracy of precision connector assembly, meets the high quality and high efficiency requirements of modern automated production, and achieves adaptability and real-time performance in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707802A_ABST
    Figure CN120707802A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of visual positioning, and provides a visual positioning method and system for precise connector assembly. Performing multi-mode HDR fusion and distortion correction on the original connector image to obtain a multi-layer fusion image matrix; and performing feature region coarse positioning in combination with a composite convolutional neural network to obtain a connector coordinate set, performing geometric feature fine positioning according to the connector coordinate set and the multilayer fusion image matrix to obtain a key point coordinate set, and performing hand-eye calibration and dynamic compensation on the key point coordinate set to generate a pose instruction. And performing vision-force control hybrid assembly processing on the pose instruction according to the sensor feedback data to obtain assembly completion state data. Through multi-modal image fusion, deep learning coarse positioning, precise geometric registration and dynamic cooperation of vision and force control, the speed, precision and stability of precise connector assembly are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of visual positioning technology, and in particular to a visual positioning method and system for precision connector assembly. Background Art

[0002] As a key electronic component in high-end manufacturing, the assembly process of precision connectors places extremely high demands on positioning accuracy, force control coordination, and environmental adaptability. Connector pins have a small pitch and a complex structure, and even the slightest deviation can cause pin bending, plugging failure, and even system failure. Therefore, the assembly system must not only achieve high-precision positioning within the submillimeter level, but also perform real-time calibration of the connector posture to ensure that the plugging direction is highly consistent with the connection end. In addition, considering the high-speed operation requirements of actual production lines, the assembly method must also have high robustness and response speed to adapt to switching between different batches and different types of connectors, and meet the dual requirements of flexible manufacturing and stable quality.

[0003] Currently, traditional precision connector assembly methods are mostly based on feature extraction from a single image combined with fixed template matching for positioning. After obtaining the approximate position of the connector through a single visual recognition system, the mechanical execution system performs the plug-in action. This method usually uses fixed threshold segmentation, edge detection, or a specific template matching algorithm, which is feasible under ideal lighting and background conditions. Although some high-end assembly lines have introduced robots and calibration-based hand-eye systems, they are still unable to achieve accurate recognition and real-time correction of complex backgrounds, diverse connector structures, and small deviations due to the limitations of image input dimensions and processing algorithms. If force control systems are used, they are mostly separate modules and fail to form an effective linkage with the vision system. Summary of the Invention

[0004] In view of this, the present application provides a visual positioning method and system for precision connector assembly to solve the problem of poor adaptability to environmental changes.

[0005] A first aspect of the present application provides a visual positioning method for precision connector assembly, the method comprising: Perform multimodal HDR fusion and distortion correction on the collected original connector image to obtain a multi-layer fused image matrix; Performing coarse positioning of feature regions on the multi-layer fused image matrix through a preset composite convolutional neural network to obtain a connector coordinate set; Performing precise positioning of geometric features according to the connector coordinate set and the multi-layer fused image matrix to obtain a key point coordinate set; Performing hand-eye calibration and dynamic compensation processing on the key point coordinate set to generate a pose instruction in the robot end-effector coordinate system; The position and posture instructions are subjected to vision-force control hybrid assembly processing according to the sensor feedback data collected in real time to obtain assembly completion status data.

[0006] In an optional embodiment, performing multimodal HDR fusion and distortion correction processing on the collected original connector image to obtain a multi-layer fused image matrix includes: The collected original connector images are weighted fused and brightness normalized according to a preset exposure weight coefficient group to obtain a preliminary HDR grayscale image; Performing geometric correction on the preliminary HDR grayscale image according to preset camera intrinsic parameters, preset radial distortion coefficients, and preset tangential distortion coefficients to obtain a distortion-free fused image matrix; Non-local mean filtering is performed on the distortion-free fused image matrix to obtain a multi-layer fused image matrix.

[0007] In an optional embodiment, performing coarse positioning of feature regions on the multi-layer fused image matrix using a preset composite convolutional neural network to obtain a connector coordinate set includes: Performing feature extraction and attention fusion processing on the multi-layer fused image matrix through a preset composite convolutional neural network to obtain an enhanced feature atlas; Performing collaborative processing of the classification branch and the regression branch on each feature map in the enhanced feature map set one by one to obtain an original candidate bounding box and a corresponding confidence value; Performing dynamic threshold calculation based on the detected real-time contrast to obtain a minimum confidence threshold, and screening the original candidate bounding boxes based on the minimum confidence threshold and the confidence value to obtain a candidate bounding box set; Non-maximum suppression is performed on the candidate bounding box set to obtain a connector coordinate set.

[0008] In an optional embodiment, performing geometric feature precise positioning according to the connector coordinate set and the multi-layer fused image matrix to obtain a key point coordinate set includes: Performing ROI region cropping processing on the multi-layer fused image matrix according to the connector coordinate set to obtain a candidate sub-atlas; Calculating the Hessian matrix and grayscale gradient of each candidate subimage, and performing edge extraction on each candidate subimage based on the Hessian matrix and the grayscale gradient to obtain an edge point set corresponding to each candidate subimage; Calculating the polarization ratio of the ROI region in each candidate sub-image, and performing SIFT feature description on each candidate sub-image according to the polarization ratio to generate a multimodal feature description subset; Performing Kalman filtering fusion on the edge point set and the multimodal feature description subset to obtain a feature state sequence; A thermal drift compensation process is performed on the coordinates in the characteristic state sequence according to the detected temperature deviation data to obtain a key point coordinate set.

[0009] In an optional embodiment, performing hand-eye calibration and dynamic compensation processing on the key point coordinate set to generate a pose instruction in the robot end effector coordinate system includes: Solving the camera observation space pose of the manipulator using a preset PnP algorithm according to a preset physical three-dimensional reference point set and the key point coordinate set to obtain observation pose data; Performing a calibration conversion solution on multiple sets of samples based on the collected actual posture data of the robotic arm and the observed posture data using a preset Tsai–Lenz algorithm to obtain an initial hand-eye transformation matrix; Performing FFT spectrum analysis based on the detected acceleration of the end of the robotic arm to obtain the main vibration frequency and phase; Performing dynamic compensation iterative optimization processing on the initial hand-eye transformation matrix according to the main vibration frequency and the phase through a preset vibration feedforward model to obtain a real-time hand-eye transformation matrix; Mapping the key point coordinate set according to the real-time hand-eye transformation matrix to obtain a three-dimensional posture point coordinate set in the robot end effector coordinate system; The pose instructions in the robot end effector coordinate system are generated according to the preset assembly trajectory planning data and the three-dimensional pose point coordinate set.

[0010] In an optional embodiment, performing vision-force hybrid assembly processing on the posture instruction according to the real-time collected sensor feedback data to obtain assembly completion status data includes: Calculating the real-time expected acceleration of each joint according to the posture instruction using a preset six-degree-of-freedom impedance model to obtain the acceleration instruction of each joint; Performing segmented planning of the interlocking stroke and setting of control thresholds according to the acceleration command to obtain an overload threshold set; Performing control deviation correction based on the real-time collected sensor feedback data and the overload threshold set, and performing closed-loop control judgment based on the sensor feedback data and a preset judgment threshold to obtain an assembly completion signal; The assembly completion signal is subjected to abnormality detection and data recording processing to obtain assembly completion status data.

[0011] A second aspect of the present application provides a visual positioning device for precision connector assembly, the device comprising: The multi-layer fusion module is used to perform multi-modal HDR fusion and distortion correction on the collected original connector images to obtain a multi-layer fused image matrix; A coarse positioning module, configured to perform coarse positioning of feature regions on the multi-layer fusion image matrix through a preset composite convolutional neural network to obtain a connector coordinate set; A precise positioning module, configured to perform precise positioning of geometric features based on the connector coordinate set and the multi-layer fused image matrix to obtain a key point coordinate set; a posture control module, configured to perform hand-eye calibration and dynamic compensation processing on the key point coordinate set to generate posture instructions in the robot end effector coordinate system; The feedback adjustment module is used to perform vision-force control hybrid assembly processing on the posture instruction according to the sensor feedback data collected in real time to obtain assembly completion status data.

[0012] A third aspect of the present application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the visual positioning method for precision connector assembly as described above when executing the computer program.

[0013] A fourth aspect of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the visual positioning method for precision connector assembly as described above.

[0014] In summary, this application has at least the following beneficial technical effects: 1. The composite convolutional neural network is used to roughly locate the feature area of ​​the multi-layer fusion image matrix, which can quickly extract the approximate area of ​​the connector, greatly improving the processing speed and meeting real-time requirements.

[0015] 2. By calculating the mapping relationship between the camera coordinate system and the robot TCP coordinate system, the hand-eye calibration of the key point coordinate set is performed to achieve system integrated calibration.

[0016] 3. Collaborative visual positioning and real-time monitoring of contact forces during assembly automatically adjust feed or offset to avoid jamming or collisions caused by minor deviations. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 This is a flow chart of a visual positioning method for precision connector assembly provided in an embodiment of the present application; Figure 2 This is a functional module diagram of a visual positioning device for precision connector assembly provided in an embodiment of the present application; Figure 3 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0019] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0020] like Figure 1 FIG. 1 is a flowchart of a method for visually locating a precision connector assembly according to an embodiment of the present invention. The method for visually locating a precision connector assembly according to an embodiment of the present invention includes the following steps.

[0021] Step S1: Perform multimodal HDR fusion and distortion correction processing on the collected original connector image to obtain a multi-layer fused image matrix.

[0022] It should be understood that the original connector image includes an image of the male end of the connector and an image of the female end of the connector. The electronic device of the present application embodiment uses a coaxial red light source (wavelength 620nm) and a circular polarized light source (polarization angle 60°) to coordinately activate to eliminate specular reflections on the metal surface of the connector and enhance edge contrast. The global shutter camera module is controlled to have an exposure time of 0.8ms to trigger the synchronous capture of three frames of the original connector image {I low ,I mid ,I high}. Among them, I low This is a low-exposure image, used to avoid blown highlights. mid This is a medium exposure image, used to preserve the details of the middle grayscale. high It is a high-exposure image used to improve the signal-to-noise ratio of dark areas. Based on the mapping relationship between the preset exposure threshold and the exposure weight coefficient, the electronic device can automatically obtain the exposure weight coefficient corresponding to each frame of the original connector image to obtain the exposure weight coefficient group for the three-frame image fusion. At the same time, for each pixel position (x, y), the local contrast metric C of the position in each frame is calculated. k (x, y), and then the contrast weighting factor C of each frame image is obtained according to the exposure weight coefficient group obtained. kEach frame is weightedly fused to eliminate inter-frame brightness offsets. The fused image is then brightness-normalized to obtain a preliminary high-dynamic range (HDR) grayscale image. Through fusion and brightness normalization, both highlight information and shadow detail are preserved, addressing the overexposure or underexposure often seen on metal-plated surfaces with a single exposure. For example, on the metal housing of a connector, a single high-exposure frame often results in complete saturation of the center area, making it impossible to distinguish the plug scale, while a single low-exposure frame fails to capture the subtle scale texture. Weighted fusion allows both the scale and the reflective edge to be presented simultaneously in the same image.

[0023] After obtaining the preliminary HDR grayscale image, it is necessary to correct the geometric distortion of the preliminary HDR grayscale image to ensure that the straight line features and known geometric landmarks in the image correspond to the actual physical positions. Specifically, the correction process is done by using the preset camera internal parameters (including focal length fx, fy and principal point position (C x ,C y ), radial distortion coefficients k1 and k2 (used to correct for distortion caused by the distance from the lens center to the edge), and tangential distortion coefficients p1 and p2 (used to correct for non-concentricity caused by lens assembly errors). Zhang Zhengyou's calibration method is used to perform inverse distortion mapping on each pixel in normalized plane coordinates (u, v). This restores all curved lines and deformed geometric patterns to their actual physical proportions and positions, ensuring that all subsequent measured lengths and angles are consistent with the real objects. For example, at the edge of the camera's field of view, an uncorrected image would make the connector pins appear bent toward the center, while after correction, the pins appear completely straight, consistent with the robot's actual trajectory planning.

[0024] Since the distortion-free fused image matrix after dedistortion may still retain sensor noise or compression artifacts, the embodiment of the present application uses non-local mean filtering technology on the distortion-free fused image matrix to take into account both noise suppression and detail retention. Non-local mean filtering weighted averages the pixel value by comparing the similarity between each pixel p in the entire image and the neighborhood N(q) within the search window Ω of the entire image. It can retain texture and edge information while smoothing noise. It is particularly suitable for suppressing random spots generated by sensor thermal noise and quantization errors without losing the clarity of the edge of the metal plug. For example, when the camera generates tiny noise during high-speed acquisition, the algorithm can identify neighborhood blocks similar to the connector outline and fuse them preferentially. The resulting multi-layer fused image matrix not only has a wide dynamic range, but also maintains high fidelity in complex reflections and subtle textures.

[0025] Step S2: performing coarse positioning of feature regions on the multi-layer fusion image matrix using a preset composite convolutional neural network to obtain a connector coordinate set.

[0026] It should be understood that the composite convolutional neural network uses CSPDarknet53-tiny as the backbone to perform hierarchical extraction of low-level features such as edges, corners, and textures of the original image, thereby dividing the input image into feature subgraphs of different resolutions. The Spatial Pyramid Pooling-Fast (SPPF) module performs multi-scale pooling on features of different scales to maintain spatial information. The Path Aggregation Network (PAN) is used to fuse shallow and deep features to construct intermediate features that contain both local details and global semantics. After the multi-layer fused image matrix is ​​input into the pre-trained composite convolutional neural network, after low-level feature extraction, feature pooling, and fusion, feature maps of five scales are generated from the multi-layer fused image matrix according to the hierarchical structure. At the same time, the composite convolutional neural network embeds the Squeeze-Excitation (SE) attention mechanism on the feature map of each scale through the Squeeze-Excitation (SE) module to compress the channel dimension into a one-dimensional description vector through global average pooling. The vector is then mapped back to the channel weights through a two-layer fully connected network and Sigmoid activation, and then multiplied with the original feature map channel by channel to highlight the key channels related to the connector structure and suppress background interference. The composite convolutional neural network structure used in the embodiment of the present application enables accurate capture of multi-scale information of the connector body area under complex reflection and occlusion conditions.

[0027] After obtaining the enhanced feature atlas through the composite convolutional neural network, the classification branch and the regression branch are collaboratively reasoned for each feature map in turn. Among them, the classification branch predicts the existence probability of each preset anchor frame with the convolution layer, and the regression branch outputs the center offset and length and width scaling of the anchor frame, and then calculates the original candidate bounding box and its confidence value based on the reference anchor frame. Among them, the reference anchor frame is pre-set according to the typical external dimensions of the connector on the production line to take into account the size difference between the male and female ends of the plug. And the classification and regression are completed independently in different channels without interfering with each other to ensure detection accuracy and positioning stability. Subsequently, in order to eliminate false alarms or missed detections caused by changes in illumination, it is necessary to dynamically calculate the minimum confidence threshold to screen candidate frames. Specifically, based on the overall gradient statistics of the current image, the standard deviation σ of the gradient amplitudes in all candidate frames is first calculated. C and the average confidence value , the minimum confidence threshold is calculated using the threshold function shown below: Where K is the contrast sensitivity coefficient. CIt is used to measure the fluctuation amplitude of pixel gradients within the candidate frame. η is a safety offset that ensures that a small number of frames are retained for subsequent analysis when the global contrast is extremely low. The exponential decay function is used to give high-contrast areas a higher threshold gain, thereby strictly screening the confidence level when the contour is clear and preventing interference from surrounding reflective areas; at the same time, when the contrast is low, due to Close to zero, the threshold is still determined by the baseline confidence +η is maintained to avoid missed detection. Only the confidence value p k Only original candidate bounding boxes with a value greater than or equal to the threshold T are retained and included in the candidate bounding box set.

[0028] Finally, non-maximum suppression is used to eliminate highly overlapping, redundant boxes. First, candidate boxes are sorted from high to low confidence. Then, each box is traversed and the intersection over union (IoU) with the remaining boxes is calculated. If the IoU of a pair of boxes exceeds a preset overlap threshold, the one with the lower confidence is removed. This process continues until the traversal is complete. Non-maximum suppression is performed on the candidate bounding box set to obtain the connector coordinate set. This process ensures that only the most representative bounding boxes are retained in scenarios where the male and female terminals are adjacent and have significant edge overlap, ensuring that both male and female terminals are accurately delineated. Finally, coordinates are extracted from the candidate bounding box set after overlapping elimination to obtain the connector coordinate set.

[0029] Step S3: performing precise positioning of geometric features based on the connector coordinate set and the multi-layer fused image matrix to obtain a key point coordinate set.

[0030] First, a candidate sub-image set for each connector region is extracted from the multi-layer fusion image matrix. The region cropping is directly mapped to the image pixel system using the frame center coordinates and size parameters in the connector coordinate set. By intercepting the corresponding columns and rows in the horizontal and vertical coordinate directions of the image matrix, a candidate sub-image containing only the male or female end of the connector is obtained, thereby eliminating background interference and significantly reducing the amount of subsequent calculations. For each cropped sub-image, sub-pixel edge extraction is performed by combining the curvature information based on the Hessian matrix and the edge intensity calculation based on the grayscale gradient. Specifically, by solving the Hessian matrix eigenvalues ​​λ1, λ2 and the gradient component I at the pixel point (u, v), the sub-pixel edge extraction is performed. u , I v And introduce the adaptive fusion index α(u,v) to construct the edge response function as shown below: in, The fusion index α(u,v) is dynamically adjusted based on the ratio of the local second-order curvature to the first-order gradient, prioritizing curvature information in smooth surface areas and gradient information in textured areas. This response function enhances detection in areas of high curvature, such as metal edges and pin notches, while maintaining robustness in shadows or low-contrast areas. After performing maximum value detection on R(u,v) and refining it through sub-pixel interpolation, a dense set of edge points corresponding to each sub-image is obtained.

[0031] On this basis, it is necessary to combine polarization imaging to enhance the texture description capability. The polarization ratio ρ of two orthogonal polarization sub-images generated by a circularly polarized light source is calculated to quantify the polarization ratio of the reflection difference of the metal surface in different polarization directions, so that the reflective texture information can be included when generating feature descriptors in areas with tiny scratches or coating thickness changes. Among them, the polarization ratio ρ is calculated by continuously irradiating the same candidate sub-image with two orthogonal polarization angles of 0° and 90° by a circularly polarized light source, and synchronously acquiring two frames of grayscale images I 0° , I 90° For each pixel position (u,v) in the image, its polarization ratio ρ(u,v) is I 0° -I 90° with I 0° +I 90° The ratio of molecule I 0° -I 90° It reflects the intensity difference of the point in the two polarization directions. The larger the difference, the more obvious the surface polarization feature. 0° +I 90° The overall brightness is normalized to prevent brightness changes from misleading the polarization ratio. To avoid the denominator being zero, the denominator I 0° +I 90° A small value ε will be superimposed. After obtaining the polarization ratio ρ, for each key point position detected based on the scale-invariant feature transform (SIFT) method, first extract the 128-dimensional grayscale descriptor d SIFT , and then concatenate the polarization ratios to form a 129-dimensional multimodal descriptor d multi =[d SIFT ,ρ] and saved together with its pixel coordinates to the multimodal feature descriptor subset.

[0032] The Kalman filter is then used to perform temporal fusion of the edge point set and the multimodal feature description subset to eliminate transient noise and estimate the dynamic pose. All feature points are regarded as observations from the same motion model, and the state vector in the feature state sequence is defined as X=[x,y,θ,s] T. Where (x, y) represents the position of the feature point in the image coordinate system, θ is the local rotation angle, and s is the scale factor. The embodiment of the present application adopts a prediction model to perform a time series fusion operation. The prediction model adopts a constant speed motion assumption, and the measurement model integrates the relative displacement generated by the edge response position and the descriptor matching, and fuses the predicted state with the measurement information through the standard Kalman filter update algorithm to obtain the optimal estimate. Specifically, first, the estimated X of the previous cycle is obtained by using the preset state transfer matrix F. t-1 Make a linear prediction and get a priori estimate , and update the estimated error covariance at the same time . Where Q describes the uncertainty of the system motion model. Then the measurement vector z t (composed of edge point detection and multimodal descriptor matching results) and the prior estimate are mapped to the same space, and the expected measurement is calculated by the observation matrix H , the difference between the two is taken as the measurement residual. Further, according to Calculate the Kalman gain, and then use the Kalman gain to modify the prior estimate of the residual weight. And update the covariance to By performing prediction and measurement updates sequentially within each frame loop, the system can capture dynamic changes while smoothing random noise, making keypoint position and orientation estimates more stable and accurate. This fusion operation first uses edge observations for updates within each frame loop, followed by a second update using descriptor observations. This allows the system to maintain tracking of feature points when they are temporarily lost due to lighting or occlusion, using motion prediction and historical information, and quickly correct them when the features reappear.

[0033] Considering that temperature fluctuations in the assembly environment may cause micron-level calibration drift, after obtaining the current temperature deviation ΔT from the external temperature sensor, the thermal expansion compensation matrix is ​​used to correct the coordinates in the characteristic state sequence output by the Kalman filter one by one. O Thermal expansion compensation is superimposed on it. After applying this transformation to each state vector (x, y, θ) output by the Kalman filter, the coordinates of the key points in the physical coordinate system can be obtained. The corrected pixel coordinates can be mapped to the actual physical size, and then summarized into a set of key point coordinates that meet the ±0.1μm level accuracy requirement. Among them, the thermal expansion compensation matrix corrects the deviation of the pixel scale caused by the change of ambient temperature based on the thermal expansion principle of the material. If the original calibration matrix (i.e., the camera calibration matrix H O ) is in homogeneous form, then the diagonal compensation matrix D(ΔT) can be constructed based on the current temperature deviation ΔT and the linear thermal expansion coefficients α and β in the horizontal and vertical directions. The state vector (x, y, θ) obtained by the Kalman filter is multiplied on the left by the diagonal compensation matrix D(ΔT) to complete the coordinate micro-scaling correction.

[0034] Step S4: performing hand-eye calibration and dynamic compensation processing on the key point coordinate set to generate a pose instruction in the robot end effector coordinate system.

[0035] First, a one-to-one mapping relationship is established between a set of physical 3D reference points mounted on the fixture and the corresponding key point coordinates on the image plane through the Perspective-n-Point (PnP) algorithm, and the rotation matrix R of the camera in the reference point system is calculated by minimizing the reprojection error. c With the translation vector t c This process ensures the correspondence accuracy between the image coordinates and the real space position. For example, the reference point at the scale hole on the side of the connector male end can be accurately mapped to the millimeter-level space position. Then, the actual pose data {(R m,k ,t m,k )} and the corresponding camera observation pose {(R c,k ,t c,k )} as a sample, and the Tsai–Lenz algorithm is used to solve the rigid body equation , to obtain the hand-eye transformation matrix X0 (i.e., the initial hand-eye transformation matrix) that meets the requirements of the rigid body equation, thereby obtaining the initial rigid body mapping relationship between the camera and the tool end of the robotic arm. Then, the acceleration signal a(t) continuously sampled by the accelerometer at the end of the robotic arm is fast Fourier transformed, and the main vibration frequency f0 and its corresponding phase Φ=arg A(f0) are identified in the spectrum to characterize the main oscillation mode of the assembly platform. Then, a vibration feedforward model is constructed based on the main vibration frequency f0 and its corresponding phase Φ to generate a six-dimensional torsion vector ξ(t) synthesized by the compensation torsion and displacement, and it is applied to the initial hand-eye transformation matrix through the Lie group-Lie algebra exponential mapping to obtain the real-time hand-eye transformation matrix X dyn Vibration feedforward dynamic compensation enables continuous compensation of both rotation and displacement, allowing the robot end to maintain micron-level stability under vibration.

[0036] Subsequently, the pixel coordinates in the keypoint coordinate set are mapped into the robot base coordinate system in the form of homogeneous vectors via a real-time matrix to obtain a series of 3D pose point coordinates. These 3D pose points constitute the spatial path that the end effector should reach. Finally, combined with the pre-set assembly trajectory planning data, an interpolation algorithm and velocity-acceleration constraint planning are used to convert the 3D pose point coordinate set composed of all the 3D pose point coordinates into detailed pose instructions for the robot end effector. These pose instructions include pose quaternions or Euler angles (α, β, γ) and Cartesian coordinates (X, Y, Z), which are then sent to the robot arm's servo driver in real time.

[0037] Step S5: performing vision-force control hybrid assembly processing on the posture instruction according to the sensor feedback data collected in real time to obtain assembly completion status data.

[0038] After collecting the sensor feedback data of the six-axis force sensor and motor drive current feedback of the robot end effector in real time, the posture command is first converted into the real-time expected acceleration of each joint based on the built-in six-degree-of-freedom impedance model to obtain the acceleration command for driving the motor. The impedance model defines the dynamic relationship between the actuator and the environment with the inertia matrix M, the damping matrix C and the stiffness matrix K. By solving When a small position deviation occurs or an external force disturbance occurs, the system can quickly generate compensation instructions by calculating the corresponding expected acceleration, thereby achieving flexible control of the insertion process. Where e represents the difference between the expected and actual postures. and are velocity error and acceleration error respectively. ext For example, during the alignment phase before the male connector is inserted into the female connector, the stiffness K can be set to a higher value to ensure quick alignment, while during the fully mated phase, the stiffness can be adjusted to a lower value to protect the pins from damage.

[0039] After receiving acceleration commands for each joint, the entire mating path is divided into three intervals: pre-alignment, initial contact, and full mating, according to the three pre-defined mating travel paths in trajectory planning. For each interval, corresponding upper and lower acceleration limits and maximum allowable external force thresholds are set, forming a segmented control threshold set. In the pre-alignment interval, when the detected external force does not exceed 1N and the velocity deviation exceeds 0.1mm / s, the system increases stiffness to quickly correct the posture. In the initial contact interval, the external force threshold is increased to 3N, and the damping C is increased to prevent excessive pin tilting due to sudden impact. In the full mating interval, the external force threshold is further adjusted to 5N to ensure that the plug does not loosen due to minor vibrations after mating.

[0040] Subsequently, the current acceleration command is dynamically corrected based on the sensor feedback force and current signal (i.e., sensor feedback data). ext Or when the motor current change rate dI / dt exceeds the threshold set of each segment, the acceleration command is modified to introduce reverse micro-motion to pull the actual action back to the set trajectory. The closed-loop control judgment outputs an assembly completion signal when two conditions are met simultaneously: one is that the axial force component of the insertion exceeds 3.5N, indicating that the pin has reliably contacted the bottom of the socket; the second is that the current change rate is less than 0.1A / s, which means that the motor has entered steady-state drive and no longer generates a significant current increase due to continued force. This dual-condition judgment effectively avoids premature termination due to single force feedback anomalies or motor overload false alarms, thereby ensuring assembly quality and mechanical safety.

[0041] After the assembly completion signal is issued, the system immediately triggers anomaly detection and data recording processing. Anomaly detection will automatically identify whether there are faults such as material jamming or excessive deflection based on the historical threshold model of the force-position curve and the current curve, and issue an alarm or initiate a return action when necessary; data recording will store the joint posture, force sensing and current data of each joint during the entire insertion process in local and cloud databases in a timestamp format, forming a traceable quality and fault log. For example, in mass production, when the assembly yield of a batch decreases, the operator can retrieve the complete force-position curve of the batch to determine the cause of the fault and adjust the parameters in time. The final output assembly completion status data not only contains a simple completion mark, but also comes with detailed process records, providing a reliable basis for subsequent quality analysis and process optimization.

[0042] This application is applied to the field of visual positioning technology, performing multimodal HDR fusion and distortion correction on the original connector image to obtain a multi-layer fused image matrix; combining a composite convolutional neural network to perform coarse positioning of feature areas to obtain a connector coordinate set, performing fine positioning of geometric features based on the connector coordinate set and the multi-layer fused image matrix to obtain a key point coordinate set, performing hand-eye calibration and dynamic compensation on the key point coordinate set to generate posture instructions, and performing vision-force control hybrid assembly processing on the posture instructions based on sensor feedback data to obtain assembly completion status data. This application improves the speed, accuracy, and stability of precision connector assembly through multimodal image fusion, deep learning coarse positioning, precise geometric alignment, and dynamic coordination of vision and force control, meeting the stringent requirements of modern automated production for high-quality and efficient assembly.

[0043] like Figure 2 , which is a functional module diagram of a visual positioning device for precision connector assembly provided in an embodiment of the present application.

[0044] In some embodiments, the visual positioning device 2 for precision connector assembly may include multiple functional modules composed of computer program segments. The computer program of each program segment in the visual positioning device 2 for precision connector assembly may be stored in a memory of a server and executed by at least one processor to execute (see Figure 1 Description) Capabilities of vision positioning methods for precision connector assembly.

[0045] In this embodiment, the visual positioning device 2 for precision connector assembly can be divided into multiple functional modules based on the functions they perform. These functional modules may include a multi-layer fusion module 21, a coarse positioning module 22, a precise positioning module 23, a posture control module 24, and a feedback adjustment module 25. As used herein, a module refers to a series of computer program segments that can be executed by at least one processor and perform a fixed function, and are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0046] The multi-layer fusion module 21 is used to perform multi-modal HDR fusion and distortion correction processing on the collected original connector image to obtain a multi-layer fused image matrix.

[0047] In an optional embodiment, the multi-layer fusion module 21 is specifically configured to: The collected original connector images are weighted fused and brightness normalized according to a preset exposure weight coefficient group to obtain a preliminary HDR grayscale image; Performing geometric correction on the preliminary HDR grayscale image according to preset camera intrinsic parameters, preset radial distortion coefficients, and preset tangential distortion coefficients to obtain a distortion-free fused image matrix; Non-local mean filtering is performed on the distortion-free fused image matrix to obtain a multi-layer fused image matrix.

[0048] The rough positioning module 22 is used to perform rough positioning of feature areas on the multi-layer fusion image matrix through a preset composite convolutional neural network to obtain a connector coordinate set.

[0049] In an optional embodiment, the coarse positioning module 22 is specifically configured to: Performing feature extraction and attention fusion processing on the multi-layer fused image matrix through a preset composite convolutional neural network to obtain an enhanced feature atlas; Performing collaborative processing of the classification branch and the regression branch on each feature map in the enhanced feature map set one by one to obtain an original candidate bounding box and a corresponding confidence value; Performing dynamic threshold calculation based on the detected real-time contrast to obtain a minimum confidence threshold, and screening the original candidate bounding boxes based on the minimum confidence threshold and the confidence value to obtain a candidate bounding box set; Non-maximum suppression is performed on the candidate bounding box set to obtain a connector coordinate set.

[0050] The precise positioning module 23 is configured to perform precise positioning of geometric features according to the connector coordinate set and the multi-layer fusion image matrix to obtain a key point coordinate set.

[0051] In an optional embodiment, the precise positioning module 23 is specifically configured to: Performing ROI region cropping processing on the multi-layer fused image matrix according to the connector coordinate set to obtain a candidate sub-atlas; Calculating the Hessian matrix and grayscale gradient of each candidate subimage, and performing edge extraction on each candidate subimage based on the Hessian matrix and the grayscale gradient to obtain an edge point set corresponding to each candidate subimage; Calculating the polarization ratio of the ROI region in each candidate sub-image, and performing SIFT feature description on each candidate sub-image according to the polarization ratio to generate a multimodal feature description subset; Performing Kalman filtering fusion on the edge point set and the multimodal feature description subset to obtain a feature state sequence; A thermal drift compensation process is performed on the coordinates in the characteristic state sequence according to the detected temperature deviation data to obtain a key point coordinate set.

[0052] The posture control module 24 is used to perform hand-eye calibration and dynamic compensation processing on the key point coordinate set to generate posture instructions in the robot end effector coordinate system.

[0053] In an optional embodiment, the posture control module 24 is specifically configured to: Solving the camera observation space pose of the manipulator using a preset PnP algorithm according to a preset physical three-dimensional reference point set and the key point coordinate set to obtain observation pose data; Performing a calibration conversion solution on multiple sets of samples based on the collected actual posture data of the robotic arm and the observed posture data using a preset Tsai–Lenz algorithm to obtain an initial hand-eye transformation matrix; Performing FFT spectrum analysis based on the detected acceleration of the end of the robotic arm to obtain the main vibration frequency and phase; Performing dynamic compensation iterative optimization processing on the initial hand-eye transformation matrix according to the main vibration frequency and the phase through a preset vibration feedforward model to obtain a real-time hand-eye transformation matrix; Mapping the key point coordinate set according to the real-time hand-eye transformation matrix to obtain a three-dimensional posture point coordinate set in the robot end effector coordinate system; The pose instructions in the robot end effector coordinate system are generated according to the preset assembly trajectory planning data and the three-dimensional pose point coordinate set.

[0054] The feedback adjustment module 25 is used to perform vision-force control hybrid assembly processing on the posture instruction according to the sensor feedback data collected in real time to obtain assembly completion status data.

[0055] In an optional embodiment, the feedback adjustment module 25 is specifically configured to: Calculating the real-time expected acceleration of each joint according to the posture instruction using a preset six-degree-of-freedom impedance model to obtain the acceleration instruction of each joint; Performing segmented planning of the interlocking stroke and setting of control thresholds according to the acceleration command to obtain an overload threshold set; Performing control deviation correction based on the real-time collected sensor feedback data and the overload threshold set, and performing closed-loop control judgment based on the sensor feedback data and a preset judgment threshold to obtain an assembly completion signal; The assembly completion signal is subjected to abnormality detection and data recording processing to obtain assembly completion status data.

[0056] It should be understood that the various variations and specific embodiments of the methods provided in the above embodiments are also applicable to the visual positioning device for precision connector assembly of this embodiment. Through the above detailed description of the visual positioning method for precision connector assembly, those skilled in the art can clearly know the implementation method of the visual positioning device for precision connector assembly in this embodiment. For the sake of brevity of the specification, it will not be described in detail here.

[0057] like Figure 3 , which is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0058] In a preferred embodiment of the present invention, the electronic device 3 may include, but is not limited to: a memory 31 , at least one processor 32 and at least one communication bus 33 .

[0059] Those skilled in the art should understand that Figure 3 The structure of the electronic device 3 shown does not constitute a limitation of the embodiment of the present invention. The electronic device 3 may also include more or less other hardware or software than shown in the figure, or a different component arrangement.

[0060] In some embodiments, the electronic device 3 is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits, programmable gate arrays, digital processors and embedded devices.

[0061] It should be noted that the electronic device 3 is only an example. Other existing or future electronic products that are suitable for this application should also be included in the scope of protection of this application and included here by reference.

[0062] In some embodiments, the memory 31 stores a computer program that, when executed by the at least one processor 32, implements all or part of the steps in the visual positioning method for precision connector assembly. The memory 31 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disk storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data. Furthermore, the computer-readable storage medium may primarily include a program storage area and a data storage area, wherein the program storage area may store an operating system, at least one application required for a function, and the like.

[0063] In some embodiments, the at least one processor 32 is the control core (Control Unit) of the electronic device 3. It connects the various components of the entire electronic device 3 using various interfaces and lines, and executes the various functions and processes data of the electronic device 3 by running or executing the programs or modules stored in the memory 31, and calling the data stored in the memory 31. For example, when the at least one processor 32 executes the computer program stored in the memory 31, it implements all or part of the steps of the visual positioning method for precision connector assembly described in the embodiments of the present application; or implements all or part of the functions of the visual positioning device for precision connector assembly. The at least one processor 32 can be composed of an integrated circuit, for example, a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips.

[0064] In some embodiments, the at least one communication bus 33 is configured to enable communication between the memory 31 and the at least one processor 32. Although not shown, the electronic device 3 may also include a power supply (e.g., a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 32 via a power management device, thereby enabling the power management device to manage charging, discharging, and power consumption. The power supply may also include one or more DC or AC power supplies, a recharging device, a power failure detection circuit, a power converter or inverter, a power status indicator, and other components. The electronic device 3 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be detailed here.

[0065] The above-mentioned integrated unit implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module stored in a storage medium includes a number of instructions for causing an electronic device (which can be a personal computer, electronic device, or network device, etc.) or a processor to execute portions of the methods described in various embodiments of the present application.

[0066] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is only a logical function division, and other division methods may be used in actual implementation.

[0067] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, and may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected to achieve the purpose of this embodiment based on actual needs.

[0068] The above are all preferred embodiments of the present application, and are not intended to limit the scope of protection of the present application. Therefore, any equivalent changes made based on the structure, shape, and principle of the present application should be included in the scope of protection of the present application.

Claims

1. A visual positioning method for precision connector assembly, characterized in that: The method comprises: Perform multimodal HDR fusion and distortion correction on the collected original connector image to obtain a multi-layer fused image matrix; Performing coarse positioning of feature regions on the multi-layer fused image matrix through a preset composite convolutional neural network to obtain a connector coordinate set; Performing precise positioning of geometric features according to the connector coordinate set and the multi-layer fused image matrix to obtain a key point coordinate set; Performing hand-eye calibration and dynamic compensation processing on the key point coordinate set to generate a pose instruction in the robot end-effector coordinate system; The position and posture instructions are subjected to vision-force control hybrid assembly processing according to the sensor feedback data collected in real time to obtain assembly completion status data.

2. The visual positioning method for precision connector assembly according to claim 1, characterized in that: The multimodal HDR fusion and distortion correction processing of the collected original connector image to obtain a multi-layer fused image matrix includes: The collected original connector images are weighted fused and brightness normalized according to a preset exposure weight coefficient group to obtain a preliminary HDR grayscale image; Performing geometric correction on the preliminary HDR grayscale image according to preset camera intrinsic parameters, preset radial distortion coefficients, and preset tangential distortion coefficients to obtain a distortion-free fused image matrix; Non-local mean filtering is performed on the distortion-free fused image matrix to obtain a multi-layer fused image matrix.

3. The visual positioning method for precision connector assembly according to claim 1, characterized in that: The performing coarse positioning of feature regions on the multi-layer fusion image matrix by a preset composite convolutional neural network to obtain a connector coordinate set includes: Performing feature extraction and attention fusion processing on the multi-layer fused image matrix through a preset composite convolutional neural network to obtain an enhanced feature atlas; Performing collaborative processing of the classification branch and the regression branch on each feature map in the enhanced feature map set one by one to obtain an original candidate bounding box and a corresponding confidence value; Performing dynamic threshold calculation based on the detected real-time contrast to obtain a minimum confidence threshold, and screening the original candidate bounding boxes based on the minimum confidence threshold and the confidence value to obtain a candidate bounding box set; Non-maximum suppression is performed on the candidate bounding box set to obtain a connector coordinate set.

4. The visual positioning method for precision connector assembly according to claim 1, characterized in that: The performing geometric feature precise positioning according to the connector coordinate set and the multi-layer fused image matrix to obtain a key point coordinate set includes: Performing ROI region cropping processing on the multi-layer fused image matrix according to the connector coordinate set to obtain a candidate sub-atlas; Calculating the Hessian matrix and grayscale gradient of each candidate subimage, and performing edge extraction on each candidate subimage based on the Hessian matrix and the grayscale gradient to obtain an edge point set corresponding to each candidate subimage; Calculating the polarization ratio of the ROI region in each candidate sub-image, and performing SIFT feature description on each candidate sub-image according to the polarization ratio to generate a multimodal feature description subset; Performing Kalman filtering fusion on the edge point set and the multimodal feature description subset to obtain a feature state sequence; A thermal drift compensation process is performed on the coordinates in the characteristic state sequence according to the detected temperature deviation data to obtain a key point coordinate set.

5. The visual positioning method for precision connector assembly according to claim 1, characterized in that: The performing hand-eye calibration and dynamic compensation processing on the key point coordinate set to generate a pose instruction in the robot end effector coordinate system includes: Solving the camera observation space pose of the manipulator using a preset PnP algorithm according to a preset physical three-dimensional reference point set and the key point coordinate set to obtain observation pose data; Performing a calibration conversion solution on multiple sets of samples based on the collected actual posture data of the robotic arm and the observed posture data using a preset Tsai–Lenz algorithm to obtain an initial hand-eye transformation matrix; Performing FFT spectrum analysis based on the detected acceleration of the end of the robotic arm to obtain the main vibration frequency and phase; Performing dynamic compensation iterative optimization processing on the initial hand-eye transformation matrix according to the main vibration frequency and the phase through a preset vibration feedforward model to obtain a real-time hand-eye transformation matrix; Mapping the key point coordinate set according to the real-time hand-eye transformation matrix to obtain a three-dimensional posture point coordinate set in the robot end effector coordinate system; The pose instructions in the robot end effector coordinate system are generated according to the preset assembly trajectory planning data and the three-dimensional pose point coordinate set.

6. The visual positioning method for precision connector assembly according to claim 1, characterized in that: The performing vision-force control hybrid assembly processing on the posture instruction according to the real-time collected sensor feedback data to obtain assembly completion status data includes: Calculating the real-time expected acceleration of each joint according to the posture instruction using a preset six-degree-of-freedom impedance model to obtain the acceleration instruction of each joint; Performing segmented planning of the interlocking stroke and setting of control thresholds according to the acceleration command to obtain an overload threshold set; Performing control deviation correction based on the real-time collected sensor feedback data and the overload threshold set, and performing closed-loop control judgment based on the sensor feedback data and a preset judgment threshold to obtain an assembly completion signal; The assembly completion signal is subjected to abnormality detection and data recording processing to obtain assembly completion status data.

7. A visual positioning device for precision connector assembly, characterized in that: The device comprises: The multi-layer fusion module is used to perform multi-modal HDR fusion and distortion correction on the collected original connector images to obtain a multi-layer fused image matrix; A coarse positioning module, configured to perform coarse positioning of feature regions on the multi-layer fusion image matrix through a preset composite convolutional neural network to obtain a connector coordinate set; A precise positioning module, configured to perform precise positioning of geometric features based on the connector coordinate set and the multi-layer fused image matrix to obtain a key point coordinate set; a posture control module, configured to perform hand-eye calibration and dynamic compensation processing on the key point coordinate set to generate posture instructions in the robot end effector coordinate system; The feedback adjustment module is used to perform vision-force control hybrid assembly processing on the posture instruction according to the sensor feedback data collected in real time to obtain assembly completion status data.

8. An electronic device, characterized in that: The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the visual positioning method for precision connector assembly according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the visual positioning method for precision connector assembly according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Connection deviation detection method and device based on machine vision

    CN121190455A

  • Connection deviation detection method and device based on machine vision

    CN121190455B

  • Drug package online detection system based on three-dimensional reconstruction

    CN121366168A

  • Online detection system for medicine packaging based on three-dimensional reconstruction

    CN121366168B

  • Mechanical arm visual guidance method and system for industrial assembly

    CN121392226A