Joint manipulator automatic calibration method and device based on visual system

By combining the multi-view three-dimensional visual measurement network with global visual tracking and local visual measurement systems, the problem of insufficient accuracy of traditional robot calibration methods in complex environments is solved, high-precision joint error correction is achieved, and the positioning and motion accuracy of five-axis robot joint manipulators is improved.

CN120503211AInactive Publication Date: 2025-08-19深圳市远望工业自动化设备有限公司
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510980461.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-08-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The traditional robot calibration method relies on contact measurement tools and is difficult to meet high-precision requirements in complex environments. The single-view vision system cannot obtain complete joint information, resulting in insufficient calibration accuracy, and poor correlation between the global visual tracking system and the local measurement system data, so it is impossible to fully utilize multi-source information for accurate calibration.

Method used

A multi-view three-dimensional visual measurement network combining a global vision tracking system and a local vision measurement system is adopted to construct an adaptive calibration equation through multi-view image acquisition, occlusion area extraction, disordered image stitching and mark point feature extraction, to solve the target transformation relationship between the global and local visual coordinate systems, and combine joint chain constraint optimization to generate joint error correction parameters.

Benefits of technology

High-precision joint error correction parameter calculation is realized, the positioning accuracy and motion accuracy of the five-axis robot joint manipulator are improved, and the problem of insufficient calibration accuracy in complex environments is overcome, ensuring that the calibration results meet the physical characteristics and practicality of the joint mechanism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120503211A_ABST
    Figure CN120503211A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of visual automatic calibration, and discloses a joint manipulator automatic calibration method and device based on a visual system. The method comprises the following steps: carrying out multi-view image acquisition on a joint manipulator of the five-axis robot to obtain original image data; performing occlusion region extraction on the original image data to obtain a joint occlusion region set; on the basis of the joint occlusion region set, disordered image splicing and mark point feature extraction are carried out on the original image data, and joint feature image data are obtained; constructing an adaptive calibration equation based on the joint feature image data, and solving a target transformation relation between a global visual coordinate system and a local visual coordinate system; multi-joint collaborative calibration and joint chain constraint optimization of the five-axis robot are executed according to the target transformation relation, and joint error correction parameters are generated, high-precision joint error correction parameter calculation is achieved, and the positioning precision and the movement precision of a joint manipulator of the five-axis robot are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automatic vision calibration, and in particular to a method and device for automatic calibration of an articulated manipulator based on a vision system. Background Art

[0002] The precise calibration and adjustment of industrial robots, especially five-axis articulated manipulators, in complex industrial environments is crucial for improving production precision. Traditional manipulator calibration methods rely primarily on contact measurement tools and manual operation, which are not only inefficient but also susceptible to interference in complex environments, making them difficult to meet the demands of high-precision industrial applications. With the development of visual measurement technology, robot calibration methods based on vision systems have gained widespread application. However, in actual industrial environments, due to occlusion issues caused by equipment layout and operation, single-viewpoint vision systems struggle to obtain complete joint information, resulting in insufficient calibration accuracy and limiting the application scope of visual calibration technology.

[0003] As a highly complex spatial mechanism, the intercoupling and cumulative error of five-axis robotic articulated manipulators complicate accurate calibration. Existing methods for articulated manipulator calibration often rely on a single perspective to acquire information, making it difficult to effectively handle visual occlusion in complex environments, leading to inaccurate calibration. Furthermore, traditional methods lack effective coordinate system fusion mechanisms, resulting in poor data correlation between the global visual tracking system and the local measurement system, making it impossible to fully utilize multi-source information for accurate calibration. Calibration accuracy significantly decreases with changes in pose during joint motion, especially during joint motion. Summary of the Invention

[0004] The present invention provides a method and device for automatic calibration of a joint manipulator based on a vision system. The present invention realizes high-precision calculation of joint error correction parameters, greatly improving the positioning accuracy and motion accuracy of the five-axis robot joint manipulator.

[0005] In a first aspect, the present invention provides a method for automatic calibration of an articulated manipulator based on a visual system, the method comprising: Perform multi-view image acquisition on the joint manipulator of the five-axis robot to obtain original image data; Extracting occlusion regions from the original image data to obtain a set of joint occlusion regions; Based on the joint occlusion area set, performing disordered image splicing and marker point feature extraction on the original image data to obtain joint feature image data; Constructing an adaptive calibration equation based on the joint feature image data and solving a target transformation relationship between a global visual coordinate system and a local visual coordinate system; Multi-joint collaborative calibration and joint chain constraint optimization of the five-axis robot are performed according to the target transformation relationship to generate joint error correction parameters.

[0006] In a second aspect, the present invention provides an automatic calibration device for an articulated manipulator based on a visual system, the automatic calibration device for an articulated manipulator based on a visual system comprising: The acquisition module is used to acquire multi-view images of the joint manipulator of the five-axis robot to obtain original image data; An extraction module, configured to extract occlusion regions from the original image data to obtain a set of joint occlusion regions; A splicing module, configured to perform unordered image splicing and marker feature extraction on the original image data based on the joint occlusion region set to obtain joint feature image data; A construction module is used to construct an adaptive calibration equation based on the joint feature image data and solve the target transformation relationship between the global visual coordinate system and the local visual coordinate system; A generation module is used to perform multi-joint collaborative calibration and joint chain constraint optimization of the five-axis robot according to the target transformation relationship, and generate joint error correction parameters.

[0007] The technical solution provided by the present invention adopts a multi-view 3D visual measurement network that combines a global visual tracking system with a local visual measurement system to achieve all-round monitoring of the five-axis robot joint manipulator, effectively solving the problem of being unable to obtain complete information from a single viewpoint. The occluded area is finely extracted through a binary tree model algorithm, and the complete information of the occluded joint is restored by combining disordered image stitching technology. This overcomes the defect of insufficient calibration accuracy of traditional methods in complex environments and makes the calibration process no longer restricted by environmental occlusion. The introduction of a scale factor correction mechanism and a multi-pose sampling strategy solves the fusion problem between the global visual coordinate system and the local visual coordinate system, establishes an accurate coordinate mapping relationship, and realizes precise data integration between systems of different scales. Through the joint chain constraint optimization model, the calibration parameters are globally optimized in combination with geometric constraints and kinematic constraints, ensuring that the calibration results meet the physical characteristics of the joint mechanism and improving the physical significance and practicality of the calibration. The combination of recursive least squares method and nonlinear optimization algorithm realizes high-precision calculation of joint error correction parameters, significantly improving the positioning accuracy and motion accuracy of the five-axis robot joint manipulator. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0009] Figure 1 A schematic diagram of a flow chart of an automatic calibration method for an articulated manipulator based on a visual system according to an embodiment of the present application; Figure 2 This is a schematic block diagram of the structure of the automatic calibration device for an articulated manipulator based on a visual system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0010] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0011] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may change based on actual circumstances.

[0012] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0013] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0014] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.

[0015] See also Figure 1 , Figure 1 The flowchart of the automatic calibration method of the joint manipulator based on the visual system provided in the embodiment of the present application is as follows: Figure 1 As shown, the automatic calibration method of the articulated manipulator based on the vision system provided in the embodiment of the present application includes steps S100 to S600.

[0016] Step S100: performing multi-view image acquisition on the joint manipulator of the five-axis robot to obtain original image data; It is understandable that the execution subject of the present invention can be a joint manipulator automatic calibration device based on a visual system, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.

[0017] Specifically, a global visual tracking system is constructed by deploying M high-frame-rate, high-resolution industrial cameras around the workspace of a five-axis robot. These cameras, mounted on a fixed structure in a three-dimensional spatial coordinate system, form a surrounding pattern, capturing the robot's overall posture, position, and trajectory changes in real time from multiple angles. Furthermore, to achieve precise recognition and three-dimensional reconstruction of local structures, particularly joints, N high-resolution industrial cameras are rigidly fixed to the end-effector of the articulated robot. These cameras form a local visual measurement system and are configured using a binocular stereo vision arrangement, enabling sub-pixel depth recovery with minimal parallax. After the cameras are physically deployed, the M cameras in the global visual system are extrinsically calibrated using a calibration plate. By capturing images of the calibration plate at multiple known positions and postures, the precise spatial positional relationships and posture transformation matrices between the cameras are repeatedly calculated, establishing a unified global three-dimensional coordinate system. Based on this, a binocular calibration operation is performed on the N cameras in the local visual system. By calculating their intrinsic parameter matrices (including focal length, principal point position, and distortion coefficients) and extrinsic parameter matrices (including rotation and translation vectors), a stereo vision measurement model is constructed. All camera systems are synchronized using a unified time base and hardware-level synchronization triggers, ensuring temporal consistency of image acquisition frames within microsecond precision to prevent motion blur or data mismatch. The robot design incorporates high-contrast, light-resistant circular markers at each joint position. Made of low-reflectivity and a stable geometry, these markers serve as core reference features for visual recognition and 3D positioning in subsequent images. To ensure consistent capture of the robot's state at the same point in time by the global and local systems during image acquisition, the global and local vision systems are jointly calibrated in space and time based on the pre-established camera spatial position relationship and the binocular camera parameter matrix. This ensures that image data acquired from different angles and depths can be subsequently fused and stitched together. After completing the initial configuration, the calibration process is initiated, controlling the articulated robot to perform rhythmic, continuous motion along a pre-set trajectory path that covers representative poses in the five-axis motion space and ensures that each joint is exposed at least once from different viewing angles during acquisition. The global vision system continuously captures the robot's overall trajectory and the relative motion paths of its joints, while the local vision system focuses on capturing joint details, landmark features, and relative depth changes. Both systems are triggered synchronously during each sampling cycle, ultimately outputting a raw image dataset consisting of a global vision tracking image sequence and a local vision measurement image sequence.

[0018] Step S200: extracting occlusion regions from the original image data to obtain a set of joint occlusion regions; Specifically, a representative image frame is selected from the global visual tracking image sequence as a reference image. This reference image covers the joint structure and landmark areas of the articulated manipulator in an unobstructed state, ensuring that it serves as the benchmark frame for subsequent image comparisons and projection transformations. Each subsequent frame in the image sequence is a secondary image, which may contain localized occlusions due to factors such as joint posture changes, manipulator occlusion, and the presence of external interfering objects. The precise boundaries of the occluded regions are identified by using the structural feature relationships between the images. To achieve precise correspondence between the images, feature extraction and matching methods are used for comparison. A stable set of image feature points is extracted from the reference image and each secondary image frame using the SIFT or SURF algorithm, ensuring that these feature points are invariant to rotation and scale changes. A matching relationship is then established between the feature points using the RANSAC algorithm, eliminating incorrectly matched point pairs and improving the robustness of the correspondence relationship. Once the matching relationship is determined, the homography matrix between the images is calculated based on the corresponding point pairs. This matrix describes the perspective projection transformation relationship from the secondary image to the reference image, allowing the content of the secondary image to be uniformly mapped into the coordinate system of the reference image. After projection, an image difference algorithm is used to compare the pixel differences between the projected second image and the reference image. Regions whose absolute pixel differences exceed a threshold are identified as suspected occlusion regions, resulting in an initial outline of the occlusion region. Because occlusions have irregular boundaries and local discontinuities, a recursive partitioning operation is performed on the target image region based on this initial outline. This process employs a binary tree model to implement recursive region splitting: the target image region is identified based on pixel intensity variance and gradient statistics. When the variance of a subregion exceeds a set threshold, the region is considered structurally complex and further partitioned into two subregions. If the variance of a subregion is below the threshold and its area is smaller than the set threshold, the region is marked as a leaf node. All leaf nodes that fall within the initial occlusion region determined by the image difference are considered to correspond to actual occlusion regions. These leaf nodes are then combined into a preliminary set of occlusion regions through a backtracking path within the tree structure. A series of morphological processing operations are then performed to improve the integrity and coherence of the occlusion boundaries. Morphological processing involves two steps: opening and closing. The former is used to eliminate pseudo-occlusion points caused by noise interference, while the latter is used to fill holes caused by occlusion edges. A circular kernel is used as the structuring element to adapt to the geometric characteristics of the marker points. The result is a high-precision set of joint occlusion regions.

[0019] In this embodiment, a binary tree structure with a spatial hierarchical relationship is constructed based on the initial range of the occlusion area, in which the root node is the complete image sub-block containing the initial occlusion area. Statistical operations are performed on the grayscale values of all pixels in the root node, including calculating the average value of their grayscale variance and gradient amplitude. These two indicators jointly reflect the texture complexity and edge feature strength of the region, and the image feature measurement parameters of the root node are generated accordingly. The pixel variance of the root node is compared with the set first target value. If the variance is greater than the threshold, it indicates that there is obvious structural inhomogeneity or local occlusion feature within the region, and further refinement is performed. The node is spatially divided according to the long axis direction of the region, and the principal component analysis method is used to estimate the main direction of the region as the reference direction for division. The root node is divided into two first-level nodes, and the two newly generated first nodes are added to the queue to be processed for subsequent analysis and judgment. During the iterative process, each first node is taken out from the queue to be processed in turn, and the variance and gradient distribution characteristics of its internal pixels are recalculated. If the variance of the first node is still greater than the set second target value, it means that the region still has a complex structure or there are occlusion points, then the deeper level division operation will continue. At this time, the division direction is no longer fixed to the long axis direction, but is based on the maximum gradient amplitude direction within the region to ensure that the image structure breaks can be more effectively segmented. The two second nodes generated after the division are added to the queue again for further analysis. For each generated second node, its area and pixel variance are calculated at the same time. If the area of the node is less than the square pixel threshold corresponding to the preset third target value, and its variance value is also lower than the fourth target value, it is inferred that the region is basically stable and no longer needs to be split, and the node is marked as a leaf node. On the contrary, if its area is still large or the variance is still high, it indicates that the region still contains multiple different substructures or occlusion areas that have not been completely separated. The system will re-queue it to continue iterative processing. As the queue iteration continues, the entire binary tree structure is gradually established. In the resulting binary tree, each leaf node represents a small region with relatively consistent pixel features in the image structure, exhibiting both spatial locality and grayscale consistency, providing an ideal analysis unit for subsequent occlusion determination. To accurately identify whether these leaf nodes are actually occluded areas, the average grayscale value in each leaf node is compared with the grayscale value at the corresponding spatial location in the reference image, and the difference between the two is calculated. If this difference exceeds the set fifth target value, it is considered that there is a significant change in light intensity or structure between the sub-region and the reference image, and the leaf node is then determined to be an occluded area.

[0020] Step S300: Based on the joint occlusion region set, perform unordered image splicing and marker feature extraction on the original image data to obtain joint feature image data; Specifically, based on the identified joint occlusion regions, a subset of images covering multi-angle information of the occluded joints is selected from the original image data. This selection process is ranked and evaluated based on the effective structural information contained in the images. Keypoints are extracted from each frame of the original multi-view image sequence using local feature extraction algorithms such as SIFT and ORB. The number of feature points per unit area in each image is calculated to obtain the image's feature point density index. Furthermore, the camera pose information is combined to estimate the viewing direction of each image relative to the joint reference point through angle calculation or extrinsic parameter solution to evaluate its visual coverage of the occluded region. Feature point density and visual coverage are combined as comprehensive indicators for evaluation, selecting the images that best cover the key structural information of the occluded region to form a candidate image set. For each image in the candidate image set, high-precision feature point extraction is performed, and the degree of feature matching between the images is calculated. The matching process uses descriptor similarity calculation and a mismatching point elimination mechanism to determine the number of stable feature point pairs between each pair of images. A similarity matrix is then constructed, in which each element represents the degree of structural similarity between the two images. Based on this, a minimum spanning tree algorithm or graph traversal strategy is used to extract the optimal path structure from the similarity matrix and determine the logical order for image stitching. This order follows the principle of the path with the least cumulative error in the feature space, minimizing the risk of global deformation during subsequent stitching. During the image stitching stage, a geometric transformation matrix is calculated for each pair of adjacent images according to the stitching order, initially establishing the mapping relationship between the images. The geometric transformation matrix solution is optimized using an enhanced RANSAC method to eliminate the impact of extreme matching errors on the accuracy of the transformation parameter solution. To improve the stitching effect, an elastic transformation model is introduced, dividing each image into multiple local grid cells (e.g., 16×16 pixel blocks). Local transformation parameters are calculated for each grid cell, enabling local adaptive adjustment in the stitching process, accurately adapting to errors caused by non-rigid deformation and viewpoint distortion, and ultimately outputting a complete image of the repaired joint. Landmark feature extraction is performed on this complete image, and a strategy combining adaptive threshold grayscale segmentation and a modified Hough transform is used to identify circular or ring-shaped landmarks in the image. After feature extraction is complete, precise stereo matching is performed between multi-view images using the previously calibrated multi-view camera internal and external parameters. The matching process introduces epipolar constraints and normalized cross-correlation coefficients to determine pixel-level corresponding points, and the 3D spatial coordinates of each marker point are calculated based on triangulation. If the candidate image set contains multiple image pairs with intersecting fields of view, the system uses a bundle adjustment algorithm to globally optimize all matching results to minimize reprojection errors and ensure that the 3D reconstruction results of each marker point have uniform and high-precision spatial consistency. This results in joint feature image data.

[0021] In this embodiment, adaptive threshold segmentation is performed on the complete image to accurately isolate circular or ring-shaped marker regions with high-contrast structures. Local image statistics are then used to achieve robust adaptation under varying lighting conditions. The complete image is divided into several sub-blocks, and the grayscale mean and standard deviation are calculated within each sub-block. A local threshold map is then generated using a dynamic threshold function to construct a binary segmented image that clearly highlights the marker features. Morphological filtering is performed on the segmented image to eliminate pseudo-target regions that are not markers. Candidate regions are then selected based on a priori parameters such as contour features, shape compactness, and area size. For each candidate marker region, the pixel distribution within it is analyzed using the centroid method or ellipse fitting method to obtain sub-pixel geometric center coordinates, which are then used to construct high-precision two-dimensional marker position information. Stereo matching is performed using epipolar constraints and a normalized cross-correlation algorithm for binocular images acquired by a local vision measurement system. An epipolar geometry model is calculated based on the intrinsic and extrinsic parameter matrices obtained from the two targets at regular intervals. Each image pair is then mapped to a unified epipolar geometry constraint using epipolar correction. The coordinates of the extracted marker points are selected from the primary view and matched with the corresponding epipolar lines in the secondary view using a fixed search window. The confidence of the matched pairs is determined by maximizing the normalized cross-correlation coefficient. Successfully matched point pairs are then associated with each other in the primary and secondary images. Based on the known extrinsic and extrinsic parameter matrices of the binocular cameras and the pixel coordinates of the marker points, triangulation is used to calculate the 3D spatial positions of the marker points in the local camera coordinate system. The 3D point positions are determined by constructing projection rays from each camera center to the image point and calculating the midpoint of the shortest distance between these rays, resulting in a high-precision 3D point set in the local vision system. Using these local 3D points as geometric anchors, bundle adjustment optimization is performed on the multi-view images in the global visual tracking system. This optimization process aims to minimize the multi-camera reprojection error. The 3D coordinates of all marker points are projected onto the image plane of each camera and compared with the corresponding 2D detection points in the actual image to construct an overall error function. A nonlinear least-squares optimization algorithm (such as the Levenberg-Marquardt algorithm) is then used to jointly adjust the extrinsic parameters of all cameras, the 3D coordinates of the marker points, and any scale factors. The entire bundle adjustment process incorporates a sparse matrix solution strategy to improve computational efficiency, and a robust loss function enhances outlier resistance, ensuring spatial consistency between data from different viewpoints. The acquired joint feature image data includes the precise 3D coordinates of each marker point.

[0022] Step S400: constructing an adaptive calibration equation based on the joint feature image data, and solving the target transformation relationship between the global visual coordinate system and the local visual coordinate system; Specifically, based on joint feature image data, three key coordinate systems involved in the system are defined. The global visual coordinate system is constructed by industrial cameras positioned around the workspace. Its origin is the reference point of the calibration plate, and its three-dimensional coordinate frame is established based on the spatial structure during the calibration process. The local visual coordinate system is constructed by the binocular vision system installed at the end of the articulated manipulator. Its reference point is set at the optical center of the main camera, and its coordinate base is established based on the extrinsic parameter matrix obtained through binocular calibration. The robot base coordinate system is defined within the control system, with the center of the robot base as its origin, and the spatial relationships between the axes are determined according to the mechanical structure. The spatial mapping relationship between these three coordinate systems constitutes the core issue of the subsequent calibration process. To establish a data foundation for calibration solutions, the five-axis robot is controlled to execute a set of preset trajectories covering its typical workspace. At these key poses, the system simultaneously collects data from three sources: the 3D coordinates of markers at each pose are obtained by the global visual tracking system; the binocular images captured by the local visual measurement system are used to perform stereo reconstruction, deriving the spatial positions of the markers in the local coordinate system; and the joint angle values and end-user pose output by the robot control system. These data are organized into coordinate sample sets in a one-to-one correspondence. During the calibration model construction process, the spatial transformation patterns between the global vision system and the robot control system, as well as between the local vision system and the robot control system, are identified based on these sample sets, forming two types of coordinate transformation models. Based on this, the calibration process analyzes scale consistency issues, particularly those caused by factors such as focal length error, projection distortion, and image size scaling across different camera systems. A scaling factor is introduced into the calibration equations to uniformly correct the spatial scale, enabling the transformation model to reflect rotation and translation relationships and accommodate differences in measurement scale between systems. To ensure high accuracy and stability of the calibration results, the constructed complete calibration equations are solved using a nonlinear optimization approach. The optimization process aims to minimize the spatial mapping error of each coordinate system under multiple poses. Rotation parameters, displacement parameters, and scaling factors are repeatedly adjusted to ensure that the coordinate reconstruction results from each viewpoint are as close as possible to the actual observations. This process not only considers visual reconstruction errors but also considers multi-view consistency and spatial structure continuity, ensuring that the resulting coordinate system transformation has good physical consistency and engineering applicability. Obtain the target transformation relationship between the global visual coordinate system and the local visual coordinate system.

[0023] Step S500 : performing multi-joint collaborative calibration and joint chain constraint optimization of the five-axis robot according to the target transformation relationship to generate joint error correction parameters.

[0024] Specifically, a calibrated motion trajectory is constructed that covers the entire workspace of the five-axis robot. This trajectory includes multiple key pose points that are highly representative and evenly distributed across space, ensuring that each joint undergoes sufficient changes during the robot's motion, thereby stimulating potential system errors. Based on this calibrated motion trajectory, the robot is controlled to gradually execute the motion task. At each pose point, a global visual tracking system collects the actual joint position and pose data in real time, obtaining observations in real physical space. These observations are not constrained by the theoretical model within the robot control system and represent the true response of the robot in the actual environment. Combining the target transformation relationship between the global and local visual coordinate systems solved in the previous stage, the three-dimensional data obtained from visual observations are uniformly transformed into the robot control coordinate system. Based on the transformation results, the current spatial position of each joint is compared with the theoretical position provided by the robot control system. By calculating position offsets and pose differences, the parameter errors exhibited by each joint during actual motion are derived. These errors include axis offsets, rotation angle deviations, zero point offsets, and other errors, forming a preliminary set of error parameters for each joint. Initial error parameters are introduced into a global joint chain constraint optimization model. Based on the geometric topological relationships within the manipulator structure and combined with continuity constraints from kinematics, this model establishes requirements for relative distances, rotational directions, and spatial coherence between joints. This ensures that the correction parameters for each joint are not only locally effective but also consistent with those of its upstream and downstream joints throughout the entire joint chain. This optimization model mathematically coordinates all joint errors across the entire joint chain, seeking a set of optimal correction values that maintains spatial structural integrity and kinematic coherence for the virtual joint chain formed by all joints during robot motion, while also minimizing the global error between observed and theoretical values in the calibration trajectory. The resulting correction parameter set, derived from this optimization model, constitutes joint error compensation information that meets the requirements of the mechanical structure logic and control system. Based on this information, the actual control parameters of each joint are corrected, and the error sources are classified and identified, decomposing them into sub-items such as zero drift, transmission backlash compensation, and shaft eccentricity correction. These are recorded as joint error correction parameters, forming an error compensation table that is directly applied to the robot controller.

[0025] In this embodiment, a kinematic model is constructed for a five-axis robot. This model includes theoretical definitions of the rotational or translational degrees of freedom for each joint, as well as basic parameters such as the connection relationship between joints, coordinate transformation structure, and connecting rod length. Preliminary joint error parameters are combined with the nominal parameters in the theoretical model to form a joint chain parameter set with error perturbations, resulting in a kinematic expression that includes error terms such as offset, rotation, and zero offset. Based on this joint chain model with error parameters, a forward kinematic equation is established, which describes the spatial position and posture that the robot's end or intermediate nodes should reach when given a set of joint angle inputs. This equation not only relies on the theoretical structure but also includes the impact of the error parameters on each motion step. Therefore, it is defined as a constraint to verify whether the spatial output of any joint combination conforms to the continuity and physical feasibility of the geometric chain structure. Using actual joint spatial position data acquired by a global visual tracking system and a local visual measurement system in multiple postures, the constructed forward kinematic model is subjected to point-by-point error analysis, with particular attention paid to whether the relative positions and postures of adjacent joints are consistent during different joint combination motions. A position closed-loop error evaluation mechanism is introduced. The position output by the theoretical model at each pose sampling point is compared with the actual position measured by the vision system. This results in a set of error-driven constraint equations. These equations reflect local incoherence, spatial misalignment, or structural offsets that occur during the motion of the actual joint chain, which are important sources of closed-loop constraints for the entire joint chain. After the closed-loop constraints are constructed, the joint chain constraints and the position closed-loop error equations are unified into a weighted least-squares optimization model. In this model, different error terms are assigned weight coefficients to emphasize the importance of structural constraints while accommodating the discrete error fluctuations of the measurement system, thus achieving a balance between spatial accuracy and mathematical stability in the optimization results. Geometric and kinematic constraints are introduced in the optimization model. The former ensures consistency in the spatial relationships between joints, such as constant connection lengths and axis distortion. The latter ensure that the range of motion and angular velocity of each joint remain within the physically feasible domain, avoiding the generation of parameter combinations that do not conform to the actual control logic during the calibration process. Based on the resulting constrained optimization problem, a hybrid nonlinear optimization strategy is employed for solution. This strategy combines global search with local convergence algorithms to effectively find globally optimal or near-globally optimal parameter solutions in a non-convex solution space, while automatically handling the activation and adjustment of constraint boundaries. The optimization process, through iterative solution of each set of error terms and dynamic adjustment of constraint residuals, gradually converges to a parameter solution set that minimizes the overall error and satisfies geometric consistency. The resulting corrected parameter set characterizes the actual error characteristics of each joint, while achieving a unified balance of structural closure, kinematic coordination, and physical feasibility across the entire joint chain.

[0026] In an embodiment of the present invention, a multi-view 3D visual measurement network combining a global visual tracking system and a local visual measurement system is employed to achieve omnidirectional monitoring of a five-axis robot articulated manipulator, effectively resolving the problem of inability to obtain complete information from a single viewpoint. Occluded regions are precisely extracted using a binary tree model algorithm, and complete information about occluded joints is restored using unordered image stitching technology. This overcomes the drawback of traditional methods with insufficient calibration accuracy in complex environments, freeing the calibration process from the limitations of environmental occlusion. A scale factor correction mechanism and a multi-pose sampling strategy are introduced to address the fusion problem between the global and local visual coordinate systems, establish an accurate coordinate mapping relationship, and achieve precise data integration between systems of different scales. A joint chain constraint optimization model is employed to globally optimize calibration parameters in conjunction with geometric and kinematic constraints, ensuring that the calibration results meet the physical characteristics of the joint mechanism and enhancing the physical significance and practicality of the calibration. A recursive least squares method combined with a nonlinear optimization algorithm is employed to achieve high-precision calculation of joint error correction parameters, significantly improving the positioning and motion accuracy of the five-axis robot articulated manipulator.

[0027] In a specific embodiment, the process of executing step S100 may specifically include the following steps: M industrial cameras are fixedly installed around the workspace of the five-axis robot to form a global visual tracking system, and N high-resolution cameras are fixedly installed on the end effector of the joint manipulator to form a local visual measurement system; The industrial cameras in the global visual tracking system are calibrated with external parameters through the calibration plate to determine the spatial position relationship between the industrial cameras. The industrial cameras in the local visual measurement system are configured with binocular stereo vision to establish the camera parameter matrix. A circular marker is set at each joint of the articulated manipulator, and the global visual tracking system and the local visual measurement system are synchronized based on the spatial position relationship and the camera parameter matrix to ensure the synchronization of image acquisition; The joint manipulator is controlled to move along a preset path, and image acquisition is synchronized through the global visual tracking system and the local visual measurement system to obtain original image data including a global visual tracking image sequence and a local visual measurement image sequence.

[0028] Specifically, an array of industrial-grade cameras with overlapping and complementary viewing angles is arranged around the workspace of a five-axis robot. This vision system, consisting of M industrial cameras, is fixedly mounted to form a global visual tracking system. The installation position of each industrial camera is determined through structural simulation and field of view analysis to ensure that the cameras form overlapping viewing angles and minimize blind spots in their spatial distribution. This ensures that the robot's joints and end-effector remain within the effective field of view of one or more cameras during any pose. Furthermore, to enhance the system's ability to capture local spatial features, N high-resolution cameras are mounted on the end-effector of the articulated manipulator. These cameras are rigidly connected to the manipulator arm, capturing close-range image information of the manipulator's joints as the end-effector moves. These local cameras are arranged in a binocular vision configuration, maintaining a fixed angle between their viewing axes and a fixed baseline length, forming a local visual measurement system capable of stereoscopic imaging and depth perception of joint features. External calibration is performed on the M industrial cameras in the global vision system to determine their relative position and pose in three-dimensional space. Using a standard industrial calibration target, each camera collects image data from the target by moving it point by point to different known spatial locations. This data is then used to calculate the position and orientation of each camera relative to a common reference coordinate system (e.g., the lower left corner of the target). Through repeated capture and matching operations, a multi-camera extrinsic parameter matrix is constructed. For the local vision measurement system, binocular calibration is performed on N cameras to establish a binocular camera parameter matrix containing intrinsic and extrinsic parameters. Intrinsic parameters primarily include focal length, principal point coordinates, and image distortion coefficients, while extrinsic parameters define the relative rotation and translation between the two cameras. This process observes a calibration pattern with known geometry and performs feature point matching and reprojection error optimization on the images captured by the two cameras to form a binocular imaging model. To ensure synchronization of image acquisition in the temporal dimension and consistency in the spatial coordinate system, a unified synchronization mechanism is established between the global and local vision systems. This mechanism includes both physical synchronization trigger signal deployment and system-level timestamp alignment and unified data flow management. At the hardware level, all cameras are connected to high-precision synchronization triggers, and frame-level image acquisition is completed at the same time as the system sends a trigger signal, so that the global image and local image acquired at any time have a strict time correspondence. At the software level, all image data are attached with metadata with a system timestamp and camera number to facilitate subsequent inter-frame alignment and data association during image fusion, feature matching, and three-dimensional calculations. At the same time, to improve the consistency of spatial data, the spatial position relationship and parameter matrix established in the aforementioned calibration process are used to embed the spatial mapping function between the global visual coordinate system and the local visual coordinate system into the image processing process, so that image acquisition and geometric reconstruction always operate under the same spatial reference, avoiding drift or mismatch of spatial information between different systems.After completing the system initialization and calibration, it enters the image acquisition phase. The five-axis robot is controlled to move precisely according to a pre-set trajectory. The trajectory consists of a series of points covering different postures, angles, and spatial positions. Its design principle is to ensure that the five joints can fully move in multiple degrees of freedom and that the captured images have structural diversity and information redundancy. During the movement of the robot along the trajectory, the control system triggers the visual system to synchronously capture images in a time-driven manner, ensuring that at each critical posture moment, the global visual system obtains multi-angle external observation images, while the local visual system focuses on capturing joint detail images under the current perspective. All image data are uniformly stored in two types of image sequences according to time sequence and identification, namely global visual tracking image sequence and local visual measurement image sequence.

[0029] In a specific embodiment, the process of executing step S200 may specifically include the following steps: A first image containing a complete view of the manipulator joints is selected from the global visual tracking image sequence of the original image data as a reference image, and each subsequent second image frame obtained in the global visual tracking image sequence is compared with the reference image; Extract features from the reference image and the second image to obtain a feature point set for each frame of the second image, and match the feature point sets to establish a correspondence between the images; The homography matrix is calculated based on the correspondence between the images and the local visual measurement image sequence, and the second image is projected into the reference image coordinate system to perform image difference calculation to determine the initial range of the occluded area; Recursively split the target area into sub-areas according to the initial range of the occluded area, and determine the leaf nodes marked as occluded areas; All leaf nodes marked as occlusion regions are combined to form a preliminary occlusion region set, and the preliminary occlusion region set is morphologically processed to obtain a joint occlusion region set.

[0030] Specifically, a frame of image is selected from the global visual tracking image sequence as a reference image. The image meets a key condition, that is, in the frame, all joints of the manipulator are completely unobstructed, and the image clarity and contrast meet the recognition requirements, providing an accurate baseline view for subsequent occlusion detection. The reference image is selected based on an image quality scoring mechanism, a marker point integrity detection, or a preset initial pose selection strategy for automatic screening. After the reference image is selected, each frame of the subsequent global visual image sequence is processed in turn. These images are second images, which have part of the joint area obscured due to reasons such as manipulator movement, obstruction intervention, or self-occlusion of the structure. Feature extraction operations are performed on the reference image and each frame of the second image. A feature extraction algorithm with scale, rotation, and affine invariance, such as SIFT or SURF algorithm, is used to ensure that stable image features can be extracted under different angles and lighting conditions. After feature extraction, high-precision matching is performed on the feature point sets in the reference image and the second image. By calculating the descriptor similarity between the feature points and combining false match elimination mechanisms such as nearest neighbor ratio filtering or the RANSAC algorithm, a stable correspondence relationship between the reference and second images is established. These matched point pairs constitute the spatial geometric association information between the images. Based on the correspondence relationship between the images and the sequence of local visual measurement images, a homography matrix is calculated. This matrix describes the planar projective transformation relationship between the images. Pixels in the second image are mapped to the coordinate system of the reference image according to this transformation relationship, spatially aligning the content of the two images. Image projection is used to unify the coordinate space of images acquired at different times, allowing occluded areas to be effectively identified within the same reference frame. Especially when the system also has local visual measurement images, the 3D structure data reconstructed in the local vision system is used to enhance the accuracy of the homography calculation, making the projection alignment more rigorous and precise. After image alignment, the projected second image and the reference image are subjected to image difference processing. The difference in the grayscale value of each pixel is calculated. Based on the distribution of the difference values, regions with significant changes are extracted. These regions are considered suspected occlusion areas. The difference calculation is combined with image smoothing preprocessing, local mean removal, and difference thresholding to reduce background noise interference and improve the clarity of occlusion boundaries. All significant difference areas are merged to form the initial range of the occlusion area. Taking the preliminary occlusion area as the root area, a binary tree structure is constructed in the image space to recursively split it. The entire occlusion area is divided into multiple sub-areas, and image texture analysis is performed on each sub-area, including the standard deviation of grayscale values, gradient amplitude, edge strength, etc. When the texture complexity (such as grayscale variance) within a sub-area exceeds a preset threshold, it indicates that the area contains a structural boundary or multiple occlusion types, and the area is further divided into two smaller sub-blocks; when the texture variance of a sub-area is less than the threshold and the area is small enough, the system marks it as a leaf node and no longer splits it.These regions that ultimately remain undivided become leaf nodes in the binary tree structure. Each leaf node represents a relatively uniform subsegment of the image in image space. The occlusion status of each leaf node is determined by comparing the average grayscale value of its internal pixels with the grayscale value of the corresponding reference image region. When the difference exceeds a set significance threshold, the leaf node is considered to be an actual occluded region. The accuracy of this process depends on the accuracy of the initial image registration and the rationality of the difference threshold. By adjusting the comparison window size, difference threshold range, and grayscale normalization method, it effectively adapts to different lighting conditions and occlusion types. All leaf nodes determined to be occluded are collected and integrated to form a preliminary occlusion region set. This preliminary occlusion region set undergoes image morphological processing to improve the continuity and regional integrity of the occlusion boundaries. Morphological processing includes opening and closing operations. Opening eliminates isolated small noise blocks, while closing fills discontinuous cracks in the occlusion boundary, making the occlusion region contour smoother and more closed, making it suitable for subsequent image inpainting. Through these steps, a set of joint occlusion regions is ultimately obtained.

[0031] In a specific embodiment, the process of executing the step of recursively splitting the occluded area into sub-areas according to the initial range of the occluded area and marking the corresponding nodes as leaf nodes may specifically include the following steps: The initial range of the occluded area is used as the root node of the binary tree, and the variance of the grayscale values of all pixels in the root node and the average value of the gradient amplitude are calculated to obtain the feature measurement of the root node; Compare the pixel variance within the root node with the first target value. When the pixel variance is greater than the first target value, split the root node into two first nodes along the long axis of the region, and add the first nodes to the queue for processing. Sequentially take out the first nodes from the queue to be processed, calculate the variance and gradient information of the pixels in each first node, and when the variance of the pixels in the first node is greater than the second target value, continue to split the first node into two second nodes along the direction of maximum gradient; Calculate the area of the second node. When the area of the second node is less than the third target value square pixel and the pixel variance is less than the fourth target value, mark the second node as a leaf node. Otherwise, add the second node back to the queue for processing. The splitting process is iteratively performed based on the queue to be processed to obtain a binary tree structure, in which each leaf node represents a sub-region with similar pixel characteristics; All leaf nodes are marked and judged. When the difference between the average grayscale value in the leaf node and the grayscale value of the corresponding area of the reference image exceeds the set fifth target value, the leaf node is marked as an occlusion area.

[0032] Specifically, the initial range of the occluded area is used as the root node of the tree structure, and image statistical feature analysis is performed inside the root node, including calculating the variance of the grayscale values of all pixels contained therein and the average value of the corresponding gradient amplitude. These two metrics reflect the degree of brightness change and the level of edge strength within the region, respectively. The above two values are jointly defined as the image feature metric of the root node to determine whether the region needs to be further divided. When the grayscale variance of the root node exceeds the first target value, it means that there are multiple image structures in the region, or there is an uneven image distribution caused by occluders, sudden changes in illumination, texture interference, etc., so the region is chosen to be initially split. In order to ensure the directionality and structural consistency of the sub-regions after splitting, the main direction of the region, that is, the long axis direction, is calculated and the region is divided into two along this direction to generate two first-level sub-regions. These two regions are the first nodes and are immediately added to the processing queue, waiting to be taken out in sequence for in-depth analysis. After entering the iterative phase, the first nodes are removed from the processing queue one by one in a first-in, first-out order. Pixel characteristic statistics are recalculated for each node, including recalculating the subregion's grayscale variance and image gradient information, with particular attention paid to the continuity and intensity distribution of edge structures. If the grayscale variance of a first node is still greater than the set second target value, the system deems that the region has not yet reached a stable state of texture homogeneity and proceeds to further split it. The splitting direction is no longer fixed to the long axis; instead, the direction with the largest gradient amplitude within the region is dynamically selected. This ensures that locations most sensitive to image structural changes are prioritized, effectively revealing structural faults caused by occlusion boundaries or complex background changes. The two second nodes generated by this splitting process are added back to the processing queue for the next round of analysis. For each second node, in addition to further analyzing its grayscale statistics, its area in image space is calculated. This controls the granularity of the image splitting process and prevents the region from being infinitely refined to the pixel level, which would increase subsequent processing costs or accumulate errors. When the area of the second node is smaller than the minimum effective processing area corresponding to the third target value, and its grayscale variance is lower than the fourth target value, it indicates that the area is not only small in size but also stable in structure. The system marks it as a leaf node and identifies it as the smallest analyzable unit in the image, and no longer participates in subsequent recursive splitting. On the contrary, if a second node still does not meet the stability requirements in terms of area or texture fluctuation, the system re-adds it to the queue to be processed and continues to participate in the next round of subdivision processing. Through this recursive iteration method, the construction of the entire binary tree structure is completed. Each leaf node represents a small area in the image occlusion area that is relatively homogeneous in structure and of reasonable size. The occlusion status of all image sub-areas that have been marked as leaf nodes is determined, and the grayscale values of the pixels in each leaf node are averaged and compared with the grayscale values at the same spatial position in the reference image.By calculating the average grayscale difference and comparing it with the fifth target value set by the system, when the difference exceeds the set threshold, it is considered that the area has been occluded in the current frame image. The basis of this judgment logic is that the joint area in the reference image is in an unoccluded state, thus providing a texture benchmark with complete structure. When there is a significant grayscale mutation or texture loss in the occluded area, the average grayscale value will shift significantly, thus serving as an effective feature for occlusion discrimination. In order to improve the recognition accuracy, a neighborhood consistency analysis mechanism is introduced to perform consistency fusion on the occlusion judgment results of multiple adjacent leaf nodes to avoid misidentification of grayscale fluctuations caused by noise or local illumination changes. All leaf nodes marked as occluded are integrated to form a high-precision occlusion area set. This set appears as an image mask with smooth boundaries, complete areas, and accurate positioning in the image structure, and can be directly used for subsequent image restoration, structure reconstruction, or 3D point cloud completion operations.

[0033] In a specific embodiment, the process of executing step S300 may specifically include the following steps: Screening an image subset containing different perspective information of the occluded joint from the multi-view image sequence of the original image data, and calculating the image feature point density and perspective coverage of the image subset to obtain a candidate image set; Extract feature points and calculate feature matching for each image in the candidate image set to construct an image similarity matrix; Determine the image splicing order based on the image similarity matrix, and calculate the geometric transformation matrix for each pair of adjacent images in the image splicing order to obtain the image mapping relationship; Based on the image mapping relationship, local transformation parameters are calculated and fine-grained splicing is performed through the elastic transformation model to obtain the complete image of the repaired joint; The repaired complete joint image is subjected to landmark feature extraction and multi-view 3D reconstruction calculation to obtain joint feature image data.

[0034] Specifically, a subset of images containing multi-view information is extracted from the original image data as the basis for stitching and restoration. This process takes as input image sequences acquired by the global and local vision systems. The system selects frames from a large number of images that capture the structural features of the occluded joint region from different angles. Therefore, an image screening and evaluation mechanism is introduced. During this stage, through position mapping analysis of the occluded region, a set of images with temporal or perspective visibility of the occluded region is identified. Two core metrics are calculated for these images: image feature point density, which is the number of feature points extracted per unit area. This metric reflects the richness of image texture information and directly affects the registration accuracy during subsequent stitching; and view coverage, which is whether the different boundary contours of the occluded region can be captured from the image's perspective. The coverage of the occluded region in each image is determined by analyzing the camera pose and gaze direction. After jointly evaluating these two metrics, image frames that both possess rich structural details and fully complement the occluded region from a perspective are selected as candidate image sets. After obtaining a set of candidate images, a detailed feature point extraction process is performed on each image. Key points and their descriptors are extracted using a robust feature description algorithm such as ORB or AKAZE. Pairwise feature point matching is then performed between all images, and a nearest neighbor ratio screening mechanism is used to eliminate low-quality matching pairs. This results in a matching strength matrix, or image similarity matrix, that reflects the structural similarity between the images. Each element of this matrix represents the strength of the structural match between the two images. Larger values indicate greater complementarity and stitchability in areas of structural overlap. By analyzing the distribution of values in the similarity matrix, an image connectivity graph is constructed. The final image stitching order is determined using a minimum spanning tree algorithm or a maximum weight path approach, ensuring a logically continuous spatial chain of images, avoiding cumulative deformation or image drift caused by jump connections. A geometric transformation matrix is calculated for each pair of adjacent images in the stitching sequence. For each pair of adjacent images, a geometric transformation matrix is calculated based on their matching point pairs. The transformation model is selected from affine, perspective, or homography transformations based on reconstruction accuracy requirements. The optimization is performed under the constraints of the RANSAC algorithm to ensure that the transformation matrix is immune to noise and maintains geometric stability. An elastic transformation model is introduced as a fine-grained compensation mechanism. Each image is divided into multiple grid blocks, each with independent local transformation parameters. These parameters are guided by the global transformation and fine-tuned with local residuals. This refinement achieves natural fusion of structural alignment and texture transitions at the stitching boundary. During the elastic transformation stitching process, a multi-scale image fusion algorithm is employed to divide each image into high-frequency details and low-frequency brightness structure in the frequency domain, taking into account stitching and brightness balance at image edges. Fusion weights are assigned based on the distance from the pixel to the stitching seam. This approach strengthens the consistency of feature contours in high-frequency regions and achieves smooth brightness transitions in low-frequency regions, ensuring that the stitched image is structurally continuous and visually natural.After stitching is complete, a high-quality joint image is output, which has been processed through multi-view fusion and fully restored in occluded areas. Joint landmarks in the inpainted image are accurately extracted, and adaptive threshold grayscale segmentation is used to highlight their outlines. Morphological operations are then used to remove noise. A modified Hough transform is then used to identify circular or ring-shaped landmarks in the image. Considering that landmark shapes may deform under different viewpoints, ellipse fitting and sub-pixel centroid localization methods are introduced to extract the center coordinates of each landmark and accurately record them as feature vectors in the image coordinate system. To convert the landmarks in the 2D image into 3D coordinates, multi-view 3D reconstruction is performed. Leveraging the previously calibrated intrinsic and extrinsic parameter matrices of the multi-camera system, epipolar constraints are used to determine the correspondence between landmarks in the image from different viewpoints. A normalized cross-correlation matching algorithm is then used to improve matching accuracy. After matching, matching point pairs are spatially localized using triangulation. Multi-view redundant data is then used to perform bundle adjustment optimization to minimize the reprojection error of each camera, resulting in the precise position of the landmarks in the global 3D coordinate system. The three-dimensional reconstruction results have millimeter-level spatial accuracy and avoid error propagation and local distortion through multi-view structural consistency constraints, so that the final joint feature image data has the integrity and reliability that can be used for calibration modeling, parameter optimization and coordinate transformation.

[0035] In a specific embodiment, the execution step performs marker feature extraction and multi-view 3D reconstruction calculation on the repaired complete joint image to obtain joint feature image data, which may specifically include the following steps: Adaptive threshold segmentation is performed on the repaired complete joint image to obtain a segmented image that highlights the features of the marker points; Generate candidate regions for marker points based on the segmented image that highlights the features of the marker points, and calculate the sub-pixel center coordinates of the marker points based on the candidate regions to obtain high-precision marker point position information; For the binocular images acquired by the local vision measurement system, epipolar constraint and normalized cross-correlation algorithm are used for stereo matching to obtain the corresponding relationship of the marker points. Based on the correspondence between the marker points and the camera parameter matrix, the local three-dimensional coordinates of the marker points are calculated using triangulation to obtain the three-dimensional point set in the local visual measurement system. According to the 3D point set in the local vision measurement system, the bundle adjustment method is applied to the multi-camera images in the global vision tracking system to perform multi-view 3D reconstruction and global optimization of camera observation data to obtain joint feature image data.

[0036] Specifically, adaptive threshold segmentation is performed on the restored joint image to accurately separate the landmark features representing the joint structure's positioning benchmarks from the complex background. Due to various interferences during image acquisition, such as uneven illumination, local shadows, and structural reflections, the entire image is divided into several local subregions. The grayscale mean and standard deviation are calculated for each subregion, and the threshold range is dynamically adjusted based on the local brightness environment. This effectively separates the landmarks from the background in both bright and dark areas. A typical segmentation strategy utilizes a local statistical threshold model to discriminate each pixel, generating a binary image that only contains strong edges and high-contrast image objects—in other words, a segmented image that highlights the landmark features. Potential landmark regions are then identified in the segmented image through contour extraction, connected domain analysis, and geometric morphology filtering. These candidate regions exhibit typical structural characteristics within the image, such as consistent size, closed edges, and concentrated grayscale. Therefore, they are filtered and sorted using metrics such as area, perimeter, and boundary tightness to retain high-confidence landmark candidates. On this basis, sub-pixel center coordinate localization is performed for each candidate region. The grayscale centroid method or ellipse fitting method is used to precisely model the grayscale distribution of pixels within the marker region, thereby determining its precise center position at sub-pixel resolution. This localization operation compensates for traditional integer pixel errors through continuous sub-pixel interpolation and gradient distribution analysis, ensuring that the localization error is within the single-pixel accuracy range. This results in a high-precision set of 2D image coordinates for 3D reconstruction. These coordinates constitute the spatial projection center of the marker in the image plane. Stereo matching is performed using epipolar constraints and normalized cross-correlation algorithms for binocular images acquired by a local vision measurement system. Based on the intrinsic and extrinsic parameters obtained from binocular camera calibration, the pixels in the image pair are projected onto their respective imaging planes, and the epipolar equations that they should geometrically satisfy are calculated. Through epipolar correction, the original image pair is reprojected into a set of row-aligned view structures, ensuring that corresponding pixels are on the same scan line, simplifying the 3D matching problem to a one-dimensional search. Based on this geometric constraint, for each high-confidence marker in the reference image, a sliding window search is performed along its epipolar line in the matching image, and the normalized cross-correlation coefficient is calculated at each position to measure the structural similarity between the image blocks. The position that reaches the maximum value in the matching score is considered to be the corresponding point, thereby obtaining the matching point pair of the marker point in the left and right image pairs, that is, the projection correspondence of the marker point in the binocular image. After establishing the binocular projection relationship of the marker point, the known camera intrinsic parameter matrix and extrinsic parameter relationship are called, and the triangulation method is used to perform spatial reconstruction on each pair of matching points to calculate its true three-dimensional position in the local visual coordinate system. This process constructs two projection rays pointing from the camera optical center to the image point and calculates their nearest intersection or the nearest distance midpoint as the three-dimensional space reconstruction result.By repeating this process for all marker points, a set of 3D points with true depth information and spatial distribution characteristics is formed, namely the 3D coordinate set of the marker points in the local vision system. Due to the limited field of view and short baseline of the local vision system itself, the overall stability of its reconstructed point cloud is limited, and slight scale drift or inconsistency may occur. Therefore, the local 3D point set is mapped and optimized to a unified framework within the global vision coordinate system. Image data collected at multiple times and from multiple perspectives in the global visual tracking system is selected to construct a multi-camera image set, and the 3D point set provided by the local vision system is used as the reference geometric anchor point. Based on this, the bundle adjustment method is introduced to perform global 3D reconstruction and joint optimization of the multi-view data. The bundle adjustment process essentially aims to minimize the reprojection error of each reconstructed point in the global camera system in each image. All camera poses, intrinsic parameters, and 3D point coordinates are jointly solved, and each parameter is adjusted through iterative updates to achieve the best approximation to the actual observation data. This process can effectively eliminate the spatial drift caused by the accumulation of local visual system errors, and use multi-redundant perspective information to improve the overall consistency, distribution stability and geometric continuity of the point cloud. The final output three-dimensional point set is the joint feature image data after global optimization.

[0037] In a specific embodiment, the process of executing step S400 may specifically include the following steps: Based on the joint feature image data, the global visual coordinate system, local visual coordinate system and robot base coordinate system are defined, and the joint manipulator of the five-axis robot is controlled to move to different positions. At each position, the three-dimensional coordinates of the joint markers in the global visual tracking system and the local visual measurement system, as well as the joint angles reported by the robot control system, are simultaneously recorded to construct a coordinate correspondence sample set; Based on the coordinate corresponding sample set, a first transformation matrix from the global to the robot coordinate system and a second transformation matrix from the local to the robot coordinate system are constructed to form an adaptive calibration equation; A scale factor is introduced into the adaptive calibration equation to obtain a complete calibration equation including the scale factor. The complete calibration equation including the scale factor is solved by nonlinear optimization to obtain the target transformation relationship between the global visual coordinate system and the local visual coordinate system.

[0038] Specifically, the logical and physical correspondences between the global visual coordinate system, the local visual coordinate system, and the robot base coordinate system are clearly defined and established in geometric space. This process relies on the three-dimensional spatial position information of the marker points provided by the joint feature image data as a unified reference basis. The global visual coordinate system is based on the lower left corner of the calibration plate or the spatial origin set by the system. Its X, Y, and Z axes extend along the horizontal, vertical, and perpendicular directions of the calibration plane, respectively. This coordinate system is established during image acquisition and 3D reconstruction by a global vision system composed of multiple fixed industrial cameras. It has the characteristics of wide spatial coverage and stable positioning. The local visual coordinate system relies on a high-resolution binocular camera installed at the end of the joint. The origin is set at its optical center, with the X-axis as the baseline direction and the Z-axis perpendicular to the imaging plane as the direction. Since it is installed at the end of the robot arm, it changes dynamically with the movement of the robot and has high-precision local perception capabilities. The robot base coordinate system is defined by the robot control system. Its origin is located at the fixed base of the robot body. The axes are strictly aligned with the degree of freedom distribution in the robot kinematic model. It serves as the basic reference for all servo command execution and posture solution. To establish the transformation relationship between these coordinate systems, the robot's articulated manipulator is controlled by the system to move to a series of known pose points. During this process, the five-axis robot is controlled to perform multiple angle changes and end-effector spatial translations according to a pre-set multi-pose calibration strategy, ensuring that all joints have a sufficient range of motion in all dimensions. This allows the global and local vision systems to fully capture the three-dimensional information of the markers from different angles, depths, and fields of view. At each pose point, three types of data are simultaneously recorded: the three-dimensional position of the marker in the global visual coordinate system, reconstructed by the global vision system; the three-dimensional measurement of the marker in the local visual coordinate system, acquired by the local vision system; and the current joint angles and the theoretical position and pose of the end effector in the robot coordinate system, output by the robot controller. These three types of data are mapped one-to-one through time synchronization and numbering, forming a coordinate correspondence sample set covering multiple pose states. After obtaining this sample set, two core transformation models are constructed based on the three sets of collected coordinates: a first transformation relationship that describes how the global visual coordinate system is transformed into the robot's base coordinate system, and a second transformation relationship that describes how the local visual coordinate system is transformed into the robot's base coordinate system. These two transformation relationships are used to calibrate the spatial references of the two independent visual systems, allowing the collected data to be normalized and integrated into the robot control space for fusion and comparison. Based on the spatial coordinates of each set of sample points, a transformation correspondence is established between the two visual coordinate systems and the robot coordinate system, and a universal coordinate transformation equation is constructed. This equation links the three-dimensional points observed in the visual system with the theoretical posture data of the robot body, forming a preliminary adaptive calibration model.Because the coordinate scales of the two vision systems may differ due to objective factors such as acquisition accuracy, lens focal length, imaging distortion, and equipment installation errors, a scale factor is introduced into the adaptive calibration equation as a compensation mechanism. This scale factor adjusts the scale of the visual measurement data to align with the actual physical scale of the robot system. This scale factor is obtained by fitting the spatial errors of multiple pose points. Its introduction transforms the calibration equation into a complete spatial mapping model with nonlinear terms. To solve this complete calibration equation including the scale factor, a nonlinear optimization process is designed. An iterative algorithm is employed to gradually adjust the spatial transformation parameters and the scale factor while minimizing the observation error. Each set of sample points in the vision coordinate system is transformed into a predicted value in the robot coordinate system using the target transformation relationship. The predicted value is then subtracted from the theoretical value provided by the robot controller, and the errors of all points are accumulated to form a total residual. The optimization goal is to minimize this total residual. By constructing an optimization objective function and iteratively updating it using algorithms such as the quasi-Newton method or the Levenberg-Marquardt algorithm, each update simultaneously adjusts the rotation matrix, translation vector, and scale factor to approach the optimal solution. Throughout the solution process, residual convergence conditions and parameter change thresholds are set as iterative stopping criteria to prevent the algorithm from oscillating or reaching a local minimum. Through iterative fitting and error minimization of multiple sets of sample points, the complete target transformation relationship between the global and local visual coordinate systems is determined. This transformation relationship includes standard spatial rotation and translation matrices and a calibrated scale correction factor to ensure that the data between the two visual systems is completely consistent in spatial structure and physical size.

[0039] In a specific embodiment, the process of executing step S500 may specifically include the following steps: Create a calibrated motion trajectory for the five-axis robot's workspace and control the robot to move according to the calibrated motion trajectory. Use the global visual tracking system to capture the actual joint position and posture data in each position in real time. Solve the joint parameter error of each joint manipulator according to the target transformation relationship to obtain the preliminary joint error parameters; The preliminary joint error parameters are introduced into the joint chain constraint optimization model. By introducing the geometric constraints and kinematic constraints between the joint manipulators, the preliminary joint error parameters are globally optimized to obtain the corrected parameter set that satisfies the integrity of the joint chain. The joint error correction parameters of each joint manipulator are calculated based on the correction parameter set that satisfies the integrity of the joint chain.

[0040] Specifically, a calibration trajectory is constructed that covers the workspace of the five-axis robot. Based on this trajectory, complete calibration data acquisition and error resolution are achieved. This calibration trajectory must cover the boundaries and core areas of the robot's workspace and ensure that all five joints can fully move in their respective degrees of freedom. This ensures comprehensive error excitation and robustness of the calibration model. During trajectory design, the system analyzes the robot's workspace boundary distribution, motion constraints, and typical operation paths to generate multiple key pose points. These pose points should exhibit spatial non-coplanarity, nonlinearity, and full pose variation, ensuring that all error types are captured and recognized by the model. Under system control, the five-axis robot moves point by point along the set calibration trajectory. During each dwell point, a global visual tracking system initiates synchronous acquisition, acquiring joint spatial position and pose data in real time. This data acquisition process is based on a calibrated multi-camera network within the vision system. It constructs a 3D reconstruction model from multi-view image data, extracts landmarks at each joint, and calculates their 3D coordinates in the global visual coordinate system. Simultaneously, the robot control system returns the theoretical joint angles and end-effector positions corresponding to the current pose in real time. The information obtained by the two systems is aligned using timestamps and pose numbers to form a complete pose dataset. The pose of each sampling point is mapped based on the target transformation relationship between the global visual coordinate system and the robot control coordinate system. By uniformly converting the joint spatial positions measured by the vision system to the robot coordinate system and directly comparing the spatial differences between the visual observations and the robot control values, the errors of each joint at different poses are decomposed and fitted. This process accounts for multiple error sources, including joint position offset, rotation angle error, zero point offset, and link length differences. By solving the error residual model, error estimates for each joint at each pose point are extracted. These errors are then statistically fused across the entire trajectory to form a preliminary set of joint error parameters. These preliminary error parameters are then incorporated into a global joint chain constraint optimization model for joint optimization. This model, based on the complete kinematic structure of the five-axis robot, constructs a topological structure that includes the connection logic and transmission paths between all joints, and embeds geometric and kinematic constraints within the model. Geometric constraints define the actual physical connection between joints, such as fixed link lengths, consistency in the relative postures of adjacent joints, and no offset in the connecting axis. Kinematic constraints ensure that during robot operation, the motion transformations of each joint satisfy the forward kinematic relationship and the continuity of the end path. For example, the forward solution to the equation satisfies the functional consistency of the input joint angle and the output end coordinate, and the inverse solution does not cause motion offset due to error accumulation. Based on this multi-constraint optimization model, a global optimization of the preliminary joint error parameters is performed. By constructing an objective function to minimize the sum of errors under the entire joint chain, the parameter offset values of each joint are adjusted simultaneously in each iteration, while maintaining the geometric structure between the links.Redundant data from multiple pose points is introduced during the solution process, ensuring that the optimization model averages residuals and compresses the error distribution based on a large number of observations. This results in a final correction that minimizes the overall spatial error while also ensuring structural continuity across joints. This optimization process uses a nonlinear optimization algorithm to achieve variable iterative convergence. Physical constraints, such as link lengths not exceeding the permitted range and rotation axis orientations remaining within the desired tolerance, are checked at each step, ensuring engineering feasibility and control system deployability. After joint chain constraint optimization, a set of correction parameters is obtained that ensures spatial continuity, structural integrity, and control consistency. This parameter set serves as the basis for the final error compensation of each joint and is decoupled into different compensation terms, including joint angle deviation corrections, link size adjustments, connection point coordinate restoration parameters, and coordinate system mapping correction coefficients. These correction parameters can be dynamically loaded into the joint control logic of the robot control system as compensation factors, participating in compensation execution during each motion planning step. They also serve as important input data for the robot's manufacturing accuracy evaluation, structural maintenance optimization, and visual recognition feedback loop. By integrating the correction parameters into the control model, the problem of reduced positioning accuracy caused by structural errors, assembly deviations or mechanical drift caused by long-term operation can be effectively reduced, and the repeatability, stability and spatial accuracy of the robot in complex tasks can be improved.

[0041] In a specific embodiment, the step of introducing preliminary joint error parameters into the joint chain constraint optimization model, and globally optimizing the preliminary joint error parameters by introducing geometric constraints and kinematic constraints between joint manipulators to obtain a corrected parameter set that satisfies the integrity of the joint chain may specifically include the following steps: A complete kinematic model of the five-axis robot joint chain is established, and the preliminary joint error parameters are combined with the theoretical parameters of the complete kinematic model to construct the joint chain parameter set; Construct the joint chain forward kinematics equation according to the joint chain parameter set, and use the joint chain forward kinematics equation as the joint chain constraint condition; Using the actual joint position data measured by the global visual tracking system and the local visual measurement system, the position closed-loop error of different joint combinations is calculated and the closed-loop constraint equation of the joint chain is established; Based on the joint chain constraint conditions and the joint chain closed-loop constraint equation, weighted least squares optimization is performed to obtain the joint chain constraint optimization model; Introducing geometric and kinematic constraints between joints into the joint chain constraint optimization model to construct a constrained optimization problem. The constrained optimization problem is solved by hybrid nonlinear optimization to obtain a modified parameter set that satisfies the integrity of the joint chain.

[0042] Specifically, a kinematic model of the joint chain is theoretically constructed, describing the connection relationships, motion patterns, and geometric topology between the five joints. A standard kinematic modeling framework is established based on theoretical parameters provided by the robot manufacturer. Each joint includes parameters such as the rotation axis direction, link length, joint offset, and joint torsion angle. The five joints are sequentially connected according to the DH parameterization rule to form a complete kinematic chain model from the base to the end effector. This yields a set of kinematic parameters under ideal error-free conditions. This allows the theoretical pose of the robot end effector to be calculated for any joint angle input using standard forward kinematic equations. To incorporate error modeling, preliminary joint error parameters estimated by a vision system combined with calibration trajectories are incorporated into the theoretical model, resulting in an error-inclusive kinematic model that more closely matches the actual structural state. This process maps the preliminary error parameters for each joint to its theoretical geometry. For example, an angle zero offset term is introduced for the rotational joints, and a length perturbation term and a connection position offset term are introduced for the link parameters. This results in a joint chain parameter set that includes both theoretical values and correction terms. Based on this parameter set, a new joint chain forward kinematic equation is constructed. This equation serves as a constraint function describing the theoretically achievable output position of the robot for different joint angle inputs. This forward kinematic equation maintains a matrix chain structure of multiplications, accumulating the end pose through the transformation matrices of each joint. The error term in each transformation is explicitly expressed, allowing it to participate in the optimization process. The core function of this equation is to provide an objective function framework that derives and corrects various parameter deviations by comparing the data measured by the vision system with the data predicted by the current model. To improve the constraint model's strength and system recognition capabilities, a joint chain closed-loop constraint is introduced. This constraint identifies kinematic consistency relationships between multiple pose combinations and converges the error search range. In practice, multiple different joint combinations are selected as inputs, and the actual measured spatial positions of the joint markers are recorded for each combination. The forward kinematic equation is then used to calculate the theoretical output pose for that combination. The difference between the two is the position closed-loop error. By analyzing error statistics and distribution under multiple joint combinations, a set of error mapping functions consisting of different input action combinations is established. These error values form the closed-loop constraint equations for the joint chain. This equation, driven by the optimization of reprojection errors and spatial transformation offsets at multiple points, spatially closes the error loops formed by each independent action. After constructing all kinematic and closed-loop constraints, they are incorporated into a weighted least squares optimization framework. The goal of the optimization process is to find an optimal set of error correction parameters that minimizes the sum of the residual errors under all constraints.Due to the varying error sources, observation accuracy, and constraint strength, a weight is assigned to each constraint type. Position errors measured by the vision system are given a higher weight due to their high accuracy, while kinematic structural terms are set as hard constraints due to their strong stability. This weighted approach improves the convergence speed and physical plausibility of the optimization results, forming a preliminary joint chain constraint optimization model. To enhance the model's physical feasibility and structural consistency, two types of explicit structural constraints are introduced into the optimization model: geometric constraints and kinematic constraints. Geometric constraints require that the spatial connections between adjacent joints conform to mechanical assembly requirements, such as maintaining parallel or coplanar axes, limiting link length variations within tolerances, and preventing connection points from drifting to non-physical locations. Kinematic constraints require that the optimized parameters still meet the robot's supported motion logic. For example, the range of joint angle variation cannot exceed the servo limits, the rotational direction between two joints must remain consistent, and the overall structural output must be continuous, non-abrupt, and usable for control system planning. These constraints are formally expressed as constraints on the optimization variables and enforced during the solution process. This optimization problem with explicit constraints is treated as a nonlinear problem and solved using a hybrid nonlinear optimization algorithm. The algorithm consists of a two-tiered mechanism: global search and local convergence. In the initial stage, random perturbations and genetic mechanisms are used to obtain a generalized optimal solution interval. In the local stage, a gradient descent-based algorithm is used for rapid convergence. The entire solution process continuously evaluates the convergence rate of the current error residual function and the trend of parameter changes. When the sum of the errors falls below a set threshold or the parameter changes stabilize, the solution is terminated and the final optimized error correction parameter set is output.

[0043] See also Figure 2 , Figure 2 The schematic block diagram of the structure of the automatic calibration device 200 of the joint manipulator based on the visual system provided in the embodiment of the present application is as follows: Figure 2 As shown, the automatic calibration device 200 of the joint manipulator based on the vision system includes: The acquisition module 210 is used to acquire multi-view images of the joint manipulator of the five-axis robot to obtain original image data; An extraction module 220 is used to extract occlusion regions from the original image data to obtain a set of joint occlusion regions; A stitching module 230 is used to perform unordered image stitching and marker feature extraction on the original image data based on the joint occlusion region set to obtain joint feature image data; A construction module 240 is used to construct an adaptive calibration equation based on the joint feature image data and solve the target transformation relationship between the global visual coordinate system and the local visual coordinate system; The generation module 250 is used to perform multi-joint collaborative calibration and joint chain constraint optimization of the five-axis robot according to the target transformation relationship, and generate joint error correction parameters.

[0044] Through the collaborative efforts of these components, a multi-view 3D visual measurement network combining a global visual tracking system and a local visual measurement system achieves comprehensive monitoring of a five-axis robot articulated manipulator, effectively addressing the inability to obtain complete information from a single viewpoint. A binary tree model algorithm is used to finely extract occluded regions, and combined with unordered image stitching techniques, complete information about occluded joints is restored. This overcomes the limitations of traditional methods in calibration accuracy in complex environments and frees the calibration process from the constraints of environmental occlusion. A scale factor correction mechanism and a multi-pose sampling strategy are introduced to address the fusion issues between the global and local visual coordinate systems, establish an accurate coordinate mapping, and achieve precise data integration between systems of different scales. A joint chain constraint optimization model is used to globally optimize the calibration parameters, combining geometric and kinematic constraints. This ensures that the calibration results meet the physical characteristics of the joint mechanism, enhancing the physical significance and practicality of the calibration. A recursive least squares method combined with a nonlinear optimization algorithm achieves high-precision calculation of joint error correction parameters, significantly improving the positioning and kinematic accuracy of the five-axis robot articulated manipulator.

[0045] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described systems, systems and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0046] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.

[0047] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for automatic calibration of joint manipulators based on a visual system, characterized in that: include: Perform multi-view image acquisition on the joint manipulator of the five-axis robot to obtain original image data; Extracting occlusion regions from the original image data to obtain a set of joint occlusion regions; Based on the joint occlusion area set, performing disordered image splicing and marker point feature extraction on the original image data to obtain joint feature image data; Constructing an adaptive calibration equation based on the joint feature image data and solving a target transformation relationship between a global visual coordinate system and a local visual coordinate system; Multi-joint collaborative calibration and joint chain constraint optimization of the five-axis robot are performed according to the target transformation relationship to generate joint error correction parameters.

2. The automatic calibration method of the joint manipulator based on the visual system according to claim 1 is characterized in that: The multi-view image acquisition of the joint manipulator of the five-axis robot to obtain original image data includes: M industrial cameras are fixedly installed around the workspace of the five-axis robot to form a global visual tracking system, and N high-resolution cameras are fixedly installed on the end effector of the joint manipulator to form a local visual measurement system; Performing external parameter calibration on the industrial cameras in the global visual tracking system through a calibration plate to determine the spatial position relationship between the industrial cameras, and performing binocular stereo vision configuration on the industrial cameras in the local visual measurement system to establish a camera parameter matrix; Setting a circular marking point at each joint of the joint manipulator, and synchronizing the global visual tracking system and the local visual measurement system based on the spatial position relationship and the camera parameter matrix to ensure image acquisition synchronization; The joint manipulator is controlled to move along a preset path, and image acquisition is performed synchronously by the global vision tracking system and the local vision measurement system to obtain original image data including a global vision tracking image sequence and a local vision measurement image sequence.

3. The automatic calibration method of the joint manipulator based on the visual system according to claim 2 is characterized in that: The extracting the occlusion region from the original image data to obtain a joint occlusion region set includes: Selecting a first image containing a complete view of the manipulator joints from the global visual tracking image sequence of the original image data as a reference image, and comparing each second image frame subsequently acquired in the global visual tracking image sequence with the reference image; Extracting features from the reference image and the second image to obtain a feature point set for each frame of the second image, and matching the feature point sets to establish a correspondence between the images; Calculating a homography matrix based on the correspondence between the images and the local vision measurement image sequence, projecting the second image into a reference image coordinate system to perform image difference calculation to determine an initial range of the occluded area; Recursively splitting the target area into sub-areas according to the initial range of the occlusion area, and determining leaf nodes marked as occlusion areas; All leaf nodes marked as occlusion regions are combined to form a preliminary occlusion region set, and the preliminary occlusion region set is morphologically processed to obtain a joint occlusion region set.

4. The automatic calibration method of the joint manipulator based on the visual system according to claim 3 is characterized in that: The step of recursively splitting the target area into sub-areas according to the initial range of the occlusion area and determining leaf nodes marked as occlusion areas comprises: Taking the initial range of the occluded area as the root node of the binary tree, calculating the variance of the grayscale values of all pixels in the root node and the average value of the gradient amplitude to obtain the feature metric of the root node; Comparing the pixel variance within the root node with a first target value, and when the pixel variance is greater than the first target value, splitting the root node into two first nodes along the long axis of the region, and adding the first nodes to a queue to be processed; Sequentially taking out first nodes from the queue to be processed, calculating variance and gradient information of pixels within each first node, and when the variance of pixels within the first node is greater than a second target value, continuing to split the first node into two second nodes along the direction of maximum gradient; Calculating the area of the second node; when the area of the second node is less than the third target value square pixel and the pixel variance is less than the fourth target value, marking the second node as a leaf node; otherwise, adding the second node back to the queue for processing; Iteratively executing a splitting process based on the queue to be processed to obtain a binary tree structure, wherein each leaf node in the binary tree structure represents a sub-region with similar pixel characteristics; All leaf nodes are marked and judged. When the difference between the average grayscale value in the leaf node and the grayscale value of the corresponding area of the reference image exceeds the set fifth target value, the leaf node is marked as an occlusion area.

5. The automatic calibration method of the joint manipulator based on the visual system according to claim 4 is characterized in that: The method of performing disordered image splicing and marker feature extraction on the original image data based on the joint occlusion area set to obtain joint feature image data includes: Screening an image subset containing different perspective information of the occluded joint from the multi-perspective image sequence of the original image data, and calculating the image feature point density and perspective coverage of the image subset to obtain a candidate image set; Extracting feature points and calculating feature matching for each image in the candidate image set to construct an image similarity matrix; Determining an image stitching sequence based on the image similarity matrix, and calculating a geometric transformation matrix for each pair of adjacent images in the image stitching sequence to obtain an image mapping relationship; Calculating local transformation parameters based on the image mapping relationship and performing fine-grained splicing through an elastic transformation model to obtain a complete image of the repaired joint; The repaired complete joint image is subjected to marker point feature extraction and multi-view three-dimensional reconstruction calculation to obtain joint feature image data.

6. The automatic calibration method of joint manipulator based on visual system according to claim 5 is characterized in that: The method of performing marker point feature extraction and multi-view 3D reconstruction calculation on the repaired complete joint image to obtain joint feature image data includes: Performing adaptive threshold segmentation on the repaired complete joint image to obtain a segmented image with prominent marker features; Generating a candidate region of a marker point based on the segmented image of the prominent marker point feature, and calculating the sub-pixel center coordinates of the marker point according to the candidate region of the marker point to obtain high-precision marker point position information; For the binocular images acquired by the local vision measurement system, stereo matching is performed using epipolar constraint and normalized cross-correlation algorithm to obtain the corresponding relationship of the marker points; Based on the correspondence relationship of the marker points and the camera parameter matrix, a triangulation method is used to calculate the local three-dimensional coordinates of the marker points to obtain a three-dimensional point set in the local visual measurement system; According to the three-dimensional point set in the local vision measurement system, the multi-camera images in the global vision tracking system are subjected to multi-view three-dimensional reconstruction and global optimization of camera observation data by applying the bundle adjustment method to obtain joint feature image data.

7. The automatic calibration method of the joint manipulator based on the visual system according to claim 6 is characterized in that: The method of constructing an adaptive calibration equation based on the joint feature image data and solving a target transformation relationship between a global visual coordinate system and a local visual coordinate system includes: Based on the joint feature image data, a global visual coordinate system, a local visual coordinate system, and a robot base coordinate system are defined, and the joint manipulator of the five-axis robot is controlled to move to different positions. At each position, the three-dimensional coordinates of the joint markers in the global visual tracking system and the local visual measurement system, as well as the joint angles reported by the robot control system, are simultaneously recorded to construct a coordinate corresponding sample set; Constructing a first transformation matrix from the global to the robot coordinate system and a second transformation matrix from the local to the robot coordinate system based on the coordinate corresponding sample set to form an adaptive calibration equation; A scale factor is introduced into the adaptive calibration equation to obtain a complete calibration equation including the scale factor, and a nonlinear optimization solution is performed on the complete calibration equation including the scale factor to obtain a target transformation relationship between the global visual coordinate system and the local visual coordinate system.

8. The method for automatic calibration of joint manipulator based on visual system according to claim 7, characterized in that: The method of performing multi-joint collaborative calibration and joint chain constraint optimization of the five-axis robot according to the target transformation relationship to generate joint error correction parameters includes: Creating a calibrated motion trajectory for the workspace of the five-axis robot, and controlling the five-axis robot to move according to the calibrated motion trajectory, and capturing the actual position and posture data of the joints in each posture in real time through the global visual tracking system; Solve the joint parameter error of each joint manipulator according to the target transformation relationship to obtain preliminary joint error parameters; The preliminary joint error parameters are introduced into the joint chain constraint optimization model, and the preliminary joint error parameters are globally optimized by introducing geometric constraints and kinematic constraints between joint manipulators to obtain a correction parameter set that satisfies the integrity of the joint chain; The joint error correction parameters of each joint manipulator are calculated according to the correction parameter set that satisfies the integrity of the joint chain.

9. The method for automatic calibration of joint manipulator based on visual system according to claim 8, characterized in that: The preliminary joint error parameters are introduced into the joint chain constraint optimization model, and the preliminary joint error parameters are globally optimized by introducing geometric constraints and kinematic constraints between joint manipulators to obtain a correction parameter set that satisfies the integrity of the joint chain, including: Establishing a complete kinematic model of a five-axis robot joint chain, and combining the preliminary joint error parameters with theoretical parameters of the complete kinematic model to construct a joint chain parameter set; Constructing a joint chain forward kinematics equation according to the joint chain parameter set, and using the joint chain forward kinematics equation as a joint chain constraint condition; Utilizing the actual joint position data measured by the global visual tracking system and the local visual measurement system, the position closed-loop errors during the combined motions of different joints are calculated, and a closed-loop constraint equation for the joint chain is established; Performing weighted least squares optimization based on the joint chain constraint conditions and the joint chain closed-loop constraint equation to obtain a joint chain constraint optimization model; Introducing geometric constraints and kinematic constraints between joints into the joint chain constraint optimization model to construct a constrained optimization problem; The constrained optimization problem is solved by hybrid nonlinear optimization to obtain a modified parameter set that satisfies the integrity of the joint chain.

10. An automatic calibration device for joint manipulator based on a visual system, characterized in that: The method for automatically calibrating an articulated manipulator based on a vision system according to any one of claims 1 to 9 is configured to include: The acquisition module is used to acquire multi-view images of the joint manipulator of the five-axis robot to obtain original image data; An extraction module, configured to extract occlusion regions from the original image data to obtain a set of joint occlusion regions; A splicing module, configured to perform unordered image splicing and marker feature extraction on the original image data based on the joint occlusion region set to obtain joint feature image data; A construction module is used to construct an adaptive calibration equation based on the joint feature image data and solve the target transformation relationship between the global visual coordinate system and the local visual coordinate system; A generation module is used to perform multi-joint collaborative calibration and joint chain constraint optimization of the five-axis robot according to the target transformation relationship, and generate joint error correction parameters.

Citation Information

Cited By

  • Industrial robot optical navigation anti-shielding tracking system, method and device and medium

    CN120755891A

  • Cooperative control method and system for acoustic module injection molding production line

    CN121277138A

  • Six-axis manipulator control method based on global interior point iteration multi-starting-point solution

    CN121290440A

  • Storage safety monitoring method and system based on visual inspection

    CN121564635A

  • A warehouse safety monitoring method and system based on visual detection

    CN121564635B