Special machine hand-eye calibration error compensation method and system based on residual full connection network, storage medium and program product

CN122820863APending Publication Date: 2026-09-25CHANGSHU INSTITUTE OF TECHNOLOGY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611186445.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-06
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0004]针对上述现有技术的缺陷,本发明提供了一种基于残差全连接网络的专机手眼标定误差补偿方法,解决一次性刚体标定难以完全吸收位置相关残余误差,进而影响坐标映射精度的问题

Benefits of technology

[0031]本发明将传统刚体手眼标定与数据驱动残差补偿结合起来,并针对三维坐标残差回归任务设计了适配的残差全连接网络结构,利用经训练的残差全连接网络学习工作空间内非线性误差场,通过将预测残差叠加至初始预测坐标实现二次修正,既保留了手眼标定的几何基础,又能够补偿传统刚体模型难以描述的空间非线性残余误差,提高3D视觉坐标转换精度与一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820863A_ABST
    Figure CN122820863A_ABST
Patent Text Reader

Abstract

The application discloses a special machine hand-eye calibration error compensation method based on a residual full connection network, the special machine is a translational motion mechanism, and the method comprises the following steps: establishing an initial rigid body mapping between a camera coordinate system and a special machine coordinate system by using chessboard corner points; constructing a residual full connection network with a predicted coordinate as input and a three-dimensional residual error between a real coordinate and the predicted coordinate as a supervised label; outputting a predicted three-dimensional error vector from the residual full connection network; and finally adding the predicted three-dimensional error vector and the predicted coordinate to obtain a coordinate of a target point in the special machine coordinate system after compensation. The application also discloses a corresponding compensation system and a storage medium and a program product. The application can eliminate position-related residual errors and improve precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method, system, storage medium, and program product for hand-eye calibration error compensation for dedicated machines based on residual fully connected networks. Background Technology

[0002] In 3D vision-guided systems, hand-eye calibration is fundamental to establishing the spatial mapping relationship between the camera coordinate system and the motion coordinate system. Existing technologies often employ one-time rigid body calibration correction. For example, in patent document CN114519738A, the point cloud coordinates in the camera are transformed to transformed coordinates in the robot coordinate system using matrix transformation relationships. The actual coordinates of the point cloud in the robot coordinate system are then acquired. The ICP algorithm is used to match the two sets of point sets—the transformed coordinates and the actual coordinates—to obtain the hand-eye calibration error matrix for translation and rotation between the two sets of point sets, thus yielding the corrected hand-eye calibration matrix.

[0003] However, in a Cartesian coordinate system special machine system that performs X, Y, and Z translation axes, one-time rigid body calibration still cannot completely absorb factors such as camera installation error, mechanism adjustment error, kinematic nonlinearity, visual measurement noise, and human touch error. These factors will form position-related residual errors in the workspace, which will ultimately affect the coordinate mapping accuracy. Summary of the Invention

[0004] To address the shortcomings of the existing technology, this invention provides a special-purpose aircraft hand-eye calibration error compensation method based on residual fully connected networks, solving the problem that one-time rigid body calibration cannot completely absorb position-related residual errors, thus affecting coordinate mapping accuracy. This invention also provides a special-purpose aircraft hand-eye calibration error compensation system, system, storage medium, and program product based on residual fully connected networks.

[0005] The technical solution of the present invention is as follows:

[0006] A method for hand-eye calibration error compensation for a special-purpose aircraft based on a residual fully connected network, wherein the special-purpose aircraft is a translational motion mechanism, the method comprising:

[0007] Step 1: Perform hand-eye calibration based on the corner points of the chessboard grid, establish the initial rigid body mapping from the camera coordinate system to the special aircraft coordinate system, and obtain the hand-eye calibration matrix;

[0008] Step 2: Calculate the predicted coordinates of the target point in the dedicated machine coordinate system based on the coordinates of the target point in the camera coordinate system, the position of the machine end in the dedicated machine coordinate system when the target point is sampled, and the hand-eye calibration matrix;

[0009] Step 3: Input the predicted coordinates into the trained residual fully connected network to obtain the predicted 3D error vector output by the residual fully connected network;

[0010] The residual fully connected network is used to output a corresponding predicted 3D error vector based on the input predicted coordinates. The residual fully connected network includes at least two feature extraction modules, feature compression modules, and linear output layers connected in sequence. The feature extraction module includes an input expansion layer and a multi-scale residual feature extraction layer connected in sequence. The feature compression module includes a feature compression layer and an activation layer connected in sequence.

[0011] Step 4: Add the predicted 3D error vector and the predicted coordinates to obtain the compensated coordinates of the target point in the special aircraft coordinate system.

[0012] Furthermore, the input extension layer is a fully connected structure used to map low-dimensional input features to a high-dimensional feature space. The input extension layers in different feature extraction modules map the input features from three dimensions to high dimensions step by step according to the connection order.

[0013] By mapping low-dimensional coordinates to a high-dimensional feature space through the input extension layer, the feature representation capability is enhanced, providing a foundation for the subsequent multi-scale residual feature extraction layer to extract complex nonlinear error relationships.

[0014] Furthermore, the input expansion layer progressively maps the input features from three dimensions to higher dimensions, sequentially passing through 32 dimensions, 64 dimensions, and up to 128 dimensions.

[0015] Furthermore, the multi-scale residual feature extraction layer includes at least two cascaded residual blocks. Each residual block includes a main branch, an identity jump connection, a residual summing node, and an output activation layer. The main branch is used to process the input features of the residual block sequentially through a first fully connected layer, a first main branch activation layer, a second fully connected layer, and a second main branch activation layer to obtain the main branch output. The residual summing node adds the main branch output to the input features of the residual block passed through the identity jump connection element-wise. The output activation layer activates the output of the element-wise added features through an activation function.

[0016] Residual connections enable the network to retain basic information in the input features while learning complex nonlinear mappings, effectively alleviating the gradient degradation problem in deep network training and improving convergence consistency; multi-scale residual feature extraction layers extract nonlinear features of the error field step by step under different feature dimensions, enhancing the ability to learn the changing laws of complex error fields.

[0017] Furthermore, the feature compression layer is a fully connected structure used to reduce the dimensionality of the high-dimensional feature map obtained by the last feature extraction module. The dimensionality-reduced features are non-linearly activated by the activation layer and finally further reduced to three-dimensional output by the linear output layer.

[0018] The high-dimensional feature mapping is reduced to a smaller dimension by a feature compression layer, which retains the core information that is most discriminative for error prediction and reduces the impact of high-dimensional redundancy on small sample generalization performance. After compression, nonlinear activation is used to further enhance the ability of compact features to represent complex error patterns.

[0019] Furthermore, the trained residual fully connected network is trained in the following way: using the predicted coordinates as training input, and using the three-dimensional error vector formed by the difference between the actual coordinates of the target point in the special aircraft coordinate system and the predicted coordinates as supervision labels, the residual fully connected network is trained under supervision, so that the residual fully connected network learns the nonlinear mapping relationship from the predicted coordinates to the three-dimensional error vector.

[0020] Furthermore, during the training of the residual fully connected network, each dimension of the training input and each dimension of the supervision label are subjected to Min-Max normalization. During the inference phase, the residual fully connected network outputs the normalized result and undergoes inverse normalization to obtain the predicted three-dimensional error vector.

[0021] The numerical ranges of input coordinates and 3D residuals vary greatly across different dimensions. Directly inputting them into the network can easily lead to uneven gradient scales in each dimension, affecting training consistency and convergence speed. Therefore, Min-Max normalization is used to ensure training consistency.

[0022] Another technical solution of the present invention is: a special-purpose aircraft hand-eye calibration error compensation system based on a residual fully connected network, wherein the special-purpose aircraft is a translational motion mechanism, and the system includes:

[0023] The hand-eye calibration module is used to perform hand-eye calibration based on the corner points of the chessboard grid, establish an initial rigid body mapping from the camera coordinate system to the special aircraft coordinate system, and obtain the hand-eye calibration matrix.

[0024] The predicted coordinate calculation module is used to calculate the predicted coordinates of the target point in the dedicated machine coordinate system based on the coordinates of the target point in the camera coordinate system, the position of the machine end in the dedicated machine coordinate system when the target point is sampled, and the hand-eye calibration matrix.

[0025] The residual prediction module is used to input the predicted coordinates into a trained residual fully connected network to obtain the predicted three-dimensional error vector output by the residual fully connected network.

[0026] The residual fully connected network is used to output a corresponding predicted 3D error vector based on the input predicted coordinates. The residual fully connected network includes at least two feature extraction modules, feature compression modules, and linear output layers connected in sequence. The feature extraction module includes an input expansion layer and a multi-scale residual feature extraction layer connected in sequence. The feature compression module includes a feature compression layer and an activation layer connected in sequence.

[0027] The error compensation module is used to add the predicted three-dimensional error vector and the predicted coordinates to obtain the compensated coordinates of the target point in the special aircraft coordinate system.

[0028] Another technical solution of the present invention is: a computer-readable storage medium storing a computer program thereon, wherein when the computer program is executed by a processor, the aforementioned method for compensation of hand-eye calibration error of a special machine based on a residual fully connected network is implemented.

[0029] Another technical solution of the present invention is: a computer program product, including a computer program, which, when executed by a processor, implements the aforementioned method for compensation of hand-eye calibration error of a special-purpose machine based on a residual fully connected network.

[0030] Compared with the prior art, the present invention has the following advantages:

[0031] This invention combines traditional rigid body hand-eye calibration with data-driven residual compensation, and designs an adapted residual fully connected network structure for 3D coordinate residual regression tasks. The trained residual fully connected network learns the nonlinear error field in the workspace, and achieves secondary correction by superimposing the predicted residuals onto the initial predicted coordinates. This not only preserves the geometric basis of hand-eye calibration, but also compensates for spatial nonlinear residual errors that are difficult to describe by traditional rigid body models, thereby improving the accuracy and consistency of 3D visual coordinate transformation. Attached Figure Description

[0032] Figure 1 This is a flowchart illustrating the hand-eye calibration error compensation method for dedicated aircraft based on residual fully connected networks, as an example.

[0033] Figure 2 This is a schematic diagram of the structure of a residual fully connected network.

[0034] Figure 3 This is a schematic diagram of the structure of a residual block in a fully connected residual network.

[0035] Figure 4 A histogram of hand-eye calibration error statistics.

[0036] Figure 5 This is the error statistics histogram after error compensation.

[0037] Figure 6 Box plot showing the distribution before and after error compensation.

[0038] Figure 7 Heat map of spatial error distribution for hand-eye calibration.

[0039] Figure 8 This is a heatmap showing the spatial error distribution after error compensation. Detailed Implementation

[0040] The present invention will be further described below with reference to embodiments, but these are not intended to limit the scope of the invention.

[0041] Please combine Figure 1 As shown, the embodiment of the present invention provides a method for hand-eye calibration error compensation for a special-purpose aircraft based on a residual fully connected network. The special-purpose aircraft is a translational motion mechanism. In this embodiment, the translational motion mechanism is a three-axis Cartesian coordinate system special-purpose aircraft executed by three translation axes (X, Y, and Z). The method specifically includes the following steps:

[0042] Step 1: Perform hand-eye calibration based on the corner points of the chessboard grid, establish the initial rigid body mapping from the camera coordinate system to the special aircraft coordinate system, and obtain the hand-eye calibration matrix.

[0043] During the calibration process, the camera is moved within the workspace using the three linear axes X, Y, and Z to complete corner sampling and machine coordinate recording at different planar positions and heights.

[0044] To improve the adequacy of sample distribution in the workspace, this embodiment employs a "three-layer height + nine-grid sampling" method for data acquisition. The specific process is as follows:

[0045] (1) Set three different sampling heights by adjusting the Z-axis position;

[0046] (2) Within each height layer, perform single-layer 3×3 grid position sampling with the target corner point as the center;

[0047] (3) At each sampling location, the coordinates of the corner point in the camera coordinate system are extracted using a 3D surface structured light camera. ;

[0048] (4) Simultaneously record the position of the machine end in the dedicated machine coordinate system during this sampling. ;

[0049] (5) After completing all 27 sets of sampling, control the calibration needle tip to accurately touch the same target corner point and obtain its true coordinates in the special machine coordinate system. .

[0050] (6) By , , Construct a rigid body transformation model and solve the rotation matrix using the least squares method. Translation vector Thus, a hand-eye calibration matrix is ​​established.

[0051] .

[0052] The principle of this calibration process is: the target corner point coordinates in the machine coordinate system are transformed by a rigid body. and Then, it can be expressed as the displacement increment of the end calibration needle tip relative to the current end position. Therefore, the predicted coordinates of this corner point in the special aircraft coordinate system in the i-th sample can be expressed as:

[0053]

[0054] Ideally, predict coordinates Should be consistent with the actual coordinates Therefore, the hand-eye alignment problem can be expressed as:

[0055]

[0056] Rearranging the above equation, we get:

[0057]

[0058] To facilitate the subsequent solution using the least squares method, we define auxiliary quantities:

[0059]

[0060] The above formula can then be written as:

[0061]

[0062] Therefore, hand-eye calibration can be transformed into a problem of solving rigid body transformations between two sets of three-dimensional point sets.

[0063]

[0064] Solving for the optimal rotation matrix using the least squares criterion Translation vector .

[0065] The solution process is as follows:

[0066] (1) First, calculate the centroids of the two sets of points:

[0067]

[0068] (2) Decentralized processing:

[0069]

[0070] (3) Construct the covariance matrix:

[0071]

[0072] (4) Perform singular value decomposition:

[0073]

[0074] right After performing singular value decomposition, the optimal rotation matrix and translation vector can be obtained; when det( When ) < 0, it is necessary to adjust The last column undergoes sign correction to ensure... It is a valid rotation matrix.

[0075] Step 2: Calculate the predicted coordinates of the target point in the dedicated machine coordinate system based on the coordinates of the target point in the camera coordinate system, the position of the end effector of the machine in the dedicated machine coordinate system when sampling the target point, and the hand-eye calibration matrix.

[0076] Specifically, according to the description in step 1, the coordinates of the target point in the camera coordinate system are: The corresponding position of the machine end in the dedicated machine coordinate system The predicted coordinates are obtained by the following formula.

[0077] .

[0078] Step 3: Input the predicted coordinates into the trained residual fully connected network to obtain the predicted 3D error vector output by the residual fully connected network.

[0079] Since the input data in this embodiment is three-dimensional spatial coordinates, and the output data is the three-dimensional error vector at the corresponding position, it is essentially a nonlinear regression problem between low-dimensional vectors. It lacks the local spatial texture structure found in image data and the temporal correlation features found in sequence data. Therefore, convolutional networks or recurrent networks are not suitable as the main modeling structure for this task. Based on this characteristic, this embodiment uses a fully connected neural network as the basic framework to fully learn the nonlinear mapping relationship between spatial coordinates and error vectors. Furthermore, considering the potential for gradient degradation and training instability as network depth increases, a residual connection mechanism is introduced to construct a residual fully connected network, thereby improving the model's convergence stability and deep feature representation capability.

[0080] Please combine Figure 2 As shown, the residual fully connected network is used to output a corresponding predicted 3D error vector based on the input predicted coordinates. In this embodiment, the residual fully connected network includes three sequentially connected feature extraction modules 1, feature compression modules 2, and a linear output layer 3. Feature extraction module 1 includes a sequentially connected input extension layer 1a and a multi-scale residual feature extraction layer, which includes two cascaded residual blocks 1b. Feature compression module 2 includes a sequentially connected feature compression layer 2a and an activation layer 2b.

[0081] The general form of a fully connected layer is defined as follows:

[0082]

[0083] Among them Input feature vector, This is the weight matrix. The bias vector is the input dimension of this layer. The output dimension is , , , .

[0084] The network activation function uses the modified linear unit. Its function is to introduce nonlinear expressive power after fully connected mapping and suppress negative responses.

[0085] For the input of this network, let the initial predicted coordinates of the i-th sample obtained through hand-eye calibration be:

[0086]

[0087] in This represents the predicted coordinates of the same spatial location in the aircraft's coordinate system. Let... Because the numerical range of input coordinates can vary significantly across different dimensions, directly inputting them into the network can easily lead to uneven gradient scales across dimensions, thus affecting training consistency and convergence speed. Therefore, the input coordinates... A stepwise Min-Max normalization process is employed. To avoid premature leakage of validation and test set information during preprocessing, after dividing the training, validation, and test sets, this paper only uses the minimum and maximum values ​​of each component of the 3D coordinates from the training set samples, and fixes this set of statistics for normalization and denormalization during the training, validation, and testing phases. Then the... The normalized network input for each sample can be written as: ,in To avoid the special case where the denominator is zero, when a dimension satisfies the condition that the maximum and minimum values ​​are equal, the normalization result of that dimension can be set to 0.

[0088] Ultimately, the neural network does not actually receive the original coordinates as input. Instead, it is the normalized 3D input vector. In subsequent training and inference processes, the network performs feature extraction and error prediction based on this normalized input.

[0089] The input expansion layer 1a is a fully connected structure used to map low-dimensional input features to a high-dimensional feature space. The input expansion layers 1a in different feature extraction modules 1 progressively map the input features from three dimensions to higher dimensions according to their connection order. In this embodiment, the network has three feature extraction modules 1, and therefore three input expansion layers 1a, sequentially expanding the three-dimensional input from 3D to 32D to 64D to 128D.

[0090] Please combine again Figure 3 As shown, the residual block 1b in the multi-scale residual feature extraction layer includes a main branch, an identity jump connection 1b1, a residual summing node 1b2, and an output activation layer 1b3. The main branch processes the input features of the residual block sequentially through a first fully connected layer 1b4, a first main branch activation layer 1b5, a second fully connected layer 1b6, and a second main branch activation layer 1b7 to obtain the main branch output. The residual summing node 1b2 adds the main branch output element-wise to the input features of the residual block 1b passed through the identity jump connection 1b1. The output activation layer 1b3 activates the element-wise added features using an activation function. The first fully connected layer 1b4 and the second fully connected layer 1b6 are identical-dimensional fully connected mappings, and no dimensionality increase or decrease operations are performed within the residual block 1b. The first main branch activation layer 1b5, the second main branch activation layer 1b7, and the output activation layer 1b3 all use the ReLU activation function.

[0091] After the network input undergoes 128-dimensional high-level feature extraction in the final feature extraction module 1, directly feeding these high-dimensional features into the output layer would result in a large output parameter size and potentially increase the risk of overfitting. Therefore, the 128-dimensional high-dimensional feature mapping is first reduced to 32 dimensions using feature compression layer 2a (fully connected mapping) in feature compression module 2. The reduced features are then non-linearly activated by ReLU activation layer 2b, and finally further reduced to a three-dimensional output by linear output layer 3 (fully connected mapping), yielding the normalized 3D error prediction value, denoted as...

[0092]

[0093] , , These represent the normalized error prediction values ​​on the X, Y, and Z axes, respectively. To obtain the error prediction results at the actual physical quantity scale, it is necessary to perform inverse normalization using the output error statistics, restoring the 3D error prediction values ​​in the normalized space to physical quantities in mm. In other words, the inferred... Then reverse normalization is obtained For ease of subsequent explanation, Recorded as .

[0094] Step 4: Add the predicted 3D error vector and the predicted coordinates to obtain the compensated coordinates of the target point in the special aircraft coordinate system. : .

[0095] To verify the effectiveness of the method of this invention, corresponding experiments were conducted. The experimental platform consisted of a dedicated machine (a three-axis Cartesian coordinate system machine executing X, Y, and Z translation axes), a 3D surface structured light camera, an end-point calibration pin, and a checkerboard calibration plate. The dedicated machine has a maximum effective working range of 405×440×205mm, with an X, Y, and Z axis repeatability of 0.03mm and a maximum working speed of 500mm / s. A WiSight P180P area-array grating structured light 3D camera was also used, with a resolution of 230W, a sampling frequency ≥4 Hz, a ranging height of 180mm, an optical field of view of 110×70×80 mm, an absolute measurement accuracy of 0.018mm, and a repeatability of 0.005mm. The calibration plate used a 30×30mm grid, with a total of 10×12 checkerboard patterns, covering an effective XY working range of 300×360mm. Within the effective working space of hand-eye calibration, the true coordinates of corner points in the special-purpose machine coordinate system are obtained through manual teaching and pin-tip touching. The corresponding predicted coordinates are then calculated by combining the three-dimensional coordinates measured by the camera with the hand-eye calibration matrix, thereby constructing the original data required for error compensation modeling.

[0096] In sample construction, the predicted coordinates from hand-eye calibration output are used as network input, and the 3D error vector formed by the difference between the true coordinates and the predicted coordinates is used as the supervision output, ultimately constructing 256 input-output sample pairs. After the dataset is loaded, it is first randomly shuffled (random seed 42), and then divided into training, validation, and test sets at a ratio of 70% / 15% / 15%, resulting in 179, 38, and 39 sets respectively; the normalized statistics of input and output are calculated only from the training set.

[0097] The loss function for the training phase is the multi-output mean squared error (MSE). In the normalized space, the single-sample loss is defined as the component-wise average squared error:

[0098]

[0099] in the formula To iterate through the three components (x, y, z). The true label (normalized true value) of the i-th sample on the k-th component. The network prediction (normalized prediction) for the i-th sample on the k-th component. The loss for each sample is the average of the squared prediction errors of the three output components (in the normalized space).

[0100] The loss for mini-batch (size B) is:

[0101] .

[0102] The training parameters are: initial learning rate 0.01, minimum learning rate 1×10. -4 Batch size 8, maximum training epochs 1000. The validation set is used only for convergence monitoring and does not employ early stopping or optimal weight rollback.

[0103] The evaluation indicators and methods are explained below:

[0104] Let the true error vector of the i-th sample be... The prediction error vector output by the network is The compensated residual error vector is defined as follows: To characterize the magnitude of spatial bias in a single sample, the three-dimensional Euclidean error (spatial error) is defined as follows: In the model evaluation phase, the residual error vector and Euclidean error are first calculated, and then the model performance is quantitatively evaluated using indicators such as Euclidean RMSE, MAE, P95, P99, maximum error, and axis-by-axis RMSE.

[0105] The calculations for each indicator are as follows:

[0106] , , , , ,

[0107] Where N is the number of evaluation samples, and P95 and P99 represent respectively... The 95th and 99th percentiles.

[0108] Hand-eye calibration error statistical histogram as follows Figure 4 As shown, the spatial error before compensation is mainly concentrated in the range of 0.9–1.3 mm, with a relatively large overall amplitude (RMSE = 1.0949 mm), indicating that the system has a stable structural bias. The statistical histogram of residual error after error compensation is shown below. Figure 5 As shown, after error compensation, the residual error is mainly distributed in the range of 0.06 to 0.22 mm, with the highest error being about 0.27 mm. The overall RMSE is reduced to 0.1407 mm, which is 87.15% lower than before compensation.

[0109] Box plot comparing distribution before and after error compensation as shown in the figure Figure 6 As shown, the median error and interquartile range are significantly reduced after compensation, indicating that the model not only reduces the average error but also improves the consistency of the error distribution.

[0110] Heatmap of spatial error distribution in hand-eye calibration as shown Figure 7 As shown, the spatial error after hand-eye calibration exhibits significant regional differences within the workspace, indicating that the residual error possesses significant spatial correlation and non-uniformity. The heatmap of spatial error distribution after error compensation is shown below. Figure 8 As shown, it can be found that after compensation, the errors of most grid elements are concentrated in the range of 0.08 to 0.20 mm, the high error area shrinks significantly, and there are still relatively large residual errors only at a few boundary positions.

[0111] The following table compares the overall statistics before and after error compensation.

[0112]

[0113] The table shows that after error compensation, the overall RMSE, MAE, P95, P99 and Max all decreased significantly, and the spatial error was compressed from the order of about 1 mm to the order of 0.1 to 0.3 mm.

[0114] The following table compares the RMSE values ​​of each axis before and after error compensation.

[0115]

[0116] It can be seen that the error in the hand-eye calibration stage is mainly in the Z direction; after compensation, the RMSE in the X, Y and Z directions all decreased significantly, with the Z direction showing the most significant decrease.

[0117] To verify the effectiveness of the key network structures, two ablation experiments were designed, one involving the removal of residual connections and the other involving the removal of feature compression layers, under the same dataset, normalization strategy, training parameters, and loss function settings.

[0118] The following table compares the error compensation performance under different structural configurations.

[0119]

[0120] In the two sets of structural ablation experiments, the complete residual fully connected error compensation model performed best in terms of RMSE, MAE, P95, P99, maximum error, and triaxial RMSE, indicating that it can more effectively learn the nonlinear residual error distribution in the special aircraft workspace and achieve better compensation results in overall error control, tail error suppression, and triaxial error balance.

[0121] In contrast, removing residual connections significantly degraded model performance, with the overall RMSE increasing to 0.1696 mm, and the increase in RMSE in the Y direction being particularly pronounced. This indicates that residual connections help retain effective information in shallow layers and are a key structure for improving training stability and error compensation accuracy. While removing the feature compression layer improved model performance compared to the structure without residual connections, the RMSE increased to 0.1542 mm, and P95, P99, and maximum error all increased. This suggests that the lack of compact feature representation introduces feature redundancy, weakening generalization ability under small sample conditions.

[0122] In summary, the performance advantage of residual fully connected networks mainly comes from the synergistic effect of residual connections and feature compression layers. Residual connections are the core factor in improving error compensation capabilities, while feature compression layers further enhance the model's generalization and robustness.

[0123] This embodiment of the method, while preserving the geometric interpretability of the chessboard hand-eye calibration, introduces a residual fully connected network for secondary error compensation. Residual connections can alleviate the training degradation problem of deep networks and improve the ability to represent nonlinear features. Using the predicted coordinates obtained through hand-eye calibration as input, and the three-dimensional residual between the true and predicted coordinates as the supervision label, the method learns the mapping relationship between coordinate positions and error vectors. The predicted three-dimensional error vector is then superimposed onto the initial predicted coordinates to achieve secondary coordinate correction.

[0124] Another embodiment of the present invention is a special-purpose aircraft hand-eye calibration error compensation system based on residual fully connected networks, comprising:

[0125] The hand-eye calibration module is used to perform hand-eye calibration based on the corner points of the chessboard grid, establish an initial rigid body mapping from the camera coordinate system to the special aircraft coordinate system, and obtain the hand-eye calibration matrix.

[0126] The predicted coordinate calculation module is used to calculate the predicted coordinates of the target point in the dedicated machine coordinate system based on the target point's coordinates in the camera coordinate system, the position of the machine end in the dedicated machine coordinate system when sampling the target point, and the hand-eye calibration matrix.

[0127] The residual prediction module is used to input the predicted coordinates into the trained residual fully connected network to obtain the predicted three-dimensional error vector output by the residual fully connected network.

[0128] The error compensation module is used to add the predicted 3D error vector and the predicted coordinates to obtain the compensated coordinates of the target point in the special aircraft coordinate system.

[0129] The specific working process of each of the above modules is the same as that of the first embodiment, and will not be repeated here.

[0130] It is readily understood that the implementation of the methods in the above embodiments can be based on a computer program, which is a set of instructions that can be executed by a computer (i.e., by a processor). When the computer program is executed by the processor, it implements a dedicated machine hand-eye calibration error compensation method based on a residual fully connected network.

[0131] Furthermore, at least some of the computer programs associated with the methods of the embodiments can be distributed in a computer program product including a computer-readable storage medium carrying computer-usable instructions for one or more processors. This computer-readable storage medium can be provided in various forms, including non-transitory forms, such as, but not limited to, one or more disks, optical discs, magnetic tapes, chips, and magnetic and electronic storage. Further, the computer program can also be stored in a memory, which is part of an electronic device that also has a processor electrically connected to the memory. The computer program stored in the memory can be executed by the processor, thereby realizing a dedicated machine hand-eye calibration error compensation method based on a residual fully connected network.

Claims

1. A method for compensating hand-eye calibration error of a special-purpose aircraft based on a residual fully connected network, wherein the special-purpose aircraft is a translational motion mechanism, characterized in that... The method includes: Step 1: Perform hand-eye calibration based on the corner points of the chessboard grid, establish the initial rigid body mapping from the camera coordinate system to the special aircraft coordinate system, and obtain the hand-eye calibration matrix; Step 2: Calculate the predicted coordinates of the target point in the dedicated machine coordinate system based on the coordinates of the target point in the camera coordinate system, the position of the machine end in the dedicated machine coordinate system when the target point is sampled, and the hand-eye calibration matrix; Step 3: Input the predicted coordinates into the trained residual fully connected network to obtain the predicted 3D error vector output by the residual fully connected network; The residual fully connected network is used to output a corresponding predicted 3D error vector based on the input predicted coordinates. The residual fully connected network includes at least two feature extraction modules, feature compression modules, and linear output layers connected in sequence. The feature extraction module includes an input expansion layer and a multi-scale residual feature extraction layer connected in sequence. The feature compression module includes a feature compression layer and an activation layer connected in sequence. Step 4: Add the predicted 3D error vector and the predicted coordinates to obtain the compensated coordinates of the target point in the special aircraft coordinate system.

2. The method for compensating hand-eye calibration error of a special-purpose aircraft based on a residual fully connected network according to claim 1, characterized in that, The input extension layer is a fully connected structure used to map low-dimensional input features to a high-dimensional feature space. The input extension layers in different feature extraction modules map the input features from three dimensions to high dimensions step by step according to the connection order.

3. The method for compensating hand-eye calibration error of a special-purpose aircraft based on a residual fully connected network according to claim 2, characterized in that, The input expansion layer maps the input features from three dimensions to higher dimensions step by step, from three dimensions through 32 dimensions, 64 dimensions to 128 dimensions.

4. The method for compensating hand-eye calibration error of a special-purpose aircraft based on a residual fully connected network according to claim 1, characterized in that, The multi-scale residual feature extraction layer includes at least two cascaded residual blocks. Each residual block includes a main branch, an identity jump connection, a residual summing node, and an output activation layer. The main branch is used to process the input features of the residual block sequentially through a first fully connected layer, a first main branch activation layer, a second fully connected layer, and a second main branch activation layer to obtain the main branch output. The residual summing node adds the main branch output to the input features of the residual block passed through the identity jump connection element-wise. The output activation layer activates the output of the element-wise added features through an activation function.

5. The method for compensating hand-eye calibration error of a special-purpose aircraft based on a residual fully connected network according to claim 1, characterized in that, The feature compression layer is a fully connected structure used to reduce the dimensionality of the high-dimensional feature map obtained by the last feature extraction module. The dimensionality-reduced features are non-linearly activated by the activation layer and finally further reduced to three-dimensional output by the linear output layer.

6. The method for compensating hand-eye calibration error of a special-purpose aircraft based on a residual fully connected network according to claim 1, characterized in that, The trained residual fully connected network is obtained by training the network in the following way: using the predicted coordinates as training input and the three-dimensional error vector formed by the difference between the actual coordinates of the target point in the special aircraft coordinate system and the predicted coordinates as supervision labels, the residual fully connected network is trained under supervision, so that the residual fully connected network learns the nonlinear mapping relationship from the predicted coordinates to the three-dimensional error vector.

7. The method for compensating hand-eye calibration error of a special-purpose aircraft based on a residual fully connected network according to claim 6, characterized in that, When training the residual fully connected network, the dimensions of the training input and the dimensions of the supervision label are respectively subjected to Min-Max normalization. During the inference phase, the residual fully connected network outputs a normalized result, which is then denormalized to obtain the predicted three-dimensional error vector.

8. A special-purpose aircraft hand-eye calibration error compensation system based on a residual fully connected network, wherein the special-purpose aircraft is a translational motion mechanism, characterized in that, The system includes: The hand-eye calibration module is used to perform hand-eye calibration based on the corner points of the chessboard grid, establish an initial rigid body mapping from the camera coordinate system to the special aircraft coordinate system, and obtain the hand-eye calibration matrix. The predicted coordinate calculation module is used to calculate the predicted coordinates of the target point in the dedicated machine coordinate system based on the coordinates of the target point in the camera coordinate system, the position of the machine end in the dedicated machine coordinate system when the target point is sampled, and the hand-eye calibration matrix. The residual prediction module is used to input the predicted coordinates into a trained residual fully connected network to obtain the predicted three-dimensional error vector output by the residual fully connected network. The residual fully connected network is used to output a corresponding predicted 3D error vector based on the input predicted coordinates. The residual fully connected network includes at least two feature extraction modules, feature compression modules, and linear output layers connected in sequence. The feature extraction module includes an input expansion layer and a multi-scale residual feature extraction layer connected in sequence. The feature compression module includes a feature compression layer and an activation layer connected in sequence. The error compensation module is used to add the predicted three-dimensional error vector and the predicted coordinates to obtain the compensated coordinates of the target point in the special aircraft coordinate system.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the special-purpose machine hand-eye calibration error compensation method based on residual fully connected networks as described in any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the special-purpose machine hand-eye calibration error compensation method based on residual fully connected networks as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Hand-eye calibration error correction method based on ICP algorithm

    CN114519738A