Target pose measurement method fusing TOF camera intensity image and depth information
By fusing the intensity map and depth information of the TOF camera and using deep learning methods to measure the target posture, the problems of large calculation amount, long processing time and large cumulative error in the prior art are solved, and the pose measurement effect with high accuracy and strong robustness are achieved.
Patent Information
- Application Number
- CN202510104759.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
AI Technical Summary
When using a TOF camera to measure the target position, the calculation amount is large and the processing time is long, making it difficult to apply in real time, and the cumulative error is large, making it difficult to obtain high-precision position position information.
By fusing the TOF camera intensity map and depth information, network training is completed on the ground using deep learning methods and inference is performed on the orbit, which enhances the robustness of pose recognition and shortens the online detection time.
High-precision target position measurement is achieved, and the problem of traditional optical cameras being susceptible to complex lighting conditions is overcome, and the robustness and detection efficiency of the algorithm are improved.
Smart Images

Figure CN120031964A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of on-orbit services in the aerospace field, and in particular relates to a target posture measurement method integrating a TOF camera intensity map and depth information. Background Art
[0002] Space on-orbit operation missions are generally divided into four stages: measurement and planning, approach, capture or operation, and post-capture stabilization. High-precision measurement of the target's position and posture is a necessary condition for successful capture. Visual sensors are widely used in target recognition and position estimation due to their low cost, low power consumption, and lack of cooperative information. Since visible light cameras are sensitive to complex lighting conditions in space environments, TOF cameras have been used in recent years to measure target position and posture.
[0003] ToF (Time-of-Flight) camera is a new type of measurement payload that can output scene depth and intensity information at the same time. It has many excellent characteristics, such as being able to measure the depth value of each pixel of the array detector in parallel, obtaining accurate three-dimensional structural information of the object, and having the ability to suppress ambient light. The methods commonly used for posture measurement based on TOF cameras include: relative posture measurement method based on three-dimensional point cloud matching and relative posture measurement method based on feature matching. The former involves point cloud processing, which has a large amount of calculation and a long processing time, and is difficult to apply in real time on orbit; the measurement method based on feature matching usually uses the feature correspondence between two images to calculate the relative posture of the target between the previous and next two moments. Its disadvantage is that the cumulative error is relatively large, and compared with the relative posture between two moments, most applications require obtaining the posture information of the target relative to the tracking spacecraft.
[0004] Therefore the prior art needs to be improved. Summary of the invention
[0005] Purpose of the invention: In order to overcome the above shortcomings, the purpose of the present invention is to provide a target pose measurement method that integrates the TOF camera intensity map and depth information. The TOF camera intensity map and depth information are integrated through a deep learning method, network training is completed on the ground, and the trained network is inferred on orbit, which can enhance the robustness of pose recognition and shorten the online detection time.
[0006] Technical solution: In order to achieve the above-mentioned purpose, the present invention provides a method for measuring target pose by fusing TOF camera intensity map and depth information, comprising:
[0007] S1): Select a set of target key points in the target model, collect TOF camera images, and mark the selected target key points;
[0008] S2): Construct a training data set, use the TOF camera to image the target at different distances and different postures, and obtain the position of the target key points in the TOF camera intensity image and the corresponding depth information;
[0009] S3): construct a deep network for target key point recognition, add the two-dimensional target key point prediction error, the target key point 3D-2D projection error and the target key point 3D position error into the loss function, and use the training data set constructed in S2) for training;
[0010] S4): Obtain the position of the target key point, that is, obtain the image input target key point recognition deep network in real time, and predict the two-dimensional and three-dimensional position of the target key point;
[0011] S5): Calculate the relative position and posture parameters of the target, and calculate the relative position and posture parameters of the target according to the corresponding relationship between the two-dimensional position of the key point on the intensity map and the three-dimensional position in the target body coordinate system;
[0012] S6): Filter bad prediction results. Use the target relative pose parameters calculated in S5) to predict the 2D and 3D positions of the target key points, compare them with the target key point positions predicted in S4), and calculate the error; if the error is greater than the threshold, the result is eliminated; when the retained results exceed the smoothing quantity, calculate the mean and variance of the image projection errors in all retained sets, eliminate samples with errors greater than 3 times the variance, and obtain more accurate pose parameters.
[0013] In the target pose measurement method of the present invention that integrates the TOF camera intensity map and depth information, the two-dimensional target key point prediction error in S3) is defined as:
[0014]
[0015] in, represents the key point position predicted by the neural network, and (u, v) represents the actual position of the key point.
[0016] In the target pose measurement method of the present invention that integrates the TOF camera intensity map and depth information, the 3D-2D projection error of the target key point is:
[0017]
[0018] (u est , v est ) represents the estimated projection coordinates of the key points using the positions and postures calculated using the predicted key points;
[0019] Assuming that the projection position The target relative posture matrix is calculated as The relative position is The position of each predicted key point in the camera coordinate system is calculated using the following formula:
[0020]
[0021] in, is the coordinate of the i-th key point in the camera coordinate system, is the position of the i-th key point in the target body coordinate system; thus,
[0022] Calculate the predicted projection position of each key point on the camera plane:
[0023]
[0024] f is the focal length of the camera; (c x , c y ) is the coordinate of the camera principal point on the imaging plane.
[0025] In the target pose measurement method of the present invention that integrates the TOF camera intensity map and depth information, the 3D position error of the target key point is:
[0026]
[0027] in, Predict feature point locations for the network;
[0028] Therefore, the loss function can be determined as:
[0029] L=β 1 l uv +β 2 l proj +β 3 l 3d .
[0030] Among them, β 1 ,β 2 ,β 3 is the weight coefficient.
[0031] In the target pose measurement method of the present invention that integrates the TOF camera intensity map and depth information, when predicting the 2D and 3D positions of the target key points in S5), the following formula is used
[0032] The conditions for accepting the forecast result are:
[0033]
[0034] Among them, ε is the definition threshold, K is the proportional coefficient, and other symbols are defined as before.
[0035] The target pose measurement method for integrating the TOF camera intensity map and depth information described in the present invention, the target key point recognition deep network constructed in S3) is a U-shaped network, whose input is an intensity image containing the target, and the output is a target boundary box and two-dimensional projection position information and three-dimensional position information of each target key point. After the target key point recognition deep network is constructed, the data in the TOF camera intensity image collected in S1) is used to train the network until the loss function value converges to a stable state or the number of training times reaches a preset value; such training can be repeated many times until the loss function is small enough; the accuracy of testing the verification set reaches more than 90%.
[0036] It can be seen from the above technical solution that the present invention has the following beneficial effects:
[0037] 1. The target pose measurement method of the present invention that integrates the intensity map and depth information of a TOF camera utilizes the characteristic that the TOF camera is less affected by light, obtains the two-dimensional position of the key points of the target wheel frame through the intensity image, obtains the three-dimensional position of the key points through the depth image, and trains the fusion of the target intensity map and the depth information to obtain the spatial target position and pose. This method can well overcome the weakness of traditional optical cameras that are easily affected by complex lighting conditions in space and are difficult to stably obtain target features, and obtain stable feature information. By fusing the target depth information, the algorithm has strong robustness and can obtain high-precision pose estimation results.
[0038] 2. The target pose measurement method of the present invention that integrates the TOF camera intensity map and depth information completes network training on the ground through a deep learning method, and infers the trained network on track to obtain feature point position information, which can solve the problem of feature point shielding and greatly shorten the online detection time. It well solves the problems of instability, large amount of calculation and long time consumption of traditional target image feature point extraction.
[0039] 3. The method of eliminating bad estimation results by calculating the target projection error described in the present invention can ensure that the present invention obtains high-precision posture measurement results. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is a structural schematic diagram of the target pose measurement method integrating the TOF camera intensity map and depth information according to the present invention;
[0041] Figure 2 This is an example of comparing the intensity image of a TOF camera under strong light conditions with the RGB image of a common CCD camera in the present invention;
[0042] Figure 3 This is a diagram of the deep learning network structure in the present invention;
[0043] Figure 4A schematic diagram showing the target recognition box and key points predicted by the deep learning network of the present invention. DETAILED DESCRIPTION
[0044] The present invention is further explained below in conjunction with the accompanying drawings and specific embodiments.
[0045] Example
[0046] like Figures 1 to 4 A target pose measurement method for fusing a TOF camera intensity map and depth information as shown, comprising:
[0047] S1): Select a set of (8-17) target key points in the target model, collect TOF camera images, and mark the selected target key points;
[0048] S2): Construct a training data set, use the TOF camera to image the target at different distances and different postures, and obtain the position of the target key points in the TOF camera intensity image and the corresponding depth information;
[0049] S3): construct a deep network for target key point recognition, add the two-dimensional target key point prediction error, the target key point 3D-2D projection error and the target key point 3D position error into the loss function, and use the training data set constructed in S2) for training;
[0050] S4): Obtain the position of the target key point, that is, obtain the image input into the target key point recognition deep network in real time, and obtain the two-dimensional and three-dimensional position of the target key point;
[0051] S5): Calculate the relative position and posture parameters of the target, and calculate the relative position and posture parameters of the target according to the corresponding relationship between the two-dimensional position of the key point on the intensity map and the three-dimensional position in the target body coordinate system;
[0052] S6): Filter bad prediction results. Use the target relative pose parameters calculated in S5) to predict the 2D and 3D positions of the target key points, compare them with the target key point positions predicted in S4), and calculate the error; if the error is greater than the threshold, the result is eliminated; when the retained results exceed the smoothing quantity, calculate the mean and variance of the image projection errors in all retained sets, eliminate samples with errors greater than 3 times the variance, and obtain more accurate pose parameters.
[0053] In the process of selecting target key points, the corner points of the target are first considered, and they are distributed on different faces as much as possible. It should be noted that when selecting target key points, different requirements require different characteristics of the target points. In actual application, appropriate target points can be selected according to actual needs.
[0054] The characteristic of TOF intensity graph is that it is not affected by light. Whether it is backlit or frontlit, or the light is shining from the front or side, the intensity graph can maintain stable characteristics, such as Figure 2 shown.
[0055] In the target pose measurement method of integrating TOF camera intensity map and depth information described in this embodiment, the two-dimensional target key point prediction error in S3) is defined as:
[0056]
[0057] in, represents the key point position predicted by the neural network, and (u, v) represents the actual position of the key point.
[0058] In the target pose measurement method of the embodiment described in the fusion TOF camera intensity map and depth information, the 3D-2D projection error of the target key point is:
[0059]
[0060] (u est , v est ) represents the estimated projection coordinates of the key points using the positions and postures calculated using the predicted key points;
[0061] Assuming that the projection position The target relative posture matrix is calculated as The relative position is The position of each predicted key point in the camera coordinate system is calculated using the following formula:
[0062]
[0063] in, is the coordinate of the i-th key point in the camera coordinate system, is the position of the i-th key point in the target body coordinate system; from this, the predicted projection position of each key point on the camera plane can be calculated:
[0064]
[0065] f is the focal length of the camera; (c x , c y ) is the coordinate of the camera principal point on the imaging plane.
[0066] In the target pose measurement method of the embodiment described in the fusion TOF camera intensity map and depth information, the 3D position error is:
[0067]
[0068] in, Predict feature point locations for the network;
[0069] Therefore, the loss function can be determined as:
[0070] L=β 1 l uv +β 2 l proj +β 3 l 3d .
[0071] Among them, β 1 ,β 2 ,β 3 is the weight coefficient.
[0072] In the target pose measurement method of integrating the TOF camera intensity map and depth information described in this embodiment, when predicting the 2D and 3D positions of the target key points in S5), the following formula is used as a condition for accepting the prediction result:
[0073]
[0074] Among them, ε is the definition threshold, K is the proportional coefficient, and other symbols are defined as before.
[0075] The target pose measurement method of integrating the TOF camera intensity map and depth information described in this embodiment, the target key point recognition deep network constructed in S3) is a U-shaped network, whose input is an intensity image containing the target, and the output is a target bounding box and two-dimensional projection position information and three-dimensional position information of each target key point, and the target key point recognition deep network is constructed as follows Figure 3 As shown in the figure, the number of target key points is 8-17. After the target key point recognition deep network is built, the network is trained using the data in the TOF camera intensity image collected in S1) until the loss function value converges to a stable state or the number of training times reaches a preset value; such training can be repeated many times until the loss function is small enough; the accuracy of the test on the verification set reaches more than 90%.
[0076] The above description is only a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements without departing from the principle of the present invention, and these improvements should also be regarded as within the protection scope of the present invention.
Claims
1. A method for measuring target pose by fusing TOF camera intensity map and depth information, characterized in that: include: S1): Select a set of target key points in the target model, collect TOF camera images, and mark the selected target key points; S2): Construct a training data set, use the TOF camera to image the target at different distances and different postures, and obtain the position of the target key points in the TOF camera intensity image and the corresponding depth information; S3): construct a deep network for target key point recognition, add the two-dimensional target key point prediction error, the target key point 3D-2D projection error and the target key point 3D position error into the loss function, and use the training data set constructed in S2) for training; S4): Obtain the position of the target key point, that is, obtain the TOF camera image in real time, input the target key point recognition deep network, and predict the two-dimensional and three-dimensional position of the target key point; S5): Calculate the relative position and posture parameters of the target, and calculate the relative position and posture parameters of the target according to the corresponding relationship between the two-dimensional position of the key point on the intensity map and the three-dimensional position in the target body coordinate system; S6): filtering bad prediction results, i.e. using the target relative posture parameters calculated in S5), calculating the 2D and 3D positions of the target key points, comparing them with the target key point positions predicted in S4), and calculating the error; if the error is greater than the threshold, the result is eliminated; When the retained results exceed the smoothing quantity, the mean and variance of the image projection errors in all retained sets are calculated, and samples with errors greater than 3 times the variance are eliminated to obtain more accurate pose parameters.
2. The target pose measurement method according to claim 1, wherein: The two-dimensional target key point prediction error in S3) is defined as: in, represents the key point position predicted by the neural network, and (u, v) represents the actual position of the key point.
3. The target pose measurement method according to claim 1, wherein: The target key point 3D-2D projection error is: (u est , v est ) represents the estimated projection coordinates of the key points using the positions and postures calculated using the predicted key points; Assuming that the projection position The target relative posture matrix is calculated as The relative position is The position of each predicted key point in the camera coordinate system is calculated using the following formula: in, is the coordinate of the i-th key point in the camera coordinate system, is the position of the i-th key point in the target body coordinate system; from this, the predicted projection position of each key point on the camera plane can be calculated: f is the focal length of the camera; (c x , c y ) is the coordinate of the camera principal point on the imaging plane.
4. The target pose measurement method according to claim 1, wherein: The 3D position error of the target key point is: in, Predict feature point locations for the network; Therefore, the loss function can be determined as: L=β1l uv +β2l proj +β3l 3d 。 Among them, β1, β2, β3 are weight coefficients.
5. The target pose measurement method of fusing TOF camera intensity map and depth information according to claim 1, characterized in that: When predicting the 2D and 3D positions of the target key points in S6), the following formula is used as a condition for accepting the prediction result: Among them, ε is the definition threshold, K is the proportional coefficient, and other symbols are defined as before.
6. The target pose measurement method of fusing TOF camera intensity map and depth information according to claim 1, characterized in that: The target key point recognition deep network constructed in S3) is a U-shaped network, whose input is an intensity image containing the target, and whose output is a target bounding box and two-dimensional projection position information and three-dimensional position information of each target key point. After the target key point recognition deep network is constructed, the network is trained using data from the TOF camera intensity image collected in S1) until the loss function value converges to a stable state or the number of training times reaches a preset value; such training can be repeated many times until the loss function is small enough; the accuracy of testing the verification set reaches more than 90%.