A Deep Learning-Based Spacecraft Pose Estimation Method
By using normal distribution initialization parameters and normalized preprocessing in spacecraft posture estimation, combined with the minimizing geometric residual optimization algorithm, the problem of large error in posture parameter estimation in the existing technology is solved, and pose estimation with high accuracy, low complexity and real-timeness is achieved.
Patent Information
- Application Number
- CN202111517636.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2041-12-13
AI Technical Summary
The existing deep learning-based spacecraft posture estimation method is prone to large errors in rotation parameter estimation due to the order of magnitude difference between translation parameters and rotation parameters during training, and the network is too large, the running time is long, and the hardware requirements are high.
The distribution differences between translation and rotation parameters are reduced by using normal distribution initialization parameters in neural network models and normalized preprocessing of training images and pose labels. At the same time, minimize the geometric residuals of the projection profile of the spacecraft CAD model in the image to optimize the posture parameters and obtain the maximum likelihood estimate of the spacecraft posture parameters.
It improves the estimation accuracy of pose parameters, reduces the algorithm run time and hardware requirements, and ensures high accuracy and real-time pose estimation.
Smart Images

Figure CN114419149B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a method for estimating the pose of a spacecraft based on deep learning. Background Art
[0002] With the development of artificial intelligence technology, deep learning has attracted much attention due to its excellent performance in tasks such as image classification, detection, and segmentation. Deep learning algorithms extract target features in images based on the powerful feature extraction ability of convolutional networks, and then use fully connected networks to achieve classification output. The fitting training based on a large amount of data enables the entire deep learning model to have a high classification accuracy in actual application scenarios. Given the powerful classification and recognition effect of deep learning, many scholars have begun to study the application of deep learning to the pose estimation of non-cooperative spacecraft.
[0003] The methods for estimating the pose of non-cooperative spacecraft based on deep learning are mainly divided into two categories according to the implementation principle: pose estimation algorithms based on classification and pose estimation algorithms based on regression. The pose estimation algorithms based on classification discretize the pose space into different categories. The deep learning network takes the spacecraft image as the input, and maps the input image to different categories. The selected category is the pose parameter of the spacecraft in the input image at this moment. The classification accuracy of this method is relatively high. For example, the accuracy of the pose estimation algorithm proposed by Sharma et al. can reach more than 95%. However, the pose estimation accuracy of this type of method depends severely on the number of categories. Too few categories result in low algorithm accuracy, while too many categories make the training of the deep learning network extremely difficult. The pose estimation algorithms based on regression establish a mapping relationship between the image space and the pose space through a deep learning network. The network takes the spacecraft image as the input data, and then directly outputs the pose parameters of the spacecraft in the image through the mapping relationship. This type of method avoids discretizing the pose space, and thus is favored by more and more scholars. However, the related technologies mainly have the following disadvantages: 1) Because of the large difference in the order of magnitude between the translation parameters and the rotation parameters, the optimization direction of the neural network during training tends to fit the translation parameters and ignore the rotation parameters, ultimately resulting in a large estimation error of the rotation parameters of the algorithm; 2) The network scale is too large, the algorithm running time is long, and there are high requirements for hardware. Summary of the Invention
[0004] To solve the problems in the prior art, the present invention proposes a deep learning-based spacecraft pose estimation method, which can effectively solve the difference in the order of magnitude between the translation parameter and the rotation parameter in the pose parameters, thereby ensuring that the position parameter and the rotation parameter predicted by the deep learning network have high accuracy. Moreover, by minimizing the geometric residual of the projection contour of the spacecraft CAD model in the image, the pose parameters obtained by the deep learning network are optimized, so as to give a maximum likelihood estimation of the spacecraft pose parameters, which has very excellent statistical characteristics, thus ensuring the accuracy of pose estimation. The network scale of this algorithm is small, the algorithm running time is short, and the hardware requirements are low.
[0005] To achieve the above object, the present invention provides a deep learning-based spacecraft pose estimation method, including: first, randomly initialize the neural network model parameters according to the normal distribution, then normalize the training images and the corresponding pose labels in the dataset respectively and use them to train the neural network model, then input the test images into the trained neural network model, use the inverse normalization to obtain the pose predicted by the neural network model, and finally, according to the pose parameters predicted by the neural network model, use the geometric residual of minimizing the projection contour of the spacecraft CAD model in the image to optimize the pose parameters of the spacecraft, and finally obtain the maximum likelihood estimation of the spacecraft pose parameters.
[0006] Further, the normalization of the training images includes:
[0007] First, create a 7×7 grid called the Gaussian kernel W, and initialize the weight values of the grid with a two-dimensional normal distribution based on the grid center:
[0008]
[0009] where w(x,y) is the weight value of each grid in the Gaussian kernel, x,y are the positions of each grid relative to the center grid, and σ is a constant, set to 2;
[0010] Then normalize the Gaussian kernel W so that the sum of the weight values of all grids is 1:
[0011]
[0012] Next, calculate the weighted average value m and the weighted standard deviation of each pixel point in the image
[0013]
[0014] where p is each pixel value in the original image;
[0015] Finally, calculate the pixel value of each pixel point after normalization:
[0016]
[0017] Further, the normalization and denormalization of the pose label include: The parameters of the pose label include translation parameters and rotation parameters. The rotation parameters are represented by unit quaternions. The three translation parameters of the i-th image are represented by x i , y i , z i respectively. During the training stage, the translation parameters in the pose label corresponding to each training image are normalized:
[0018]
[0019] where are the mean values corresponding to x i , y i , z i respectively in the training label set, and σ x , σ y , σ z are the corresponding standard deviations respectively, and are the corresponding output parameters after normalization;
[0020] During the testing stage, the normalized prediction parameters are denormalized to obtain the finally predicted translation parameters:
[0021]
[0022] where x out , y out , z out are the normalized prediction parameters respectively, and is the predicted translation parameter after denormalization.
[0023] Further, the neural network model includes an encoder and a decoder. The encoder adopts a 34-layer residual network layer framework and uses an adaptive average pooling layer to replace the fully connected layer in the residual network. The decoder includes a translation decoder for predicting the translation vector and a rotation decoder for predicting the rotation quaternion. The translation decoder includes two 1×2048 fully connected layers with dropout and finally connects to a 1×3 output vector. The rotation decoder includes two 1×2048 fully connected layers with dropout and connects to a 1×4 fully connected layer. Finally, the output of the 1×4 fully connected layer is connected to a normalization layer so that the output is a 1×4 vector with a unit length of 1 as the predicted value of the rotation quaternion.
[0024] Further, the settings of the neural network model include:
[0025] a) The neural network parameters are initialized using a normal distribution, and the ReLU function is selected as the activation function;
[0026] b) The optimizer uses the Adam optimizer, and the learning rate is adjusted to 0.001. The first-order decay rate β 1 and the second-order decay rate β 2 of the optimizer are set to 0.9 and 0.009 respectively;
[0027] c) The mini-batch size for training is set to 8;
[0028] d) Dropout is set to 0.5.
[0029] Furthermore, in the training stage, the loss functions of the translation parameters and the rotation parameters are minimized simultaneously:
[0030]
[0031] where p t and p pre are the true value and the predicted value of the translation parameters respectively, and q t and q pre are the true value and the predicted value of the rotation parameters respectively.
[0032] Furthermore, in the algorithm, a perspective model of a vacuum camera is adopted. Let the focal length of the camera be f, the pixel aspect ratio be k, the tilt factor be s, and the principal points of the camera be u and v. Then the internal parameter matrix of the camera is defined as:
[0033]
[0034] Furthermore, in the algorithm, a target pose model is adopted. The target pose model uses a rotation matrix R and a translation vector t to describe the pose of the target relative to the camera coordinate system. Assume that the coordinates of a point in space in the camera coordinate system are X. Then the corresponding image point x of X in the image is:
[0035]
[0036] where and are the homogeneous coordinates of x and X respectively, and the symbol means equal up to a scale factor;
[0037] The rotation matrix adopts a representation method of the rotation matrix based on Lie groups and Lie algebras. Let the rotation vector be Ω. Then the rotation matrix is represented as:
[0038]
[0039] where I is the unit vector, and [Ω] x is the skew-symmetric matrix composed of the components of Ω:
[0040]
[0041] When the rotation matrix R is given, the rotation vector is calculated from R:
[0042] ||Ω|| = arccos((trace(R) - 1) / 2)
[0043] [Ω] x = ||Ω|| / (2sin||Ω||)(R - R T )
[0044] The pose of the target is uniquely determined by six parameters, p = [Ω T t T T , and p is called the pose parameter vector.
[0045] Furthermore, the maximum likelihood estimation of the spacecraft pose parameters based on geometric residual minimization in the algorithm includes:
[0046] a) Control point selection: Select several control points on the projected contour of the spacecraft CAD model, and then use these control points to find useful image points from the spacecraft image. The projected contour of the spacecraft CAD model on the image plane is a closed polygon, which is called c m , let S be an edge of the CAD model, and its two endpoints be P 1 and P 2 . Given the pose vector p, obtaining the rotation matrix R and the translation vector t, then the homogeneous coordinates of the projected points of P1 and P2 on the image plane are:
[0047]
[0048]
[0049] The image of S is a straight line segment with g 1 and g 2 as endpoints. The straight line equation passing through g 1 and g 2 is:
[0050]
[0051] Calculate the midpoint of the image segment corresponding to each edge of the CAD model. For each midpoint, determine whether the back-projection ray of the midpoint has one and only one intersection with the CAD model. If so, the midpoint must be located on the contour c m ; regard the midpoint as a control point; if not, discard the midpoint so that all control points are located on different edges of the contour polygon c m ;
[0052] b) Matching between the projection contour of the CAD model and the image pixels: At the current moment, read an image of the spacecraft, use the Canny operator method to obtain the edge image of the image, and use the control points to find the image edge points belonging to the projection contour c of the CAD model from the edge image. m For each control point, search for image edge points within a certain range on the normal line of the control point. If an edge point satisfies the following two conditions within the search interval of a certain control point, it is considered that the edge point belongs to the projection contour c of the spacecraft CAD model. m : (I) The angle between the normal line of the edge point and the normal line of the given control point is less than a certain threshold; (II) The edge point is the edge point closest to the given control point among all edge points that satisfy (I).
[0053] c) Pose parameter optimization: Let {q i , i = 1, 2,..., n} be the edge points found to belong to c m . The control point corresponding to each q i is located on a certain side of the polygon c m . Denote this straight line as l i . l i is a function vector of the pose parameter p. For the correct pose parameter p, q i is located on the straight line l i . Obtain the pose parameter p of the spacecraft by minimizing the geometric distance between {q i} and the corresponding straight line {l i}:
[0054]
[0055] where, U i = q i q i T , A = diag{1 1 0}, d i represents the geometric distance between q i and l i ;
[0056] Use the LM method to optimize the non - linear least - squares problem. At the s - th iteration, let the estimated value of p be p(s), and define the residual vector as: d (s) = [d 1 (s) … d n (s) T , then the update process of p is as follows:
[0057] p (s+1) = p (s) - (J (s)T J(s) + ζI n ) -1 J (s)T d (s)
[0058] where ζ is a non - negative factor that is updated as the optimization process proceeds; I n is an n - order identity matrix; J (s) is the Jacobian matrix of J (s) and is calculated as follows:
[0059]
[0060] where is calculated as follows:
[0061]
[0062] where is obtained by taking the partial derivative of ;
[0063] d) Optimization of spacecraft attitude parameters: Using the spacecraft attitude parameters predicted by the deep - learning neural network model as the initial values, successively perform a) control point selection, b) matching of the CAD model projection contour and image pixels, and c) attitude parameter optimization to obtain the estimated values of the spacecraft attitude parameters. Then, execute the above process multiple times, with the attitude estimation of the result of the previous execution as the initial value each time. After multiple executions, the estimated result of the spacecraft attitude parameters obtained is the final estimated value of the spacecraft attitude parameters.
[0064] Compared with the prior art, the present invention first randomly initializes the neural network model parameters according to the normal distribution, then preprocesses the training images and the corresponding pose labels in the dataset by normalization respectively for training the neural network model. Then, the test images are input into the trained model, and the poses predicted by the model are obtained by denormalization. Finally, according to the pose parameters predicted by the model, the geometric residuals of the projection contour of the spacecraft CAD model in the image are minimized to optimize the pose parameters of the spacecraft, and finally the maximum likelihood estimation of the spacecraft pose parameters is obtained. By preprocessing the pose parameters through the neural network model for estimating the spacecraft pose, the present invention makes the distribution ranges of the translation parameters and the rotation parameters relatively close. Therefore, the fitting of the translation and rotation parameters can tend to be balanced during the training process, thereby reducing the translation and rotation errors of the pose estimation algorithm. The floating-point operation count (FLOPs) of the deep learning network proposed by the present invention is 22.5G, and the prediction time for a single image is about 7ms. The pose calculation time is short, which can meet the real-time requirements. Based on the pose prediction values provided by the deep learning network, the present invention optimizes the pose parameters of the spacecraft by minimizing the geometric residuals of the projection contour of the spacecraft CAD model in the image, and finally provides the maximum likelihood estimation of the spacecraft pose parameters. Compared with the prior art, to solve the balance problem between the translation loss and the rotation loss, a hyperparameter is usually added to the loss function to make the losses of the two parameters tend to be balanced. The disadvantage of this approach is that when the distribution of the parameters changes, the value of the hyperparameter needs to be readjusted. Therefore, a large amount of time and computing resources are required to adjust the parameters to reach a relatively ideal approximation value, so as to make the prediction results more accurate. The label normalization preprocessing adopted by the present invention can constrain the translation parameters in the same distribution. Therefore, this hyperparameter can be set to a constant value, making the training of the neural network model more efficient and obtaining more accurate prediction values. The local contrast normalization is used to preprocess the images to highlight the contours in the images, enabling the neural network model to learn richer contour information and making the prediction results more accurate. The geometric residual algorithm for minimizing the projection contour of the spacecraft CAD model in the image is used to optimize the predicted poses to obtain more accurate pose parameters. It has very excellent statistical characteristics, thus ensuring the accuracy of the pose estimation. The attitude angle error of the algorithm is about 2.68 degrees, and the relative translation error is about 4%. Moreover, the network scale is small, the algorithm running time is short, and the hardware requirements are low. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 is the algorithm flowchart of the present invention;
[0066] Figure 2The effect diagrams of the algorithm of the present invention for processing three different CAD models are shown. Among them, the three rows are sample frames of three different CAD models respectively. The green contour is the projection contour corresponding to the pose predicted by the neural network model, and the red contour is the projection contour corresponding to the optimized pose.
[0067] Figure 3 The display diagram of the image preprocessing by the algorithm of the present invention is shown. Among them, the image in the first row is the original image in the dataset, and the image in the second row is the image after preprocessing by local contrast normalization.
[0068] Figure 4 The neural network model diagram of the present invention. Specific implementation manners
[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0070] The present invention provides a deep learning-based spacecraft pose estimation method, which is a six-degree-of-freedom relative pose estimation method for spacecraft based on deep learning. Refer to Figure 1 , and the algorithm workflow includes: First, randomly initialize the neural network model parameters according to the normal distribution. Then, normalize and preprocess the training images and the corresponding pose labels in the dataset for training the neural network model. Then, input the test images into the trained neural network model, and use the inverse normalization to obtain the poses predicted by the model. Finally, according to the pose parameters predicted by the neural network model, use the geometric residuals that minimize the projection contour of the spacecraft CAD model in the image to optimize the pose parameters of the spacecraft, and finally obtain the maximum likelihood estimation of the spacecraft pose parameters.
[0071] The pose predicted by the neural network model of the present invention and the optimized pose effects are as Figure 2 shown. The effect diagrams of the algorithm of the present invention for processing three different CAD models are shown. Among them, the three rows are sample frames of three different CAD models respectively. The green contour is the projection contour corresponding to the pose predicted by the neural network model, and the red contour is the projection contour corresponding to the optimized pose. From Figure 2It can be seen from this that the present invention preprocesses the pose parameters, making the distribution ranges of the translation parameters and rotation parameters relatively close. Therefore, during the training process, the fitting of the translation and rotation parameters can tend to be balanced, thereby reducing the translation and rotation errors of the pose estimation algorithm. The floating-point operation count (FLOPs) of the deep learning network proposed by the present invention is 22.5G, and the prediction time for a single image is about 7ms. The pose calculation time is short, which can meet the real-time requirements. Based on the pose prediction values provided by the deep learning network, the present invention optimizes the spacecraft pose parameters by minimizing the geometric residuals of the projected contour of the spacecraft CAD model in the image, and finally provides a maximum likelihood estimate of the spacecraft pose parameters. The data preprocessing algorithm of label normalization can effectively solve the difference in the order of magnitude between the translation parameters and rotation parameters in the pose parameters, thereby ensuring high precision of the position parameters and rotation parameters predicted by the deep learning network. In addition, this method optimizes the pose parameters obtained by the deep learning network by minimizing the geometric residuals of the projected contour of the spacecraft CAD model in the image, thereby giving a maximum likelihood estimate of the spacecraft pose parameters. This makes this method have very excellent statistical characteristics, thus ensuring the accuracy of pose estimation, with a small network scale, short algorithm running time, and low hardware requirements.
[0072] Specifically, the present invention includes:
[0073] 1) Image and label normalization:
[0074] 1a) Image normalization: Before each image is input into the neural network, local contrast normalization is performed on the image. However, local contrast normalization ignores regions with similar intensities. The present invention makes the edge features of the image more prominent. Specifically, the effect is as Figure 3 shown. The data preprocessing algorithm can effectively solve the difference in the order of magnitude between the translation parameters and rotation parameters in the pose parameters, thereby ensuring high precision of the position parameters and rotation parameters predicted by the deep learning network. Specifically, it includes:
[0075] First, create a 7×7 grid called the Gaussian kernel W. Based on the grid center, initialize the weight values of the grid with a two-dimensional normal distribution, as shown in formula (1):
[0076]
[0077] where w(x,y) is the weight value of each grid in the Gaussian kernel, x,y are the positions of each grid relative to the central grid, and σ is a constant, set to 2.
[0078] Next, normalize the Gaussian kernel W through formula (2) so that the sum of the weight values of all grids is 1:
[0079]
[0080] Then calculate the weighted average value m and the weighted standard deviation of each pixel point in the image where p is each pixel value in the original image, as shown in formulas (3) and (4):
[0081]
[0082] Finally, use formula (5) to calculate the pixel value of each pixel point after normalization:
[0083]
[0084] 1b) Label normalization and denormalization:
[0085] In the dataset, each image corresponds to a label, and its parameters consist of two parts, namely the translation parameter and the rotation parameter. The rotation parameter is represented by a unit quaternion, so no normalization is required. Denote the three translation parameters of the i-th image as x i , y i , z i . During the training stage, normalize the translation parameters in the label corresponding to each training image according to formula (6):
[0086]
[0087] where are the average values of x i , y i , z i corresponding to the training label set respectively, σ x , σ y , σ z are the corresponding standard deviations respectively, are the corresponding output parameters after normalization.
[0088] During the testing stage, when the test image is used as the input of the neural network model, the predicted parameters output consist of two parts, namely the unit quaternion parameter and the normalized translation parameter. Therefore, it is necessary to denormalize the normalized predicted parameters as follows to obtain the finally predicted translation parameter, as shown in formula (7):
[0089]
[0090] where x out , y out , z out are the normalized predicted parameters respectively, is the predicted translation parameter after denormalization.
[0091] 2) Construction of the neural network model
[0092] The neural network model architecture of the present invention is as shown in Figure 4 . It is designed by the Pytorch open-source framework. The entire model consists of two parts: an encoder and a decoder. The encoder adopts the framework of a 34-layer residual network (ResNet34), and replaces the last fully connected layer in the residual network with an adaptive equalization layer. The decoder consists of two parts: a translation decoder for predicting the translation vector and a rotation decoder for predicting the rotation quaternion. The translation decoder is composed of two fully connected layers of 1×2048 with dropout, and finally connects to a 1×3 output vector. The rotation decoder is composed of two fully connected layers of 1×2048 with dropout, connects to a 1×4 fully connected layer, and finally connects the output of the 1×4 fully connected layer to a normalization layer so that its output is a 1×4 vector with a unit length of 1 as the predicted value of the rotation quaternion. The settings of the entire neural network model are as follows:
[0093] 2a) The neural network parameters are initialized using a normal distribution, and the ReLU function is selected as the activation function;
[0094] 2b) The Adam optimizer is used, the learning rate is adjusted to 0.001, and the first-order decay rate β 1 and the second-order decay rate β 2 of the optimizer are set to 0.9 and 0.009 respectively;
[0095] 2c) The mini-batch size for training is set to 8;
[0096] 2d) Dropout is set to 0.5.
[0097] 3) Loss function:
[0098] In the training stage, the loss functions of the translation parameters and the rotation parameters are minimized simultaneously using formula (8):
[0099]
[0100] where p t and p pre are the true value and the predicted value of the translation parameter respectively, and q t and q pre are the true value and the predicted value of the rotation parameter respectively.
[0101] 4) Construct a camera model and a target pose model:
[0102] 4a) Camera model:
[0103] The present invention adopts a perspective model of a vacuum camera. Let the focal length of the camera be f, the pixel aspect ratio be k, the tilt factor be s, and the principal points of the camera be u and v. Then, the internal parameter matrix of the camera is defined as:
[0104]
[0105] 4b) Pose model of the target:
[0106] The pose model of the target in the present invention uses a rotation matrix R and a translation vector t to describe the pose of the target relative to the camera coordinate system. Assume that the coordinates of a point in space in the camera coordinate system are X. Then, the corresponding image point x of X in the image is:
[0107]
[0108] where and are the homogeneous coordinates of x and X respectively. The symbol means equality up to a scale constant;
[0109] The rotation matrix has different representation forms. To represent the rotation of the target unambiguously, the present invention adopts a rotation matrix representation method based on Lie groups and Lie algebras. Let the rotation vector be Ω. Then, the rotation matrix is represented as:
[0110]
[0111] where I is the identity vector, and [Ω] x is the skew-symmetric matrix formed by the components of Ω:
[0112]
[0113] It can be seen from formula (11) that when the rotation matrix R is given, the rotation vector is calculated through R:
[0114] ||Ω|| = arccos((trace(R) - 1) / 2)
[0115] [Ω] x = ||Ω|| / (2sin||Ω||)(R - R T ) (13)
[0116] Therefore, the pose of the target is uniquely determined by six parameters, p = [Ω T t T T , and p is called the pose parameter vector.
[0117] 5) Estimation of spacecraft pose parameters based on geometric residual minimization:
[0118] 5a) Selection of control points:
[0119] To optimize the attitude and position parameters of the spacecraft, first, several control points need to be selected on the projection contour of the spacecraft CAD model, and then, using these control points, useful image points are found from the spacecraft images. The CAD model of the spacecraft is a polyhedron constructed by triangular patches. Therefore, the projection contour of the spacecraft CAD model on the image plane is a closed polygon, which is called c m Let S be an edge of the CAD model, and its two endpoints are P 1 and P 2 . Given the attitude vector p, the rotation matrix R and the translation vector t are obtained through formula (11). Then, the homogeneous coordinates of the projection points of P1 and P2 on the image plane are:
[0120]
[0121] Therefore, the image of S is a straight-line segment with g 1 and g 2 as endpoints. The straight-line equation passing through g 1 and g 2 is:
[0122]
[0123] Use formula (14) to calculate the midpoint of the image line segment corresponding to each edge of the CAD model. For each midpoint, judge whether the back-projection ray of the midpoint has one and only one intersection with the CAD model. If so, the midpoint must be located on the contour c m and regard this midpoint as a control point. If not, discard this midpoint. From the above process, it can be seen that all control points are located on different edges of the contour polygon c m .
[0124] 5b) Matching between the projection contour of the CAD model and image pixels:
[0125] The algorithm reads in an image of the spacecraft at the current moment and uses the Canny operator method to obtain the edge image of the image. To accurately estimate the attitude and position of the spacecraft at the current moment, it is necessary to use the control points to find the image edge points belonging to the projection contour c m of the CAD model from the edge image. For this purpose, centered on each control point, image edge points are searched within a certain range on the normal line of the control point. This search range is called the search interval. The normal vector of the control point can be calculated by formula (15). Within the search interval of a certain control point, if an edge point satisfies the following two conditions, it is considered that the edge point belongs to the projection contour c m: (I) The angle between the normal of the edge point and the normal of the given control point is less than a certain threshold (recommended value 15 degrees); (II) The edge point is the edge point closest to the given control point among all edge points satisfying (I).
[0126] 5c) Pose parameter optimization:
[0127] Let {q i , i = 1, 2,..., n} be the edge points found in 5b) that belong to c m . From 5a), it can be seen that for each q i , the corresponding control point is located on a certain side of the polygon c m . The straight-line equation of this side can be calculated by formula (15), and this straight line is denoted as l i . From formula (15), it can be seen that l i is a function vector of the pose parameter p. For the correct pose parameter p, q i is located on the straight line l i . The pose parameter p of the spacecraft is obtained by minimizing the geometric distance between {q i} and the corresponding straight line {l i}:
[0128]
[0129] where, U i = q i q i T , A = diag{1 1 0}, d i represents the geometric distance between q i and l i ; formula (16) is a non-linear least squares problem, and the present invention uses the LM algorithm to optimize the non-linear least squares problem. At the s-th iteration, let the estimated value of p be p(s), and the residual vector is defined as: d (s) = [d 1 (s) … d n (s) T , then the update process of p is as follows:
[0130] p (s+1) = p (s) - (J (s)T J (s) + ζI n ) -1 J (s)T d (s) (17)
[0131] where, ζ is a non-negative factor, which is updated as the optimization process progresses and plays a role in increasing the robustness of the algorithm. In is the n - order identity matrix; J (s) is the Jacobian matrix of J (s) and is calculated as follows:
[0132]
[0133] wherein, can be calculated by the following formula:
[0134]
[0135] wherein, in formula (19), is obtained by taking the partial derivative of formula (15).
[0136] 5d) Optimization of spacecraft pose parameters:
[0137] Taking the spacecraft pose parameters predicted by the deep learning network as the initial values, sequentially execute 5a) control point selection, 5b) image point matching, and 5c) pose parameter optimization, so as to obtain the estimated values of the spacecraft pose parameters. Next, execute the above process 6 times, and each execution uses the pose estimation of the previous execution result as the initial value. After the sixth execution, the spacecraft pose parameter estimation result given by the algorithm is the final estimated value of the spacecraft pose parameters.
[0138] The pseudo - code of the algorithm is as follows:
[0139]
[0140] In the prior art, to solve the balance problem between translation loss and rotation loss, a hyper - parameter is usually added to the loss function to make the losses of the two parameters tend to be balanced. The disadvantage of this approach is that when the distribution of the parameters changes, the value of the hyper - parameter needs to be readjusted. Therefore, a large amount of time and computing resources are required to adjust the parameters to reach a more ideal approximate value, so as to make the prediction result more accurate. The label normalization pre - processing adopted in the present invention can constrain the translation parameters in the same distribution. Therefore, this hyper - parameter can be set as a constant value, making the training of the neural network model more efficient and obtaining more accurate prediction values. Using local contrast normalization to pre - process the image can highlight the contours in the image, enabling the neural network model to learn richer contour information and making the prediction result more accurate. Using the geometric residual algorithm that minimizes the projected contour of the spacecraft CAD model in the image to optimize the predicted pose, more accurate pose parameters are obtained. The present invention has been preliminarily verified, and the attitude angle error of the algorithm is about 2.68 degrees, and the relative translation error is about 4%.
[0141] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for spacecraft pose estimation based on deep learning, characterized in that, it includes: First, randomly initialize the neural network model parameters according to the normal distribution. Then, normalize the training images and the corresponding pose labels in the dataset respectively and use them to train the neural network model. Next, input the test images into the trained neural network model, and use denormalization to obtain the pose predicted by the neural network model. Finally, according to the pose parameters predicted by the neural network model, use the geometric residual that minimizes the projected contour of the spacecraft CAD model in the image to optimize the pose parameters of the spacecraft, and finally obtain the maximum likelihood estimation of the spacecraft pose parameters; The normalization of the training images includes: First, create a 7×7 grid called the Gaussian kernel W, and based on the grid center, initialize the weight values of the grid with a two-dimensional normal distribution: where w(x,y) is the weight value of each grid in the Gaussian kernel, x,y are the positions of each grid relative to the central grid, and σ is a constant, set to 2; Then unitize the Gaussian kernel W so that the sum of the weight values of all grids is 1: Next, calculate the weighted average value m and the weighted standard deviation of each pixel point in the image m = ∑ ixy w'(x, y)·p i,j+x,k+y where p is each pixel value in the original image; Finally, calculate the pixel value of each normalized pixel point: The normalization and denormalization of the pose label include: The parameters of the pose label include translation parameters and rotation parameters. The rotation parameters are represented by unit quaternions. The three translation parameters of the i-th image are represented by x i , y i , z i respectively. During the training stage, the translation parameters in the pose label corresponding to each training image are normalized: Among them, are the average values corresponding to x i , y i , z i in the training label set respectively, and σ x , σ y , σ z are the corresponding standard deviations respectively, are the corresponding output parameters after normalization; In the test stage, denormalize the normalized prediction parameters to obtain the finally predicted translation parameters: where x out , y out , z out are respectively normalized prediction parameters, is the predicted translation parameter after denormalization.
2. A method for spacecraft pose estimation based on deep learning according to claim 1, characterized in that, the neural network model includes an encoder and a decoder. The encoder adopts a 34-layer residual network layer network framework and uses an adaptive equalization layer to replace the fully connected layer in the residual network. The decoder includes a translation decoder for predicting the translation vector and a rotation decoder for predicting the rotation quaternion. The translation decoder includes two 1×2048 fully connected layers with dropout and finally connects to a 1×3 output vector. The rotation decoder includes two 1×2048 fully connected layers with dropout and connects to a 1×4 fully connected layer. Finally, the output of the 1×4 fully connected layer is connected to a normalization layer so that the output is a 1×4 vector with a unit length of 1 as the predicted value of the rotation quaternion.
3. A method for spacecraft pose estimation based on deep learning according to claim 2, characterized in that, the setting of the neural network model includes: a) Initialize the neural network parameters using the normal distribution and select the ReLU function as the activation function; b) The optimizer uses the Adam optimizer, and the learning rate is adjusted to 0.
001. The first-order decay rate β 1 and the second-order decay rate β 2 of the optimizer are set to 0.9 and 0.009 respectively; c) Set the mini-batch size for training to 8; d) Set dropout to 0.
5.
4. A method for spacecraft pose estimation based on deep learning according to claim 1, characterized in that, in the training stage, minimize the loss function of the translation parameters and the loss function of the rotation parameters simultaneously: where p t and p pre are the true value and predicted value of the translation parameter respectively, and q t and q pre are the true value and predicted value of the rotation parameter respectively.
5. A method for spacecraft pose estimation based on deep learning according to claim 1, characterized in that, in this method, a perspective model of a vacuum camera is adopted. Let the focal length of the camera be f, the pixel aspect ratio be k, the tilt factor be s, and the principal points be u and v. Then the internal parameter matrix of the camera is defined as:
6. A method for spacecraft pose estimation based on deep learning according to claim 1, It is characterized in that in the method, a target pose model is adopted. The target pose model uses a rotation matrix R and a translation vector t to describe the pose of the target relative to the camera coordinate system. Assume that the coordinate of a certain point in space in the camera coordinate system is X, then the corresponding image point x of X in the image is: wherein, and are the homogeneous coordinates of x and X respectively, and the symbol means equal up to a proportionality constant; The rotation matrix adopts a representation method of the rotation matrix based on Lie group and Lie algebra. Let the rotation vector be Ω, then the rotation matrix is expressed as: where \(I\) is a unit vector, \([\Omega]\) x is an anti-symmetric matrix formed by the components of \(\Omega\): When the rotation matrix R is given, the rotation vector is calculated through R: ||Ω|| = arccos((trace(R) - 1) / 2) [Ω] x = ||Ω|| / (2 sin ||Ω||) (R - R T ) The pose of the target is uniquely determined by six parameters, p = [Ω T t T T , and p is called the pose parameter vector. 7. A method for spacecraft pose estimation based on deep learning according to claim 1 It is characterized in that the maximum likelihood estimation of the spacecraft pose parameters based on geometric residual minimization in the method includes: a) Control point selection: Select several control points on the projected contour of the spacecraft CAD model, and then use these control points to find useful image points in the spacecraft image. The projected contour of the spacecraft CAD model on the image plane is a closed polygon, which is called c m , Let S be an edge of the CAD model, and its two endpoints are P 1 and P 2 . Given the pose vector p, the rotation matrix R and the translation vector t are obtained. Then the homogeneous coordinates of the projection points of P1 and P2 on the image plane are: The image of S is a straight line segment with endpoints g 1 and g 2 as endpoints. The straight line equation passing through g 1 and g 2 is: Calculate the midpoint of the image line segment corresponding to each edge of the CAD model. For each midpoint, determine whether the back-projected ray of the midpoint has exactly one intersection with the CAD model. If so, the midpoint must lie on the contour c m . Treat this midpoint as a control point. If not, discard the midpoint so that all control points lie on different edges of the contour polygon c m . b) Matching between the projection contour of the CAD model and the image pixels: At the current moment, an image of the spacecraft is read in, and the edge image of the image is obtained using the Canny operator method. The image edge points belonging to the projection contour c of the CAD model are found from the edge image using the control points. m Taking each control point as the center, within a certain range on the normal line of the control point, search for image edge points. Within the search interval of a certain control point, if an edge point satisfies the following two conditions, then it is considered that the edge point belongs to the projection contour c of the spacecraft CAD model. m : (I) The angle between the normal line of the edge point and the normal line of the given control point is less than a certain threshold; (II) The edge point is the edge point closest to the given control point among all edge points that satisfy (I). c) Pose parameter optimization: Let {q i , i = 1, 2, …, n} be the edge points found to belong to c m . The control point corresponding to each q i lies on a certain side of the polygon c m . The straight-line equation of this side can be calculated from the formula of the straight-line equation passing through g 1 and g 2 . Denote this straight line as l i . l i is a function vector of the pose parameter p. For the correct pose parameter p, q i lies on the straight line l i . The pose parameter p of the spacecraft is obtained by minimizing the geometric distance between {q i} and the corresponding straight line {l i}: Among them, A = diag{1 1 0}, d i represents q i and l i the geometric distance between them; The LM method is used to optimize the non - linear least - squares problem. At the s - th iteration, let the estimated value of p be p(s), and the residual vector is defined as: Then the update process of p is as follows: p (s+1) .= (s) -(J (s)T J (s) ++ζI n ) -1 -J (s)T d (s) where ζ is a non - negative factor that is updated as the optimization process proceeds; I n is the n - order identity matrix; J (s) is the Jacobian matrix of d (s) and is calculated as follows: Among them, The calculation is as follows: Among them, obtained by taking the partial derivative; d) Optimization of spacecraft pose parameters: Taking the spacecraft pose parameters predicted by the deep learning neural network model as the initial value, successively execute a) control point selection, b) matching of the CAD model projection contour and image pixels, and c) pose parameter optimization to obtain the estimated value of the spacecraft pose parameters. Then execute the above process multiple times, with the pose estimation of the result of the previous execution as the initial value each time. After multiple executions, the obtained spacecraft pose parameter estimation result is the final estimated value of the spacecraft pose parameters.
Citation Information
Patent Citations
Six-degree-of-freedom attitude tracking method for target object and terminal equipment
CN112435294A
Multi-label classification method and device based on deep learning and storage medium
CN112651412A
Space target relative pose estimation method based on deep learning
CN113034581A
Cited By
Spacecraft pose estimation and uncertainty modeling method based on intrinsic space
CN121980964A
Spacecraft pose estimation and uncertainty modeling method based on intrinsic space
CN121980964B