A 6D pose estimation system for space targets

The system addresses robustness and power consumption issues in space 6D pose estimation by integrating causal reasoning and low-bit quantization, enhancing feature extraction and reducing network size for efficient space mission deployment.

CN115457368BActive Publication Date: 2025-07-15FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211159308.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-22
Publication Date
2025-07-15
Estimated Expiration
2042-09-22

AI Technical Summary

Technical Problem

The existing space target 6D pose estimation system is poorly robust under harsh imaging conditions, and has large network parameters and high power consumption, making it difficult to effectively deploy in space environments with limited computing resources.

Method used

Counterfactual analysis technology is used to generate counterfactual images, weaken the influence of background features through causal reasoning units, unbiased features are extracted in combination with DarkNet-53 and feature pyramid networks, low-bit quantization network is used for low-bit quantization, and convolutional layers are deployed on the PIM chip, and pose estimation is performed by combining PnP and RANSAC algorithms.

Benefits of technology

It improves the robustness and accuracy of the system, significantly reduces memory footprint and power consumption, and achieves efficient space target 6D pose estimation, suitable for actual deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115457368B_ABST
    Figure CN115457368B_ABST
Patent Text Reader

Abstract

The present invention relates to a 6D pose estimation system for space targets, which includes a feature extraction unit, a causal reasoning unit, a key point detection unit, a pose estimation unit and a neural network quantization unit. The image features are extracted through the DarkNet-53 model, and a feature pyramid network is introduced to complete feature enhancement; the causal reasoning unit obtains unbiased features of the image based on counterfactual analysis technology; the key point detection unit predicts the 2D key point coordinates and confidence based on the unbiased features; the pose estimation unit solves the 6D pose of the satellite based on information such as 2D key points; the neural network quantization unit can significantly reduce the system energy consumption while retaining very little accuracy loss. Compared with the prior art, this system can accurately complete the 6D pose estimation of the satellite under harsh imaging conditions and complex space backgrounds, while saving resources, having lower latency and considerable accuracy, providing a strong basic guarantee for the automated task deployment and future development of the satellite in space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and more particularly to a 6D pose estimation system for space targets. Background Art

[0002] The vision navigation system of spacecraft is a key technology for unmanned space missions. The 6D pose estimation technology of space targets is a prerequisite for many tasks, such as the automatic docking of space stations and debris removal tasks. Compared with ground applications, in space scenarios, many factors should be considered, such as the harsh imaging conditions caused by the lack of atmospheric scattering, as well as limited computing resources and power consumption. In recent years, the space engineering and computer vision communities have carried out a series of studies on the 6D pose estimation system for spaceborne targets. Although these models have achieved remarkable performance, many methods directly transfer the models from ground scenes to space scenes without considering the particularity of space missions. In addition, these works mainly focus on the performance improvement of the models, while ignoring the power consumption and latency of actual deployment on actual spacecraft. Therefore, how to handle the impact of harsh imaging conditions on the 6D pose estimator and reduce power consumption has become a key research point in space intelligent navigation technology.

[0003] For the 6D pose estimation system of space targets, existing models mainly focus on two technical points: 6D pose estimation technology and low-bitwidth quantization technology of neural networks.

[0004] 6D pose estimation based on a single-object scene is a basic task in computer vision. The structures of existing 6D pose estimation networks can be roughly divided into two categories: two-stage methods and one-stage methods. The two-stage methods first complete the 2D key point detection of the object, usually using the corners of the three-dimensional object bounding box, and then use the PnP solver to solve the 6D pose through the 2D-3D correspondence. In the one-stage methods, the network is directly used to predict the object pose, that is, there is no key point detection process, but the pose information of the object is converted into a unit quaternion and a three-dimensional translation vector for direct parameter regression.

[0005] Deep neural networks have achieved excellent results in various tasks, but the huge computational cost has hindered the deployment of these models in the problem of 6D pose estimation of space targets. Researchers must balance the performance and deployment cost of the network, especially in special scenarios where computing resources are strictly limited. Currently, many studies in the academic community have explored the lightweight technologies of neural networks, such as network pruning technology and low-bitwidth quantization technology. The main idea of neural network quantization technology is to map full-precision floating-point numbers to a lower precision (8 bits or lower) through a quantizer, so as to significantly reduce the floating-point operations in matrix multiplication, thereby greatly saving memory occupancy and providing a basis for the design of low-bitwidth artificial intelligence chips.

[0006] The disadvantages of the existing technology are mainly reflected in the following two aspects:

[0007] 1. Poor algorithm robustness: Most of the current research work on 6D pose estimation of space targets only focuses on model design. Usually, the 6D pose estimation method of ground objects is directly applied to space missions. This direct transfer method will cause the model to be vulnerable to complex backgrounds, thus reducing the model robustness.

[0008] 2. Huge network parameter quantity and high system power consumption: Compared with the application of 6D pose estimation of ground objects, the 6D pose estimation task of space targets has strict navigation accuracy requirements, limited computing resources and power consumption limitations. How to reduce the network parameter quantity and operating power consumption is an issue that must be considered in the design process of the aircraft system. Therefore, low parameter quantity, low calculation amount and high calculation efficiency are the necessary characteristics of an excellent 6D pose estimation system. Summary of the Invention

[0009] The purpose of the present invention is to provide a 6D pose estimation system for space targets to overcome the defects of the above-mentioned existing technology.

[0010] The purpose of the present invention can be achieved through the following technical solutions:

[0011] A 6D pose estimation system for space targets includes an image acquisition unit, a feature extraction unit, a causal reasoning unit, a key point detection unit, and a pose estimation unit;

[0012] The image acquisition unit acquires space target images;

[0013] The feature extraction unit extracts features from the input images;

[0014] The causal reasoning unit includes a counterfactual analysis module and a total direct impact module. Among them, the counterfactual analysis module generates counterfactual images based on counterfactual analysis technology; the space target image and the counterfactual image are respectively input into the feature extraction unit to obtain two-way feature information, and the two-way feature information is subjected to a difference operation through the total direct impact calculation module to obtain unbiased features;

[0015] The key point detection unit calculates the 2D key point coordinates and key point confidence degrees of the space target based on the unbiased features;

[0016] The pose estimation unit solves the 6D pose of the space target based on the 2D key points.

[0017] Furthermore, the 6D pose estimation system for space targets further includes a neural network quantization unit, which performs low-bit quantization on the 6D pose estimation system for space targets based on the LSQ-Net quantization network.

[0018] Furthermore, the neural network quantization unit fuses the Rescaling layer, the BN layer, and the activation layer, reducing the space occupied by network parameters.

[0019] Furthermore, the feature extraction unit uses the DarkNet-53 network model as an image feature extractor.

[0020] Furthermore, the feature extraction unit uses the Feature Pyramid Network (FPN) to complete the feature enhancement function.

[0021] Furthermore, the feature extraction unit includes a first branch, a second branch, and a third branch with the same structure, which extract factual features, counterfactual features, and pseudo-counterfactual features respectively;

[0022] Among them, the input of the first branch is the space target image, and the output is the factual feature;

[0023] The input of the second branch is the counterfactual image containing only background information, and the output is the counterfactual feature; the counterfactual feature is used as the label for neural network supervised learning in the third branch;

[0024] The input of the third branch is the space target image, and the output is the pseudo-counterfactual feature containing only background information; the third branch is trained based on the counterfactual feature output by the second branch.

[0025] Furthermore, by calculating the similarity loss between the pseudo-counterfactual feature output by the third branch and the counterfactual feature output by the second branch, the training effect of the third branch is determined, and the third branch training model that meets the preset conditions is selected as the final test model for extracting pseudo-counterfactual features;

[0026] Denote the counterfactual feature as F c , and the pseudo-counterfactual feature as F pc , and use the smoothed L1 norm loss function to calculate the similarity loss L sim , and the calculation formula is as follows:

[0027] L sim = sl1(F c , F pc ).

[0028] Furthermore, the key point detection unit includes a key point coordinate detection module and a key point confidence regression module;

[0029] Among them, the key point coordinate detection module includes a 2D convolutional neural network, group normalization, and a ReLU activation function. Its input is the unbiased feature output by the causal reasoning unit, and the output is the 2D key point coordinates of the space target image;

[0030] The key point confidence regression module includes a 2D convolutional neural network, group normalization, and a ReLU activation function. Its input is the unbiased features output by the causal inference unit, and its output is the confidence of the 2D key point coordinates of the space target image.

[0031] Furthermore, the pose estimation unit uses a PnP solver to solve the 6D pose of the space target based on the 2D key points with high confidence.

[0032] Furthermore, the pose estimation unit performs parameter solving based on the RANSAC algorithm and the EPnP algorithm. Among them, the RANSAC algorithm execution includes the following steps:

[0033] Obtain multiple 3D key points provided by the satellite pose estimation dataset;

[0034] Randomly select n matching point pairs from the matching points of the 3D key points and the 2D key points, where n > 0;

[0035] According to the n matching point pairs, use the EPnP algorithm to obtain a rough pose;

[0036] According to the rough pose, reproject all 3D key points into 2D key points and calculate the reprojection error;

[0037] The reprojection error classifies the key points into inliers and outliers according to the size of the error threshold;

[0038] Judge the number of inliers. If it is less than the error threshold, reselect n points. If it is greater than the error threshold, perform refined calculation of the pose;

[0039] The EPnP algorithm execution includes the following steps:

[0040] Based on the 3D key points, select multiple control points through the PCA algorithm;

[0041] Establish a local coordinate system and represent the 3D key point coordinates with multiple control points;

[0042] Convert the multiple control points into the camera system, and establish multiple control points with the same relationship as the world system in the camera system;

[0043] Solve the coordinates of multiple control points in the camera system, and then use the ICP algorithm to solve the 6D pose.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] (1) The causal inference unit of the present invention is based on counterfactual analysis technology, which can weaken the influence of background features in the space target image and enhance the features of the space target main body such as satellites, thereby improving the robustness and stability of the system.

[0046] (2) The causal inference unit of the present invention obtains unbiased features by taking the difference between factual features and pseudo-counterfactual features. Finally, the unbiased features are used to complete accurate 6D pose estimation. This causal inference unit can solve the pose estimation problems caused by various factors such as lack of atmospheric scattering and drastic changes in object scale, and improve the 6D pose estimation accuracy of the traditional 6D pose estimation system.

[0047] (3) Based on low-bitwidth quantization technology, the neural network quantization unit of the present invention performs low-bit quantization on the networks involved in the system. Among them, performing 8-bit full quantization on the network can reduce the memory occupancy by 75%, while the prediction accuracy only decreases by 4.74%. Performing 3-bit full quantization on the network can reduce the memory occupancy by 90.63%, while the prediction accuracy only decreases by 10.17%. This shows that the present invention can significantly reduce the system energy consumption while maintaining high accuracy, which is beneficial for actual deployment and better serves space automation tasks.

[0048] (4) The present invention uses a PIM chip to complete the test of the convolutional layer on the FPGA. The results show that the PIM accelerator achieves the lowest time consumption at a clock rate of 100 MHz. The lower latency proves the high efficiency of low-bitwidth quantization and demonstrates the feasibility of the actual deployment of this system, filling the gap in the actual deployment work of the space target 6D pose estimation system and providing better guarantee for the development of space automation tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is the system principle block diagram of the present invention;

[0050] Figure 2 is the principle block diagram of the feature extraction unit of the present invention;

[0051] Figure 3 is the principle block diagram of the causal inference unit of the present invention;

[0052] Figure 4 is the principle block diagram of the key point detection unit of the present invention;

[0053] Figure 5 is the principle block diagram of the pose estimation unit of the present invention;

[0054] Figure 6 is the principle block diagram of the neural network quantization unit of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0055] The present invention will be described in detail below with reference to the drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives detailed implementation manners and specific operation processes. However, the protection scope of the present invention is not limited to the following embodiments.

[0056] As Figure 1 shown, a 6D pose estimation system for space targets includes an image acquisition unit, a feature extraction unit, a causal reasoning unit, a key point detection unit, a pose estimation unit, and a neural network quantization unit;

[0057] The image acquisition unit is responsible for acquiring space target images. In a specific implementation, the image acquisition unit includes imaging devices such as cameras;

[0058] The feature extraction unit extracts features from the input image, including multi-scale factual features, counterfactual features, and pseudo-counterfactual features, which is suitable for the case where the satellite resolutions are different;

[0059] The causal reasoning unit includes a counterfactual analysis module and a total direct impact module. Among them, the counterfactual analysis module generates counterfactual images based on counterfactual analysis technology, and the counterfactual images only contain the background information of the space target images; the space target image and the counterfactual image are respectively input into the feature extraction unit to obtain two-way feature information, and the two-way feature information is subjected to a difference operation through the total direct impact calculation module to obtain unbiased features;

[0060] The key point detection unit calculates the coordinates and confidence levels of multiple 2D key points of the space target based on the unbiased features;

[0061] The pose estimation unit solves the 6D pose of the space target based on the 2D key points with high confidence.

[0062] The neural network quantization unit performs low-bit quantization on different parts of the neural network structure in the 6D pose estimation system for space targets based on low-bitwidth quantization technology.

[0063] In this embodiment, the space target image includes a satellite image and a background image.

[0064] In this embodiment, the feature extraction unit uses the DarkNet-53 network model as an image feature extractor and cooperates with the Feature Pyramid Network (FPN) to complete the feature enhancement function. Among them, DarkNet-53 is a pre-trained model that has been trained on a subset of the ImageNet database and has been trained with more than one million images. It is the backbone network of YOLOv3 and can extract the 8-fold, 16-fold, or 32-fold downsampled features of the feature map. A large number of 1x1 convolutional and 3x3 convolutional modules are used inside the DarkNet-53 network. Among them, the 1x1 convolution is mainly used for expanding and reducing the channels of the feature map, and the 3x3 convolution is used for feature capture. The residual structure is added to the main network structure to realize the deep stacking of the network, which is used to support the network to extract higher-level semantic features and avoid the gradient disappearance and explosion phenomena during the training process. The FPN structure processes and analyzes the feature map at 5 different scales to obtain high-quality image features.

[0065] At the same time, the feature extraction unit includes a first branch, a second branch, and a third branch with the same structure. Each branch is composed of DarkNet-53 and FPN, and extracts multi-scale factual features, counterfactual features, and pseudo-counterfactual features respectively;

[0066] Among them, the input of the first branch is the space target image, and the output is the multi-scale factual features. It extracts the multi-scale features of the satellite in the complex space background. The multi-scale features correspond to the features of multiple receptive fields and are applicable to the scenarios of the satellite at different resolutions, greatly enhancing the robustness of the system and laying a solid foundation for 6D pose estimation;

[0067] The input of the second branch is the counterfactual image that only contains background information, and the output is the multi-scale counterfactual features; the counterfactual features are used as the labels for the supervised learning of the neural network in the third branch;

[0068] The input of the third branch is the space target image, and the output is the pseudo-counterfactual image that only contains background information; the third branch is trained based on the counterfactual features output by the second branch. The third branch and the first branch share the modules outside the FPN. The only difference between it and the first branch is the FPN module. The FPN module of the third branch will simulate the counterfactual features, that is, imagine an image without a satellite but only with background.

[0069] In this embodiment, the counterfactual analysis module in the causal reasoning unit uses a causal reasoning method based on counterfactual analysis technology to complete the weakening of the background features and the enhancement of the target features of the space target image. That is, by imagining a counterfactual path to erase the main object in the image and only retain the background to construct counterfactual features, and removing the influence of the background features from the total features through the calculation of the Total Direct Effect;

[0070] Specifically, first, counterfactual image generation is achieved through the target segmentation labels in the satellite pose estimation dataset, that is, only the background information is retained. Then, the original space target image and the counterfactual image are respectively input into the feature extraction unit to obtain features. The two-way feature information is subjected to a difference operation through the total direct impact module to obtain unbiased features for subsequent key point detection. This unit can effectively weaken the interference of the complex background information of the space target image, that is, weaken the influence of the background features and enhance the features of the satellite, thereby improving the robustness and stability of the system and facilitating the acquisition of high-precision 6D pose estimation results.

[0071] Among them, the satellite pose estimation dataset is a publicly available dataset on the network.

[0072] In this embodiment, after the causal inference unit completes the suppression of background interference information and the enhancement of target features, the key point detection unit, based on the unbiased features, completes the detection of the 2D key point coordinates of 8 satellites and the regression of the key point confidence. Among them, the confidence is an important basis for considering the quality of the key points, increasing the selectable range of satellite key points. Specifically, the key points use a 2D convolutional neural network to map the unbiased features into multiple 2D key point coordinates and key point confidences. The 2D key points with high confidence will be retained for the calculation of the final prediction result of the model.

[0073] The key point detection unit includes a key point coordinate detection module and a key point confidence regression module;

[0074] Among them, the key point coordinate detection module includes a 2D convolutional neural network, group normalization, and a ReLU activation function. Its input is the unbiased features output by the causal inference unit, and the output is the 2D key point coordinates of the corners of the space target.

[0075] The key point confidence regression module includes a 2D convolutional neural network, group normalization, and a ReLU activation function. Its input is the unbiased features output by the causal inference unit, and the output is the confidence of the 2D key point coordinates of the corners of the space target.

[0076] The 6D pose estimation algorithms in the prior art can be divided into two categories according to the structure, the two-stage method and the one-stage method. The two-stage method first completes the 2D key point detection and then calculates the object pose through the 6D pose estimation algorithm; the one-stage method directly uses the network to predict the object pose. The method used in this embodiment belongs to the first type of two-stage method.

[0077] In the present invention, the 6D pose of the space target refers to the translation and rotation transformation that occurs when the camera system is relative to the world system where the space target is located at the moment of obtaining the current space target image. The 6D pose of the space target is represented by the RT transformation from its world system to the camera system, that is, T c= R × T m + T, where R represents the rotation matrix and T represents the translation vector, and T c represents the 3D key points in the camera coordinate system, and T m represents the 3D key points in the world coordinate system.

[0078] In this embodiment, after obtaining the coordinates of 8 2D key points in the image, combined with the 8 3D key points T provided by the satellite pose estimation dataset m , the PnP solver is used to complete the solution of the rotation matrix R and the translation vector T. Specifically, the coordinates of the 2D key points with high confidence and the coordinates of the 3D key points provided by the satellite pose estimation dataset are paired and sent into the pose solver, and the RANSAC algorithm is used as the parameter solution algorithm.

[0079] In this embodiment, the neural network quantization unit performs low-bit quantization on different parts of the network structure in the 6D pose estimation system of space targets based on the LSQ-Net algorithm. The core idea of network quantization technology is to reduce the number of bits used to represent the network parameter weights, thereby greatly reducing the memory occupancy and bandwidth during network operation. This embodiment uses the LSQ-Net quantization strategy to perform 8-bit and 3-bit quantization on the entire network of the 6D pose estimation system, enabling floating-point parameters to be converted into low-width values for storage, greatly reducing the scale of the network, reducing the network from 205MB to 19.23MB, achieving approximately 90% memory savings, reducing the energy consumption of the network, and enabling the system to better serve various space mission applications.

[0080] At the same time, this embodiment tests a convolutional layer in the 6D pose estimation system of space targets on a PIM chip. This is the first time that low-width quantization technology has been applied to the 6D pose estimation task of space targets, bringing new contributions to the computer vision community and the space technology exploration community.

[0081] As Figure 2 shown, it is the principle block diagram of the feature extraction unit, which includes three branches with the same structure, respectively extracting factual features, counterfactual features, and pseudo-counterfactual features. Specifically as follows:

[0082] (1) The first branch: Similar to the feature extraction module of the traditional 6D satellite pose estimation system, given a satellite image with a complex background, the first branch extracts the features of the image and performs feature enhancement through a feature pyramid to obtain multi-scale features with multiple receptive fields, which is used to solve the problem of large differences in satellite resolution sizes and lay a solid foundation for subsequent accurate 6D satellite poses.

[0083] (2) The second branch: The input is an image with the satellite removed and only the space background, which is provided by the satellite pose estimation dataset pre-constructed by technicians. Its output is also the counterfactual multi-scale features enhanced by the feature pyramid, and these features will be used as the labels for the supervised learning of the third branch. The second branch is only used during the training of the entire network to help the third branch learn the complex background information. Since images containing only the background cannot be obtained in the real application scenarios of satellites, in the training stage, the present invention introduces the third branch to fit the background information, so that the background information can be directly decoupled from the original image in the actual application scenarios, thus cleverly solving the problem of difficult acquisition of background images.

[0084] (3) The third branch: Shares the weights before the FPN with the first branch. Its input is also the original image of the satellite with a complex background, and the output is the multi-scale pseudo-counterfactual features without the satellite and only with background information, that is, this branch imagines an image without the satellite and only with a complex space background, which is used to decouple the unbiased satellite features subsequently.

[0085] During the training stage, the third branch fits the output of the second branch. By calculating the similarity loss between the pseudo-counterfactual features output by the third branch and the counterfactual features output by the second branch, the training effect of the third branch can be determined, and the training model of the third branch that meets the preset conditions is selected as the final test model for extracting pseudo-counterfactual features, which can directly extract the background information in the space target image during the test stage. In a specific implementation manner, the preset conditions of the test model are formulated by professional technicians in the field.

[0086] In this embodiment, the counterfactual features are denoted as F c , and the pseudo-counterfactual features are denoted as F pc . The similarity loss L sim is calculated using the smoothed L1 norm loss sl1(.,.), and the calculation formula is as follows:

[0087] L sim = sl1(F c , F pc )

[0088] As Figure 3 shown, it is the principle block diagram of the causal reasoning unit. Through the theory in causal analysis, the unbiased features of the satellite can be obtained by taking the difference between the factual features and the pseudo-counterfactual features, and finally the unbiased features are used to complete the accurate 6D pose estimation. Due to various factors such as the lack of atmospheric scattering and the drastic change of object scale, the pose estimation task in space has more difficulties than the ordinary 6D pose estimation task. This causal reasoning unit can solve the above problems, thereby improving the accuracy of the traditional 6D satellite pose estimation network.

[0089] As Figure 4 shown, it is the principle block diagram of the key point detection unit. This unit is based on a 2D convolutional neural network, and uses unbiased features to calculate the 2D key point coordinates and their confidence levels of the space target image, which are used to select high-quality and high-confidence satellite 2D key points from a large number of low-quality key points.

[0090] As Figure 5 shown, it is the principle block diagram of the pose estimation unit. After obtaining the 2D key point coordinates, the pose estimation unit obtains the satellite 3D key points and the camera internal parameters given by the satellite pose estimation dataset, and performs steps including using the RANSAC algorithm and the EpnP algorithm, specifically as follows:

[0091] (1) Steps of the RANSAC algorithm: First, based on the 8 2D key point coordinates detected by the causal reasoning unit, combined with the 8 3D key points T m provided by the satellite pose estimation dataset, randomly select n matching point pairs among the matching points of the 3D key points and the 2D key points. In this embodiment, n≥4 because the EPnP algorithm requires at least 4 points; use the EPnP algorithm to obtain a rough pose according to these n matching point pairs. Secondly, according to the rough pose, re-project all 3D key points into 2D key points, and calculate the reprojection error. The unit of the reprojection error is pixels; the reprojection error classifies the key points into inliers and outliers according to the error threshold. Finally, judge the number of inliers. If it is less than the set threshold, re-select n points. If it is greater than the threshold, perform a refined calculation of the pose. RANSAC can robustly estimate the model parameters and can estimate high-precision parameters from a dataset containing a large number of outliers.

[0092] (2) Steps of the EpnP algorithm: Use the known 3D key points to select 4 control points through the PCA algorithm, establish a new local coordinate system, and thus represent the 3D key point coordinates with the new control points. Then, use the camera projection model and the 2D key points to convert them into the camera system, and then establish 4 control points in the camera system with the same relationship as the world system, that is, the coordinates of each point at the control points in the camera system and the world system are the same, solve the coordinates of the 4 control points in the camera system, and then use the ICP algorithm to solve the 6D pose. This step performs RANSAC screening on the result while solving the PnP problem, and thus obtains a more accurate result.

[0093] As Figure 6 shown, it is the principle block diagram of the neural network quantization unit. This unit quantizes the network structure involved in the space target 6D pose estimation system proposed by the present invention based on the LSQ-Net quantization network, and deploys it on the PIM chip, specifically as follows:

[0094] LSQ Quantization: The present invention uses the LSQ algorithm to perform 3-bit / 8-bit quantization on the proposed network. The main idea of quantization is to map full-precision floating-point numbers to low-bitwidth representations through quantizers and dequantizers, thereby achieving a significant reduction in floating-point operation volume. The LSQ algorithm introduces a new means to estimate and extend the task loss gradient of the quantizer step size for each weight and activation layer, so that it can be learned together with other network parameters.

[0095] Before the convolution operation in this embodiment, all weights and activations will complete low-bitwidth quantization through the learned step size. Considering that the convolution layer and its adjacent BN layer can be equivalent to a convolution layer with bias, the present invention fuses the Rescaling layer, BN layer, and activation layer, significantly reducing the space occupied by network parameters. Among them, 8-bit full quantization of the network can reduce the memory occupancy by 75%, while the prediction accuracy only decreases by 4.74%. 3-bit full quantization of the network can reduce the memory occupancy by 90.63%, while the prediction accuracy only decreases by 10.17%. This shows that the present invention can maintain high accuracy while significantly reducing system energy consumption, which is beneficial to the actual deployment of the network and better serves space automation tasks.

[0096] (2) PIM Chip Deployment: In the PIM structure, the quantized network can be placed inside the chip, thus eliminating the need to transfer data from DRAM. Due to the limitation of FPGA hardware resources, only a 3-bit quantized convolution layer of the network is deployed in this embodiment, with a feature map size of 128×128×64 and a kernel size of 128×64×3×3. Through actual tests and theoretical calculations, the present embodiment compares the time consumption of the PIM architecture on FPGA with that on ARM v8.2 CPU and Intel Core-i7 CPU. The results show that the PIM accelerator achieves the lowest time consumption of 5.99 ms at a clock rate of 100 MHz. Compared with ARM v8.2, the deployment of the present invention achieves a 4.4-fold speed increase. Compared with the Intel Core-i7 CPU, the speed is increased by 1.7 times. Therefore, this embodiment uses a PIM chip to complete the test of the convolution layer on FPGA, where the lower latency proves the high efficiency of low-bitwidth quantization and demonstrates the feasibility of actual deployment, filling the gap in the actual deployment work of the system and providing better guarantee for the development of space automation tasks.

[0097] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art shall fall within the protection scope determined by the claims.

Claims

1. A 6D pose estimation system for space targets, characterized in that, It includes an image acquisition unit, a feature extraction unit, a causal reasoning unit, a key point detection unit, and a pose estimation unit; The image acquisition unit acquires a space target image; The feature extraction unit extracts features from the input image; The causal reasoning unit includes a counterfactual analysis module and a total direct impact module. Among them, the counterfactual analysis module generates a counterfactual image based on counterfactual analysis technology; the space target image and the counterfactual image are respectively input into the feature extraction unit to obtain two-way feature information, and the two-way feature information is subjected to a difference operation through the total direct impact calculation module to obtain unbiased features; The key point detection unit calculates the 2D key point coordinates and key point confidence of the space target based on the unbiased features; The pose estimation unit solves the 6D pose of the space target based on the 2D key points; 2. The 6D pose estimation system for space targets according to claim 1, characterized in that, It also includes a neural network quantization unit, which performs low-bit quantization on the 6D pose estimation system of the space target based on the LSQ-Net quantization network.

3. The 6D pose estimation system for space targets according to claim 2, characterized in that, The neural network quantization unit integrates the Rescaling layer, the BN layer, and the activation layer, reducing the space occupied by network parameters.

4. The 6D pose estimation system for space targets according to claim 1, characterized in that, The feature extraction unit uses the DarkNet-53 network model as an image feature extractor.

5. The 6D pose estimation system for space targets according to claim 4, characterized in that, The feature extraction unit uses the Feature Pyramid Network (FPN) to complete the feature enhancement function.

6. The 6D pose estimation system for space targets according to claim 5, characterized in that, The feature extraction unit includes a first branch, a second branch, and a third branch with the same structure, which respectively extract factual features, counterfactual features, and pseudo-counterfactual features; Among them, the input of the first branch is the space target image, and the output is factual features; The input of the second branch is a counterfactual image containing only background information, and the output is counterfactual features; the counterfactual features are used as the labels for neural network supervised learning in the third branch; The input of the third branch is the space target image, and the output is pseudo-counterfactual features containing only background information; the third branch is trained based on the counterfactual features output by the second branch.

7. The 6D pose estimation system for space targets according to claim 6, characterized in that, By calculating the similarity loss between the pseudo-counterfactual features output by the third branch and the counterfactual features output by the second branch, the training effect of the third branch is determined, and the third branch training model that meets the preset conditions is selected as the final test model for extracting pseudo-counterfactual features; Denote the counterfactual feature as F c , and the pseudo-counterfactual feature as F pc . Calculate the similarity loss L using the smoothed L1 norm loss function sim . The calculation formula is as follows: L sim = sl1(F c , F pc ).

8. The 6D pose estimation system for space targets according to claim 1, characterized in that, The key point detection unit includes a key point coordinate detection module and a key point confidence regression module; Among them, the key point coordinate detection module includes a 2D convolutional neural network, group normalization, and a ReLU activation function. Its input is the unbiased features output by the causal reasoning unit, and the output is the 2D key point coordinates of the space target image; The key point confidence regression module includes a 2D convolutional neural network, group normalization, and a ReLU activation function. Its input is the unbiased features output by the causal reasoning unit, and the output is the confidence of the 2D key point coordinates of the space target image.

9. The 6D pose estimation system for space targets according to claim 1, characterized in that, The pose estimation unit uses a PnP solver to solve the 6D pose of the space target based on high-confidence 2D key points.

10. The 6D pose estimation system for space targets according to claim 9, characterized in that, The pose estimation unit performs parameter solution based on the RANSAC algorithm and the EPnP algorithm. Among them, the RANSAC algorithm execution includes the following steps: Obtain multiple 3D key points provided by the satellite pose estimation dataset; Randomly select n pairs of matching points from the matching points between 3D key points and 2D key points, where n > 0; Based on the n pairs of matching points, use the EPnP algorithm to obtain a rough pose; According to the rough pose, reproject all 3D key points into 2D key points and calculate the reprojection error; The reprojection error classifies the key points into inliers and outliers according to the size of the error threshold; Judge the number of inliers. If it is less than the error threshold, reselect n points. If it is greater than the error threshold, perform refined pose calculation; The execution of the EPnP algorithm includes the following steps: Based on the 3D key points, select multiple control points through the PCA algorithm; Establish a local coordinate system and represent the 3D key point coordinates with multiple control points; Convert the multiple control points into the camera system and establish multiple control points with the same relationship as the world system in the camera system; Solve the coordinates of the multiple control points in the camera system, and then use the ICP algorithm to solve the 6D pose.

Citation Information

Patent Citations

  • Real-time and efficient 6D attitude estimation network, construction method and estimation method

    CN112561995A

  • Iterative 6D pose estimation method and device based on deep learning

    CN114119999A