Monocular six-degree-of-freedom pose estimation method and apparatus for space objects

By combining the edge attention module and the depth estimation network, and using mask segmentation and attitude estimation network for attitude reconstruction, the problem of low attitude estimation accuracy in low-light and dim environments is solved, and the robustness and accuracy of spacecraft attitude estimation are improved.

CN119359799BActive Publication Date: 2025-10-10INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411221949.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-02
Publication Date
2025-10-10
Estimated Expiration
2044-09-02

AI Technical Summary

Technical Problem

The existing monocular six-degree-of-freedom attitude estimation method for space targets has low estimation accuracy in low-light and dim environments, making it difficult to fully utilize the time information in continuous images. In addition, it is difficult to obtain the real attitude data of the spacecraft, which affects the estimation accuracy.

Method used

Preprocessing is performed through the edge attention module, and the mask segmentation network and depth estimation network are used to extract edge information and depth information. Combined with the pre-trained spatial target six-degree-of-freedom pose estimation network, the pose change information is reconstructed using adjacent frame images and depth information, and the relative pose change and absolute pose information are combined to reduce the cumulative error.

Benefits of technology

It improves the feature extraction capability in low-light and dim environments, enhances the robustness of posture estimation, reduces the impact of the environment on recognition accuracy, and achieves higher posture estimation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119359799B_ABST
    Figure CN119359799B_ABST
Patent Text Reader

Abstract

The application provides a monocular six-degree-of-freedom attitude estimation method and device for a space target, wherein the method comprises the following steps: inputting an image carrying edge information into a mask segmentation network model to obtain binary mask information, so as to determine a current frame object image; inputting the current frame object image and a current frame image into a depth estimation network to obtain current frame object depth information; inputting a neighboring frame image of the current frame image and the current frame object depth information into a pre-trained space target six-degree-of-freedom attitude estimation network to obtain relative space target monocular six-degree-of-freedom attitude change information of the current frame and the neighboring frame; performing image reconstruction based on the relative space target monocular six-degree-of-freedom attitude change information and the current frame object depth information; and determining absolute space target monocular six-degree-of-freedom attitude change information of the current frame based on the relative space target monocular six-degree-of-freedom attitude change information and absolute space target monocular six-degree-of-freedom attitude change information of the neighboring frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence and machine vision technology, and in particular to a monocular six-degree-of-freedom posture estimation method and device for a space target. Background Art

[0002] Monocular six-degrees-of-freedom (6-DoF) attitude estimation of space targets is widely used in the aerospace field. 6-DoF attitude estimation can accurately determine the position and orientation of a spacecraft in space. It is the basis for real-time monitoring and attitude adjustment of spacecraft, providing key data support for precise orbital control and navigation, ensuring smooth landing, navigation and obstacle avoidance, or precise docking.

[0003] Currently, most methods for monocular 6-DoF spacecraft pose estimation mainly use single-frame images and prior information (3D detection box vertices, 3D models, radar point clouds, etc.) to supervise the spacecraft's pose information. This limits their ability to fully utilize the temporal information present in continuous images. Moreover, the above approach ignores the difficulty of obtaining the spacecraft's true pose data and prior information. In addition, in the low-light and dim environment of deep space, the detailed information of the spacecraft is difficult to capture, which has a significant impact on the final estimation accuracy.

[0004] It can be seen from this that the monocular six-degree-of-freedom pose estimation method for space targets in related technologies has a technical problem of low estimation accuracy in low-light and dim environments. Summary of the Invention

[0005] The present invention provides a monocular six-degree-of-freedom pose estimation method and device for space targets, which are used to overcome the defect of low estimation accuracy in low-light and dim environments in the existing monocular six-degree-of-freedom pose estimation methods for space targets, thereby improving the feature extraction capability of the network in dim environments and alleviating the impact of the environment on recognition accuracy.

[0006] The present invention provides a monocular six-degree-of-freedom attitude estimation method for a space target, comprising the following steps: pre-processing the input current frame image of the target spacecraft through an edge attention module to obtain an image carrying edge information; inputting the image carrying edge information into a pre-trained mask segmentation network model to obtain the binary mask information output by the mask segmentation network model, and determining the current frame object image based on the binary mask information; inputting the current frame object image and the current frame image into a depth estimation network to obtain the current frame object depth information output by the depth estimation network; inputting the adjacent frame images of the current frame image and the current frame object depth information into a pre-trained space target A six-degree-of-freedom attitude estimation network is used to obtain relative space target monocular six-degree-of-freedom attitude change information of the current frame and the adjacent frames output by the pre-trained space target six-degree-of-freedom attitude estimation network; image reconstruction is performed based on the relative space target monocular six-degree-of-freedom attitude change information and the object depth information of the current frame to obtain absolute space target monocular six-degree-of-freedom attitude change information of the adjacent frames; and absolute space target monocular six-degree-of-freedom attitude change information of the current frame is determined based on the relative space target monocular six-degree-of-freedom attitude change information and the absolute space target monocular six-degree-of-freedom attitude change information of the adjacent frames.

[0007] According to a monocular six-degree-of-freedom pose estimation method for a space target provided by the present invention, before inputting the adjacent frame images of the current frame image and the current frame object depth information into a pre-trained space target six-degree-of-freedom pose estimation network, the method further comprises: obtaining an initialized space target six-degree-of-freedom pose estimation network based on a sequence of images; inputting the current frame image sample and the adjacent frame image sample into the space target six-degree-of-freedom pose estimation network to obtain the current frame mask sample, the current frame depth map sample and the relative space target monocular six-degree-of-freedom pose change information sample between the current frame and the adjacent frame output by the space target six-degree-of-freedom pose estimation network; determining the loss function of the mask network based on the current frame mask sample and the true mask; inputting the current frame depth map sample into the space target six-degree-of-freedom pose estimation network to obtain the current frame mask sample, the current frame depth map sample and the relative space target monocular six-degree-of-freedom pose change information sample between the current frame and the adjacent frame ... The sample and the monocular six-degree-of-freedom posture change information sample relative to the spatial target are input into the image reconstruction network to obtain the reconstructed current frame image sample; based on the reconstructed current frame image sample and the current frame image sample, the structural similarity loss and the norm loss are determined as the reconstruction loss function; based on the reconstruction loss function and the edge-aware smoothing loss function, the loss function of the depth posture estimation network is determined; based on the loss function of the mask network and the loss function of the depth posture estimation network, the total loss function of the spatial target six-degree-of-freedom posture estimation network is obtained by weighting; when the depth posture estimation network reaches a preset number of training times or the loss function value of the total loss function is lower than a preset threshold, a pre-trained spatial target six-degree-of-freedom posture estimation network is obtained.

[0008] According to a monocular six-degree-of-freedom attitude estimation method for a space target provided by the present invention, the current frame image of the input target spacecraft is preprocessed by an edge attention module to obtain an image carrying edge information, including: performing edge detection on the current frame image based on Sobel convolution by the edge attention module to obtain an edge amplitude map of the current frame image; performing Fourier transform on the edge amplitude map by the edge attention module to obtain a frequency domain feature map of the current frame image; and performing discrete cosine transform based on the frequency domain feature map by the edge attention module to obtain an image carrying edge information.

[0009] According to a monocular six-degree-of-freedom pose estimation method for a space target provided by the present invention, the image carrying edge information is input into a pre-trained mask segmentation network model to obtain binary mask information output by the mask segmentation network model, including: extracting features of the image carrying edge information through the pre-trained mask segmentation network model to obtain features of different scales; inputting the features of different scales into an edge-based Fourier frequency domain attention module through the pre-trained mask segmentation network model to obtain a weighted frequency domain feature map output by the edge-based Fourier frequency domain attention module; and decoding the weighted frequency domain feature map through the pre-trained mask segmentation network model to obtain the binary mask information of the current frame image.

[0010] According to a monocular six-degree-of-freedom pose estimation method for a space target provided by the present invention, the current frame object image and the current frame image are input into a depth estimation network to obtain the current frame object depth information output by the depth estimation network, including: encoding the current frame image through the depth estimation network to obtain a current frame multi-scale feature map; continuously convolving the current frame multi-scale feature map with the current frame object image through an attention gate feature selection module of the depth estimation network to obtain a current frame fusion feature; decoding the current frame fusion feature through a depth decoder of the depth estimation network to obtain a current frame depth map, wherein the current frame depth map is used to represent the current frame object depth information.

[0011] According to a monocular six-degree-of-freedom pose estimation method for a space target provided by the present invention, the adjacent frame images of the current frame image and the depth information of the current frame object are input into a pre-trained space target six-degree-of-freedom pose estimation network to obtain relative space target monocular six-degree-of-freedom pose change information of the current frame and the adjacent frames output by the pre-trained space target six-degree-of-freedom pose estimation network, including: encoding the adjacent frame images of the current frame image through the pre-trained space target six-degree-of-freedom pose estimation network to obtain a multi-scale fusion image of adjacent frames; fusing the multi-scale fusion image of adjacent frames with the multi-scale feature map of the current frame through the attention gate feature selection module of the pre-trained space target six-degree-of-freedom pose estimation network to obtain fusion features of the current frame and adjacent frames; decoding the fusion features of the current frame and adjacent frames through the pose decoder of the pre-trained space target six-degree-of-freedom pose estimation network to obtain relative space target monocular six-degree-of-freedom pose change information of the current frame and adjacent frames.

[0012] The present invention also provides a monocular six-degree-of-freedom posture estimation device for a space target, comprising the following modules: a preprocessing module, for preprocessing the input current frame image of the target spacecraft through an edge attention module to obtain an image carrying edge information; a first determination module, for inputting the image carrying edge information into a pre-trained mask segmentation network model to obtain the binary mask information output by the mask segmentation network model, and determining the current frame object image based on the binary mask information; a first output module, for inputting the current frame object image and the current frame image into a depth estimation network to obtain the current frame object depth information output by the depth estimation network; a second output module, for inputting the adjacent image of the current frame image into the depth estimation network to obtain the current frame object depth information output by the depth estimation network; The frame image and the depth information of the object in the current frame are input into a pre-trained spatial target six-degree-of-freedom posture estimation network to obtain the relative spatial target monocular six-degree-of-freedom posture change information of the current frame and the adjacent frames output by the pre-trained spatial target six-degree-of-freedom posture estimation network; a reconstruction module is used to reconstruct the image based on the relative spatial target monocular six-degree-of-freedom posture change information and the current frame object depth information to obtain the absolute spatial target monocular six-degree-of-freedom posture change information of the adjacent frames; a second determination module is used to determine the absolute spatial target monocular six-degree-of-freedom posture change information of the current frame based on the relative spatial target monocular six-degree-of-freedom posture change information and the absolute spatial target monocular six-degree-of-freedom posture change information of the adjacent frames.

[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, a monocular six-degree-of-freedom posture estimation method for a space target as described in any one of the above is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a monocular six-degree-of-freedom posture estimation method for a space target as described in any one of the above.

[0015] The present invention also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the monocular six-degree-of-freedom pose estimation method for a space target as described in any one of the above.

[0016] The present invention provides a method and device for estimating the monocular six-degree-of-freedom posture of a space target. The method and device perform preprocessing through an edge attention module to obtain an image carrying edge information. The mask segmentation network model is called based on the image carrying edge information to output binary mask information to determine the current frame object image; the depth estimation network is called to perform depth estimation based on the current frame object image and the current frame image, and output the depth information of the current frame object; the pre-trained space target six-degree-of-freedom posture estimation network is used to determine the relative space target monocular six-degree-of-freedom posture change information between the current frame and the adjacent frames based on the adjacent frame images and the current frame object depth information; the image is reconstructed based on the relative space target monocular six-degree-of-freedom posture change information and the current frame object depth information to obtain the absolute space target of the adjacent frames. Monocular six-degree-of-freedom posture change information, thereby combining depth information and relative posture change information, making it more robust to common low-light and dim environments such as illumination changes and occlusions; based on the relative space target monocular six-degree-of-freedom posture change information and the absolute space target monocular six-degree-of-freedom posture change information of the adjacent frames, the absolute space target monocular six-degree-of-freedom posture change information of the current frame is determined. Thus, by combining the relative posture change information and the absolute posture information of the adjacent frames, not only the relative changes between frames are considered, but also the known absolute posture is used as a reference, thereby reducing the cumulative error in posture estimation, and thus solving the technical problem of low estimation accuracy in low-light and dim environments in the space target monocular six-degree-of-freedom posture estimation method in the related art. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced one by one below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0018] Figure 1 The figure is a flow chart of the monocular six-degree-of-freedom posture estimation method for a space target provided by the present invention.

[0019] Figure 2This is an overall flow chart of the monocular six-degree-of-freedom posture estimation method for a space target provided by the present invention.

[0020] Figure 3 This is a training flow chart of a six-degree-of-freedom posture estimation network for a space target based on sequence images provided by the present invention.

[0021] Figure 4 It is the original RGB image provided by the present invention.

[0022] Figure 5 This is the image after high-frequency edge processing provided by the present invention.

[0023] Figure 6 It is the mask image after segmentation provided by the present invention.

[0024] Figure 7 It is an RGB image based on mask sampling provided by the present invention.

[0025] Figure 8 It is a posture estimation visualization image provided by the present invention.

[0026] Figure 9 This is a structural flow chart of the monocular six-degree-of-freedom posture estimation device for a space target provided by the present invention.

[0027] Figure 10 It is a structural diagram of a monocular six-degree-of-freedom posture estimation system for a space target provided by the present invention.

[0028] Figure 11 It is a schematic diagram of the physical structure of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0029] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0030] Monocular six-degree-of-freedom attitude estimation of space targets is widely used in the aerospace field. Six-degree-of-freedom attitude estimation can accurately determine the position and orientation of a spacecraft in space. It is the basis for real-time monitoring and attitude adjustment of spacecraft, and provides key data support for precise orbital control and navigation, ensuring smooth landing, navigation obstacle avoidance or precise docking.

[0031] Most current methods for monocular 6-DoF spacecraft pose estimation primarily rely on supervised training of spacecraft pose information using single-frame images and prior information (such as 3D bounding box vertices, 3D models, and radar point clouds). This limits their ability to fully exploit the temporal information present in continuous images. Furthermore, this approach ignores the difficulty of acquiring both the true pose data and prior information for the spacecraft. Furthermore, in the low-light, dim environment of deep space, detailed information about the spacecraft is difficult to capture, significantly impacting the accuracy of the final estimate.

[0032] In an embodiment of the present invention, a sequence method is used to complete self-supervised training and recognition of posture, which greatly reduces the cost of obtaining prior information; for low-light and dim environments, a Fourier high-frequency attention mechanism is introduced to improve the network's feature extraction capability in dim environments and alleviate the impact of the environment on recognition accuracy.

[0033] A monocular six-degree-of-freedom attitude estimation method for a space target provided by the present invention is mainly used to solve the problem of 6-DoF attitude estimation of a spacecraft under the influence of the environment and the difficulty in obtaining real data.

[0034] refer to Figure 1 , Figure 1 : is a flow chart of the monocular six-degree-of-freedom attitude estimation method for a space target provided by the present invention, such as Figure 1 As shown, the method includes the following steps.

[0035] Step 101: pre-process the input current frame image of the target spacecraft through the edge attention module to obtain an image carrying edge information.

[0036] In the embodiment of the present invention, the current frame image is the current frame RGB image, and the current frame may be a key frame (target frame) selected according to actual application conditions.

[0037] In an embodiment of the present invention, the edge attention module is used to solve the boundary fuzziness problem in image segmentation tasks, that is, the problem that boundary pixels are difficult to accurately classify; by introducing the edge attention mechanism, the model can more accurately locate the edge areas in the image and pay more attention to these areas during the segmentation process, thereby improving the segmentation accuracy and edge clarity.

[0038] According to a monocular six-degree-of-freedom attitude estimation method for a space target provided by the present invention, the current frame image of the input target spacecraft is preprocessed by an edge attention module to obtain an image carrying edge information, which specifically includes the following steps.

[0039] Step 1011: Perform edge detection on the current frame image based on Sobel convolution through the edge attention module to obtain an edge amplitude map of the current frame image.

[0040] In an embodiment of the present invention, a current frame (RGB) image is input, and a Sobel convolution operation is used to perform edge detection on the current frame image to obtain an edge amplitude map of the current frame image.

[0041] Step 1012: Perform Fourier transform on the edge amplitude map through the edge attention module to obtain a frequency domain feature map of the current frame image;

[0042] In the embodiment of the present invention, Fourier transform is used to convert the edge amplitude map into a frequency domain feature map.

[0043] Step 1013: Perform discrete cosine transform based on the frequency domain feature map through the edge attention module to obtain an image carrying edge information.

[0044] In an embodiment of the present invention, a discrete cosine transform type II function is used to enhance high-frequency features such as edges, thereby improving subsequent image processing and analysis tasks. The discrete cosine transform type II formula can refer to the following formula (1):

[0045] (1)

[0046] in, represents a sequence of input signals or data, represents the length of the input sequence, represents the coefficient after transformation, Representation sequence length.

[0047] Through the embodiments of the present invention, high-frequency features such as edges can be enhanced, thereby improving subsequent image processing and analysis tasks.

[0048] In step 102, the image carrying edge information is input into a pre-trained mask segmentation network model to obtain binary mask information output by the mask segmentation network model, and the object image of the current frame is determined based on the binary mask information.

[0049] In an embodiment of the present invention, the mask segmentation network model can use a neural network to automatically learn and extract image features, thereby achieving accurate image segmentation. A mask is a binary image in which the pixel values ​​are usually 0 or 1 and is used to identify specific areas in the image. In the mask segmentation task, the goal of the mask segmentation network model is to generate a corresponding mask for each object in the image.

[0050] The current frame object is the target object located in the current frame image.

[0051] According to a monocular six-degree-of-freedom pose estimation method for a space target provided by the present invention, an image carrying edge information is input into a pre-trained mask segmentation network model to obtain binary mask information output by the mask segmentation network model. Based on the binary mask information, the method specifically includes the following steps:

[0052] Step 1021 , extracting features from the image carrying edge information using a pre-trained mask segmentation network model to obtain features of different scales.

[0053] In an embodiment of the present invention, a lightweight architecture for self-supervised monocular depth estimation (for example, a lightweight LiteMono model) is used to differentiate the RGB image of the input current frame, and after multiple max pooling operations, features of different scales (W / 2×H / 2, W / 4×H / 4, W / 8×H / 8, and W / 16×H / 16) are obtained.

[0054] In the embodiment of the present invention, the difference can refer to the following formula (2):

[0055] (2)

[0056] in, Representing an image In position ( ) and the color channels are The pixel value of and Represents images respectively In position ( ) position Direction and Directional gradient (a gradient is a measure of the rate of change of brightness or color intensity in an image); Representing an image In the row direction there are pixels, Representing an image In the column direction, there are pixels, Representing an image There are three color channels (e.g., RGB images).

[0057] In step 1022, features of different scales are input into the edge-based Fourier frequency domain attention module through the pre-trained mask segmentation network model to obtain a weighted frequency domain feature map output by the edge-based Fourier frequency domain attention module.

[0058] In an embodiment of the present invention, features of different scales are input into an edge-based Fourier frequency domain attention module (DFDA, Diffence Frequency Domain DCT Attention) to obtain a weighted frequency domain feature map.

[0059] Step 1023: decode the weighted frequency domain feature map through the pre-trained mask segmentation network model to obtain the binary mask information of the current frame image.

[0060] In the embodiment of the present invention, the weighted frequency domain feature map is decoded to obtain the mask (W×H, W / 2×H / 2, W / 4×H / 4) of the current frame.

[0061] After step 1023, the above-mentioned determination of the current frame object image based on the binarized mask information specifically includes:

[0062] Step 1024 , select a mask that is consistent with the size of the input current frame image, set the threshold to 0.5, perform pixel sampling on the key frame, and output the current frame object (RGB) image.

[0063] Through the embodiments of the present invention, the pre-trained mask segmentation network model can learn and recognize different objects and their boundaries in the image. By inputting an image carrying edge information, the model can more accurately capture the outline of the object, thereby generating a high-precision binary mask to clearly define the boundary between the object and the background.

[0064] Step 103: Input the current frame object image and the current frame image into a depth estimation network to obtain the current frame object depth information output by the depth estimation network.

[0065] In an embodiment of the present invention, the depth estimation network is used to output the depth (ie, the distance from the camera) information of each pixel in the image.

[0066] According to a monocular six-degree-of-freedom pose estimation method for a space target provided by the present invention, a current frame object image and a current frame image are input into a depth estimation network to obtain the current frame object depth information output by the depth estimation network, specifically comprising the following steps:

[0067] Step 1031: Encode the current frame image through a depth estimation network to obtain a multi-scale feature map of the current frame.

[0068] In an embodiment of the present invention, an RGB image of a current frame is input and encoded using a lightweight architecture (e.g., LiteMono) for self-supervised monocular depth estimation to obtain features of different sizes (i.e., a multi-scale feature map of the current frame).

[0069] Among them, LiteMono is a lightweight convolutional neural network (CNN) and transformer hybrid architecture designed for self-supervised monocular depth estimation, including: continuous expansion convolution module and local-global feature interaction module.

[0070] In step 1032 , the attention gate feature selection module of the depth estimation network is used to continuously convolve the multi-scale feature map of the current frame with the object image of the current frame to obtain the fusion feature of the current frame.

[0071] In an embodiment of the present invention, an attention gate feature selection module (ACSG, Attention CosineSimilarity Gate) is used to perform continuous Conv3x3 convolution on the multi-scale feature map of the current frame and the object image of the current frame to obtain the current frame fusion features.

[0072] Specifically, the multi-scale feature map of the current frame and the object image of the current frame are channel-weighted, and the corresponding channel correlation is calculated using the cosine similarity function. The correlation formula can refer to the following formula (3):

[0073] (3)

[0074] in, Represents correlation, A represents the first feature image (multi-scale feature map of the current frame), B represents the second feature image (object image of the current frame), and i represents the index of the pixel point.

[0075] If the channels are related, the features of the two channels are added together. Otherwise, the two feature maps are connected in the channel dimension and the convolution function is used to compress the feature channels. For details, please refer to the following formula (4):

[0076] (4)

[0077] in, represents the fusion feature (the fusion feature of the current frame), represents the first channel feature (i.e., the feature of the multi-scale feature map of the current frame), Represents the second channel feature (i.e., the feature of the object image in the current frame), represents the convolution operation, represents feature splicing, Indicates relevance.

[0078] Step 1033 : Decode the current frame fusion features through the depth decoder of the depth estimation network to obtain a current frame depth map, wherein the current frame depth map is used to represent the depth information of the current frame object.

[0079] In an embodiment of the present invention, the fusion features of the current frame are input into a depth decoder to obtain a depth map of the current frame to represent the depth information of the object in the current frame.

[0080] Step 104: input the adjacent frame images of the current frame image and the depth information of the current frame object into a pre-trained spatial target six-degree-of-freedom attitude estimation network to obtain the relative spatial target monocular six-degree-of-freedom attitude change information between the current frame and the adjacent frames output by the pre-trained spatial target six-degree-of-freedom attitude estimation network.

[0081] In an embodiment of the present invention, the pre-trained spatial object six-degree-of-freedom pose estimation network is based on sequence images.

[0082] According to a monocular six-degree-of-freedom pose estimation method for a space target provided by the present invention, before inputting adjacent frame images of a current frame image and the depth information of the current frame object into a pre-trained six-degree-of-freedom pose estimation network for a space target, the method further includes the following steps.

[0083] Step B1, obtaining an initialized spatial target six-degree-of-freedom pose estimation network based on sequence images.

[0084] In an embodiment of the present invention, the weights of a sequence image-based space target six-degree-of-freedom pose estimation network are initialized to obtain an initialized sequence image-based space target six-degree-of-freedom pose estimation network.

[0085] Step B2: Input the current frame image sample and the adjacent frame image sample into the space target six-degree-of-freedom attitude estimation network to obtain the current frame mask sample, the current frame depth map sample, and the relative space target monocular six-degree-of-freedom attitude change information sample between the current frame and the adjacent frame output by the space target six-degree-of-freedom attitude estimation network.

[0086] In an embodiment of the present invention, based on the network obtained in step B1, the current frame image (sample) and the adjacent frame image (sample) are input, and the current frame mask and the current frame depth map as well as the relative six-degree-of-freedom posture of the current frame and the adjacent frame are output.

[0087] Step B3: Determine the loss function of the mask network based on the current frame mask sample and the true mask.

[0088] In the embodiment of the present invention, the DiceLoss and BCELoss functions of the output mask (i.e., the current frame mask) and the true mask are calculated and used as the loss function of the mask network;

[0089] Step B4: input the current frame depth map sample and the monocular six-degree-of-freedom posture change information sample relative to the spatial target into the image reconstruction network to obtain the reconstructed current frame image sample.

[0090] Step B5: Based on the reconstructed current frame image sample and the current frame image sample, determine the structural similarity loss and the norm loss as the reconstruction loss function.

[0091] Step B6: Determine the loss function of the deep pose estimation network based on the reconstruction loss function and the edge-aware smoothing loss function.

[0092] In an embodiment of the present invention, the current frame depth map (sample), the relative six-degree-of-freedom posture of the current frame and the adjacent frame, and the adjacent frame image (sample) are input into the image reconstruction network to obtain a reconstructed current frame image, and the structural similarity loss and 1-norm loss of the reconstructed current frame image and the input current frame image are calculated as the reconstruction loss function. The reconstruction loss function and the edge-aware smoothing loss function together constitute the depth pose estimation network loss function.

[0093] Step B7: weighting the loss function of the mask network and the loss function of the deep pose estimation network to obtain the total loss function of the spatial target six-degree-of-freedom pose estimation network.

[0094] Step B8: When the depth pose estimation network reaches a preset number of training times or the loss function value of the total loss function is lower than a preset threshold, a pre-trained spatial target six-degree-of-freedom pose estimation network is obtained.

[0095] In an embodiment of the present invention, the total loss function is a weighted sum of the loss function of the mask network and the loss function of the depth pose estimation network, and steps B2 to B6 are repeatedly performed until a preset number of training times is reached or the loss function value of the reconstruction error is lower than a set threshold.

[0096] According to a method for estimating the monocular six-degree-of-freedom pose of a space target provided by the present invention, adjacent frame images of a current frame image and depth information of an object in the current frame are input into a pre-trained space target six-degree-of-freedom pose estimation network, and relative monocular six-degree-of-freedom pose change information of the current frame and the adjacent frames of the space target is obtained from the pre-trained space target six-degree-of-freedom pose estimation network, including the following specific steps:

[0097] Step 1041 : Encode adjacent frame images of the current frame image through a pre-trained spatial object six-degree-of-freedom pose estimation network to obtain a multi-scale fusion image of adjacent frames.

[0098] In an embodiment of the present invention, adjacent frame images (RGB images) of a current frame image are input and encoded using a lightweight architecture (e.g., LiteMono) for self-supervised monocular depth estimation to obtain features of different scales.

[0099] Step 1042: The multi-scale fusion map of adjacent frames is fused with the multi-scale feature map of the current frame through the attention gate feature selection module of the pre-trained spatial target six-degree-of-freedom pose estimation network to obtain the fusion features of the current frame and the adjacent frames.

[0100] In an embodiment of the present invention, the multi-scale fusion map of adjacent frames is fused with the multi-scale feature map of the previous frame through an attention gate feature selection module (ACSG, Attention CosineSimilarity Gate).

[0101] Step 1043 , decoding the fusion features of the current frame and the adjacent frames through the posture decoder of the pre-trained spatial target six-degree-of-freedom posture estimation network to obtain relative spatial target monocular six-degree-of-freedom posture change information between the current frame and the adjacent frames.

[0102] In an embodiment of the present invention, the fusion features of the current frame and the adjacent frames are input into a posture decoder to obtain the monocular six-degree-of-freedom posture change information of the relative spatial target between the current frame and the adjacent frames.

[0103] It should be noted that in this embodiment of the present invention, the depth estimation network and the attitude estimation network (six-degree-of-freedom attitude estimation network for spatial objects) together form a depth and attitude estimation network. The feature extraction networks of the depth estimation network and the attitude estimation network use LiteMono and share weights.

[0104] Step 105 : reconstructing an image based on the relative space target monocular six-degree-of-freedom posture change information and the object depth information of the current frame to obtain the absolute space target monocular six-degree-of-freedom posture change information of the adjacent frames.

[0105] The above-mentioned image reconstruction based on the relative space target monocular six-degree-of-freedom posture change information and the current frame object depth information is performed to obtain the absolute space target monocular six-degree-of-freedom posture change information of the adjacent frame, specifically including the following steps:

[0106] Step 1051: Obtain the two-dimensional homogeneous matrix and camera intrinsic parameters corresponding to the current frame.

[0107] Step 1052 : reconstruct the current frame based on the object depth information of the current frame and the relative space target monocular six-degree-of-freedom posture change information to obtain the absolute space target monocular six-degree-of-freedom posture change information of the adjacent frame.

[0108] For details, please refer to the following formulas (5) to (6):

[0109] (5)

[0110] (6)

[0111] in, is a two-dimensional homogeneous matrix, is the estimated 2D graph, are the camera parameters, is the attitude matrix, including the translation matrix and the rotation matrix, is the depth map of the current frame, is the RGB image of the current frame, is the estimated RGB image of the current frame.

[0112] Step 106 : Determine the absolute space target monocular six-degree-of-freedom pose change information of the current frame based on the relative space target monocular six-degree-of-freedom pose change information and the absolute space target monocular six-degree-of-freedom pose change information of the adjacent frames.

[0113] Among them, the above-mentioned determination of the absolute space target monocular six-degree-of-freedom posture change information of the current frame based on the relative space target monocular six-degree-of-freedom posture change information and the absolute space target monocular six-degree-of-freedom posture change information of the adjacent frames specifically includes the following steps.

[0114] Step 1061 : Obtain the absolute space target monocular six-degree-of-freedom posture change information and the relative space target monocular six-degree-of-freedom posture change information of adjacent frames.

[0115] Step 1062: Determine the absolute space target monocular six-degree-of-freedom posture change information of the current frame.

[0116] For details, please refer to the following formulas (7) to (8):

[0117] (7)

[0118] (8)

[0119] in, is the estimated object relative rotation change matrix, is the estimated object relative translation change matrix, is the true rotation matrix of the adjacent frames, is the real translation matrix of adjacent frames, is the estimated keyframe rotation matrix, is the estimated keyframe translation matrix.

[0120] Through the embodiments of the present invention, by combining relative posture change information and absolute posture information of adjacent frames, not only the relative changes between frames are considered, but also the known absolute posture is used as a reference, thereby reducing the cumulative error in posture estimation.

[0121] Through the above-mentioned embodiments of the present invention, semi-supervised acquisition of spacecraft pose information is possible without obtaining true pose and depth labels. The edge-based Fourier frequency domain attention module (DFDA) is designed to enhance the texture features of images in the dim environment of deep space, further improving the accuracy of mask segmentation. The attention gate feature selection module (ACSG) directs the network's focus to the spacecraft for feature extraction and fusion, effectively preventing the influence of background changes on pose estimation.

[0122] refer to Figure 2 , Figure 2 This is an overall flow chart of the monocular six-degree-of-freedom posture estimation method for a space target provided by the present invention.

[0123] The following describes an embodiment of a monocular six-degree-of-freedom pose estimation method for a space target proposed by the present invention in a practical application scenario, which specifically includes the following steps.

[0124] Step S10: Select an image of a spacecraft as a key frame image and input it into the edge attention module for image preprocessing.

[0125] Specifically, an image of a spacecraft is selected as a key frame image and input into the edge attention module for image preprocessing.

[0126] First, the current frame is segmented as input, and the Sobel convolution operation is used to detect the edge of the entire image. After obtaining an edge amplitude map of the image, Fourier transform is performed to convert the edge amplitude map into a frequency domain feature map, and discrete cosine transform type II function is used to enhance high-frequency features such as edges.

[0127] The edge information attention module uses an edge extraction algorithm to enhance the image edge information as a weight coefficient and increase the texture information of the edge pixels in the original image.

[0128] In step S20, the image with edge information obtained in the previous step is input into the trained mask segmentation network model to obtain its binary mask information and calculate the RGB image of the object in the current frame.

[0129] Specifically, the image with edge information obtained in the previous step is input into the trained mask segmentation network model to obtain its binary mask information and calculate the RGB image of the current frame object.

[0130] The mask segmentation network uses a lightweight LiteMono model as a feature extraction network, inputs the RGB image of the current frame, and performs difference, multiple maxpooling to obtain different scales, and inputs the edge-based Fourier frequency domain attention module DFDA to obtain a weighted frequency domain feature map. The weighted frequency domain feature map is obtained by the decoder to obtain the mask of the current frame. The mask with the same size as the input current frame RGB image is selected, the threshold is set to 0.5, the key frame is pixel sampled, and the RGB image of the object is output.

[0131] The edge-based Fourier frequency domain attention module (DFDA, Diffence Frequency Domain DCT Attention) aims to enhance the texture features of images in deep space dark and weak environment, and further improves the accuracy of mask segmentation.

[0132] The mask segmentation network takes the current frame RGB image as input to obtain binary mask information. Subsequently, the current frame target RGB image is obtained by sampling according to the mask information. Specifically, the threshold is set to 0.5, and when the binary mask information is greater than 0.5, the pixel value of the corresponding position of the current frame RGB image is retained, and when the binary mask information is less than 0.5, the pixel value of the corresponding position of the current frame RGB image is discarded.

[0133] The mask segmentation network is a supervised network, and the DiceLoss and BCELoss functions of the output mask and the real mask are used as the loss function of the mask network.

[0134] Here, Dice Loss is a loss function commonly used in image segmentation tasks, especially in cases where foreground and background pixels are imbalanced. BCELoss, or Binary Cross-Entropy Loss, is a loss function used in machine learning and deep learning for binary classification problems, used to measure the difference between the model's predicted probability distribution and the true label.

[0135] The relevant formulas can be referred to as formulas (9)-(10) below:

[0136] (9)

[0137] where, represents the Dice loss function, and represent the set of non-zero (or foreground) pixels in the predicted result and the true label, respectively. is a very small positive number to prevent the denominator from being zero.

[0138] Here, when the prediction is completely accurate (ie, X=Y), the Dice coefficient is 1 and the Dice Loss is 0. When the prediction does not overlap with the true label at all, the Dice coefficient is 0 and the Dice Loss is 1.

[0139] (10)

[0140] in, Represents the BCE loss function, when When it is equal to 1 (that is, the true label is positive), the loss function is simplified to If the model predicts The closer it is to 1, the The closer it is to 0, the smaller the loss; on the contrary, if The smaller it is (that is, the greater the probability that the model predicts the negative class), the greater the loss.

[0141] when When it is equal to 0 (that is, the true label is negative), the loss function is simplified to If the model predicts The closer it is to 0, the The closer it is to 0, the smaller the loss; on the contrary, if The larger is (i.e., the greater the probability that the model predicts the positive class), the greater the loss.

[0142] Step S30: Input the RGB image of the current frame object obtained in the previous step and the current frame RGB image into a depth estimation network to obtain depth information of the object.

[0143] Specifically, the current frame RGB image is input and encoded with LiteMono to obtain features of different sizes. The multi-scale feature map of the key frame obtained by the attention gate feature selection module ACSG is fused with the features obtained by continuous Conv3x3 convolution of the object RGB image. The fused features are then input into the depth decoder to obtain the depth map of the current frame.

[0144] Step S40: Select the adjacent frame image as the key frame image in the previous step, and input the features obtained by the depth estimation in the previous step into the pose estimation network to obtain the relative 6-DoF pose change information between the key frame and the adjacent frame.

[0145] Specifically, adjacent frame RGB images are input and encoded by LiteMono to obtain features of different sizes. The fused multi-scale feature maps of adjacent frames are fused with the multi-scale feature maps of key frames output from the depth estimation network through the attention gate feature selection module ACSG. The fused features are input into the pose decoder to obtain the relative 6-DoF pose information of the current frame and the adjacent frames.

[0146] Step S50 , using the relative 6-DoF posture change information obtained in the previous step and the current frame depth information obtained in step S30 , image reconstruction is performed to reconstruct the adjacent frame RGB image into the current frame RGB image.

[0147] Specifically, steps S30 through S50 involve the deep pose estimation network. This is a self-supervised network that does not require true pose and depth labels during training. Instead, the network constrains the current frame using a reconstruction loss and an edge-aware smoothing loss. The reconstruction loss consists of a structural similarity loss (SSIM) and a 1-norm (L1) loss between the current frame and the reconstructed current frame.

[0148] The reconstruction loss formula can refer to the following formula (11):

[0149] (11)

[0150] in, Represents the predicted image With real images The loss between is a hyperparameter used to balance the weight between SSIM loss and L1 loss. The value range is usually between 0 and 1. It represents the structural similarity index, which is used to evaluate the structural similarity between two images. The closer the SSIM value is to 1, the more similar the two images are; the closer it is to 0, the greater the difference. Represents the L1 norm, also known as Manhattan distance or minimum absolute error (MAE), which is used to calculate the sum of the absolute values ​​of the pixel differences between the predicted image and the true image.

[0151] In an embodiment of the present invention, The default is 0.85.

[0152] The structural similarity loss formula can refer to the following formula (12):

[0153] (12)

[0154] in, Represent two image windows (usually image blocks or the entire image), and Respectively The mean of and Respectively The variance of express The covariance of and are two constants used to avoid zero denominators and improve numerical stability.

[0155] Edge-aware smoothness loss is used to optimize the edge smoothness of the depth map. The edge-aware smoothness loss formula can refer to the following formula (13):

[0156] (13)

[0157] in, represents the edge-aware smoothing loss function, Represents images respectively The width and height of represents the set of all pixels in the image, represents the displacement field ( )exist The absolute value of the gradient in the direction, and Representing an image exist Exponential decay term in the direction (adaptive weight).

[0158] Step S60 , calculating the absolute posture information of the current frame based on the relative 6-DoF posture change information obtained in the previous step and the absolute 6-DoF posture information of the adjacent frames.

[0159] Specifically, the estimated rotation error and translation error of the posture can refer to the following formula (14):

[0160] (14)

[0161] in, represents the total attitude error, represents the rotation error, represents the estimated or measured rotation matrix, represents the true rotation matrix, represents the translation error, where is the estimated or measured translation vector, is the real translation vector, Represents the translation vector The Euclidean norm (i.e., length) of .

[0162] The rotation error formula can refer to the following formula (15):

[0163] (15)

[0164] in, represents the rotation error, represents the estimated or measured rotation matrix, represents the true rotation matrix, represents the transposed matrix, The trace of a matrix is ​​the sum of its diagonal elements, Represents the inverse cosine function.

[0165] The translation error formula can refer to the following formula (16):

[0166] (16)

[0167] in, represents the translation error, The estimated translation vector is calculated and the true translation vector The difference between.

[0168] refer to Figure 3 , Figure 3 This is a training flow chart of a six-degree-of-freedom posture estimation network for a space target based on sequence images provided by the present invention.

[0169] In an embodiment of the present invention, a training method of a six-degree-of-freedom pose estimation network for a spatial target based on sequence images can refer to the following steps:

[0170] Step B10: Initialize the weights of the spatial target six-degree-of-freedom pose estimation network based on the sequence images.

[0171] In an embodiment of the present invention, an initialized sequence image-based six-degree-of-freedom pose estimation network for a spatial target is obtained;

[0172] Step B20: input the current frame image and the adjacent frame image, and output the current frame mask, the current frame depth map, and the relative six-degree-of-freedom posture of the current frame and the adjacent frame.

[0173] In this embodiment of the present invention, based on the network obtained in step B10, the current frame image and the adjacent frame image are input, the current frame mask and the current frame depth map and the relative six-degree-of-freedom posture of the current frame and the adjacent frame are output, and the DiceLoss and BCELoss functions of the output mask and the true mask are calculated for the loss function of the mask network;

[0174] Step B30: Input the current frame depth map and relative posture information obtained in the previous step and the adjacent frame images into the image reconstruction network to obtain a reconstructed current frame image.

[0175] In an embodiment of the present invention, the structural similarity loss and the 1-norm loss between the reconstructed current frame image and the input current frame image are calculated as the reconstruction loss function. The reconstruction loss function and the edge-aware smoothing loss function together constitute the depth pose estimation network loss function.

[0176] Step B40, calculate the loss function, the total loss function is the weighted sum of the loss function of the mask network and the loss function of the depth pose estimation network, and update the network parameters by gradient descent.

[0177] In an embodiment of the present invention, the total loss function is a weighted sum of the loss function of the mask network and the loss function of the depth pose estimation network, and steps B20 and B30 are repeatedly performed until a preset number of training times is reached or the loss function value of the reconstruction error is lower than a set threshold.

[0178] In the simulation experiment of the embodiment of the present invention, the software used is: Python 3.8, processor: E5, graphics card RTX3090, and operating system: Ubuntu.

[0179] In this embodiment of the present invention, an NVIDIA RTX3090 was used in a Linux environment to train and evaluate the Swisscube dataset based on the PyTorch platform. The dataset consisted of 33,000 images, 350 sequences, and 12 batches. The image size was 416x416. The feature extraction networks for the deep network and the pose network were pre-trained on the KITTY dataset. Both networks used the same Adam optimizer with a weight decay of 1e−2 and an initial learning rate of 1e-4. Training was performed for a total of 30 epochs, after which the learning rate decayed to 1e-5. The feature extraction network used in the mask segmentation network was a ResNet18 pre-trained on ImageNet. The optimizer also used Adam with a weight decay of 1e-2 and an initial learning rate of 1e-3. The learning rate decayed to 1e-4 and 1e-5 after 10 and 20 epochs, respectively.

[0180] The key frame images were segmented by masks, sampled according to the masks, and finally the relative change was estimated. During verification, the true pose of the previous frame was used to calculate the rotation matrix and translation matrix to obtain the predicted pose of the key frame, and finally the error was calculated with the true pose of the key frame. We selected one frame of the final result for visualization and numerical display, where the angle error was 0.051 and the translation error was 0.016. For details, please refer to Figures 4 to 8 , Figures 4 to 8 To visualize the results: Figure 4 is the original RGB image, Figure 5 is the image after high-frequency edge processing, Figure 6 is the mask image after segmentation, Figure 7 is the RGB image based on mask sampling, Figure 8 Pose estimation visualization image.

[0181] The following describes a monocular six-degree-of-freedom attitude estimation device for a space target provided by the present invention. The monocular six-degree-of-freedom attitude estimation device for a space target described below and the monocular six-degree-of-freedom attitude estimation method for a space target described above can be referenced to each other.

[0182] refer to Figure 9 , Figure 9 This is a structural flow chart of a monocular six-degree-of-freedom posture estimation device for a space target provided by the present invention, which includes: a preprocessing module 901, a first determination module 902, a first output module 903, a second output module 904, a reconstruction module 905, and a second determination module 906.

[0183] A preprocessing module 901 is configured to preprocess the input current frame image of the target spacecraft through an edge attention module to obtain an image carrying edge information;

[0184] A first determination module 902 is configured to input the image carrying edge information into a pre-trained mask segmentation network model, obtain binary mask information output by the mask segmentation network model, and determine the object image of the current frame based on the binary mask information;

[0185] A first output module 903 is configured to input the current frame object image and the current frame image into a depth estimation network to obtain the current frame object depth information output by the depth estimation network;

[0186] The second output module 904 is configured to input adjacent frame images of the current frame image and the depth information of the object of the current frame into a pre-trained spatial target six-degree-of-freedom pose estimation network, and obtain relative spatial target monocular six-degree-of-freedom pose change information of the current frame and the adjacent frames output by the pre-trained spatial target six-degree-of-freedom pose estimation network;

[0187] A reconstruction module 905 is configured to perform image reconstruction based on the relative space target monocular six-degree-of-freedom posture change information and the current frame object depth information to obtain the absolute space target monocular six-degree-of-freedom posture change information of the adjacent frame;

[0188] The second determining module 906 is configured to determine the absolute space target monocular six-degree-of-freedom posture change information of the current frame based on the relative space target monocular six-degree-of-freedom posture change information and the absolute space target monocular six-degree-of-freedom posture change information of the adjacent frame.

[0189] Specifically, the monocular six-degree-of-freedom attitude estimation device for the above-mentioned space target provided by the present invention can implement all the method steps implemented in the above-mentioned monocular six-degree-of-freedom attitude estimation method embodiment for the space target, and can achieve the same technical effect. The parts and beneficial effects that are the same as those in the method embodiment will not be described in detail here.

[0190] refer to Figure 10 , Figure 10 : This is a schematic diagram of the structure of a monocular six-degree-of-freedom attitude estimation system for a space target provided by the present invention, comprising:

[0191] The mask segmentation module inputs the RGB image of the current frame, obtains the corresponding mask, and outputs the RGB image of the target in the current frame according to the set threshold;

[0192] The depth estimation module inputs the current frame RGB image and the current frame target RGB image. The former is extracted by LiteMono, and the latter is extracted by continuous Conv3x3 convolution. The features extracted by the two are fused and the fused features are decoded to obtain the depth map of the current frame.

[0193] The pose estimation module inputs the RGB images of adjacent frames, extracts features through LiteMono, and fuses them with the fusion features obtained by the depth estimation network. After decoding by the pose decoder, it outputs the relative pose information of the current frame and the adjacent frames;

[0194] The image reconstruction module inputs the current frame depth map obtained by the depth estimation module, the adjacent frame RGB image and the relative posture information obtained by the posture estimation module, and outputs the reconstructed current frame RGB image.

[0195] Figure 11 This is a schematic diagram of the physical structure of the electronic device provided by the present invention, such as Figure 11As shown, the electronic device can include a processor 1110, a communications interface 1120, a memory 1130, and a communications bus 1140, wherein the processor 1110, the communications interface 1120, and the memory 1130 complete mutual communication through the communications bus 1140. The processor 1110 can invoke a logical instruction in the memory 1130 to execute the monocular six-degree-of-freedom attitude estimation method of a space target, which includes: preprocessing an input current frame image of a target spacecraft through an edge attention module to obtain an image carrying edge information; inputting the image carrying edge information into a pre-trained mask segmentation network model to obtain binary mask information output by the mask segmentation network model, and determining a current frame object image based on the binary mask information; inputting the current frame object image and the current frame image into a depth estimation network to obtain current frame object depth information output by the depth estimation network; inputting a neighboring frame image of the current frame image and the current frame object depth information into a pre-trained six-degree-of-freedom attitude estimation network of a space target to obtain relative spatial target monocular six-degree-of-freedom attitude change information of the current frame and the neighboring frame output by the pre-trained six-degree-of-freedom attitude estimation network of the space target; performing image reconstruction based on the relative spatial target monocular six-degree-of-freedom attitude change information and the current frame object depth information to obtain absolute spatial target monocular six-degree-of-freedom attitude change information of the neighboring frame; and determining absolute spatial target monocular six-degree-of-freedom attitude change information of the current frame based on the relative spatial target monocular six-degree-of-freedom attitude change information and the absolute spatial target monocular six-degree-of-freedom attitude change information of the neighboring frame.

[0196] In addition, the logical instruction in the memory 1130 described above can be implemented in the form of a software function unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0197] On the other hand, the present invention also provides a computer program product, which includes a computer program, which can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the monocular six-degree-of-freedom attitude estimation method of the space target provided by the above methods, and the method includes: preprocessing the current frame image of the input target spacecraft through an edge attention module to obtain an image carrying edge information; inputting the image carrying edge information into a pre-trained mask segmentation network model to obtain the binary mask information output by the mask segmentation network model, and determining the current frame object image based on the binary mask information; inputting the current frame object image and the current frame image into a depth estimation network , obtain the current frame object depth information output by the depth estimation network; input the adjacent frame images of the current frame image and the current frame object depth information into the pre-trained spatial target six-degree-of-freedom attitude estimation network, and obtain the relative spatial target monocular six-degree-of-freedom attitude change information of the current frame and the adjacent frames output by the pre-trained spatial target six-degree-of-freedom attitude estimation network; reconstruct the image based on the relative spatial target monocular six-degree-of-freedom attitude change information and the current frame object depth information, and obtain the absolute spatial target monocular six-degree-of-freedom attitude change information of the adjacent frames; determine the absolute spatial target monocular six-degree-of-freedom attitude change information of the current frame based on the relative spatial target monocular six-degree-of-freedom attitude change information and the absolute spatial target monocular six-degree-of-freedom attitude change information of the adjacent frames.

[0198] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a monocular six-degree-of-freedom attitude estimation method for a space target provided by the above-mentioned methods, the method comprising: preprocessing the current frame image of the input target spacecraft through an edge attention module to obtain an image carrying edge information; inputting the image carrying edge information into a pre-trained mask segmentation network model to obtain binary mask information output by the mask segmentation network model, and determining the current frame object image based on the binary mask information; inputting the current frame object image and the current frame image into a depth estimation network to obtain the current frame image output by the depth estimation network. frame object depth information; input the adjacent frame images of the current frame image and the current frame object depth information into the pre-trained spatial target six-degree-of-freedom attitude estimation network to obtain the relative spatial target monocular six-degree-of-freedom attitude change information of the current frame and the adjacent frames output by the pre-trained spatial target six-degree-of-freedom attitude estimation network; reconstruct the image based on the relative spatial target monocular six-degree-of-freedom attitude change information and the current frame object depth information to obtain the absolute spatial target monocular six-degree-of-freedom attitude change information of the adjacent frames; determine the absolute spatial target monocular six-degree-of-freedom attitude change information of the current frame based on the relative spatial target monocular six-degree-of-freedom attitude change information and the absolute spatial target monocular six-degree-of-freedom attitude change information of the adjacent frames.

[0199] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0200] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A monocular six-degree-of-freedom pose estimation method for a space target, characterized in that: include: The edge attention module is used to preprocess the input current frame image of the target spacecraft to obtain an image with edge information; Inputting the image carrying edge information into a pre-trained mask segmentation network model to obtain binary mask information output by the mask segmentation network model, and determining the current frame object image based on the binary mask information; Inputting the current frame object image and the current frame image into a depth estimation network to obtain the current frame object depth information output by the depth estimation network; Inputting adjacent frame images of the current frame image and the depth information of the object of the current frame into a pre-trained space target six-degree-of-freedom attitude estimation network, obtaining relative space target monocular six-degree-of-freedom attitude change information of the current frame and the adjacent frames output by the pre-trained space target six-degree-of-freedom attitude estimation network; Performing image reconstruction based on the relative space target monocular six-degree-of-freedom posture change information and the current frame object depth information to obtain the absolute space target monocular six-degree-of-freedom posture change information of the adjacent frame; Based on the relative space target monocular six-degree-of-freedom posture change information and the absolute space target monocular six-degree-of-freedom posture change information of the adjacent frame, the absolute space target monocular six-degree-of-freedom posture change information of the current frame is determined.

2. The monocular six-degree-of-freedom pose estimation method for a space target according to claim 1, wherein: Before inputting the adjacent frame images of the current frame image and the current frame object depth information into a pre-trained spatial target six-degree-of-freedom pose estimation network, the method further includes: Obtain an initialized six-degree-of-freedom pose estimation network for spatial targets based on sequence images; Inputting the current frame image sample and the adjacent frame image sample into the space target six-degree-of-freedom attitude estimation network, obtaining the current frame mask sample, the current frame depth map sample, and the relative space target monocular six-degree-of-freedom attitude change information sample between the current frame and the adjacent frame output by the space target six-degree-of-freedom attitude estimation network; Determining a loss function of a mask network based on the current frame mask sample and the true mask; Inputting the current frame depth map sample and the relative space target monocular six-degree-of-freedom posture change information sample into the image reconstruction network to obtain a reconstructed current frame image sample; Determining a structural similarity loss and a norm loss based on the reconstructed current frame image sample and the current frame image sample as a reconstruction loss function; Determining a loss function of a deep pose estimation network based on the reconstruction loss function and the edge-aware smoothing loss function; Obtaining a total loss function of the space target six-degree-of-freedom pose estimation network based on weighting the loss function of the mask network and the loss function of the depth pose estimation network; When the depth pose estimation network reaches a preset number of training times or the loss function value of the total loss function is lower than a preset threshold, a pre-trained space object six-degree-of-freedom pose estimation network is obtained.

3. The monocular six-degree-of-freedom pose estimation method for a space target according to claim 1, wherein: The edge attention module is used to preprocess the input current frame image of the target spacecraft to obtain an image carrying edge information, including: Performing edge detection on the current frame image based on Sobel convolution using an edge attention module to obtain an edge amplitude map of the current frame image; Performing Fourier transform on the edge amplitude map through the edge attention module to obtain a frequency domain feature map of the current frame image; The edge attention module performs discrete cosine transform based on the frequency domain feature map to obtain an image carrying edge information.

4. The monocular six-degree-of-freedom pose estimation method for a space target according to claim 1, wherein: Inputting the image carrying edge information into a pre-trained mask segmentation network model to obtain binary mask information output by the mask segmentation network model includes: Extracting features from the image carrying edge information using a pre-trained mask segmentation network model to obtain features of different scales; Inputting the features of different scales into an edge-based Fourier frequency domain attention module through the pre-trained mask segmentation network model to obtain a weighted frequency domain feature map output by the edge-based Fourier frequency domain attention module; The weighted frequency domain feature map is decoded by the pre-trained mask segmentation network model to obtain the binary mask information of the current frame image.

5. The monocular six-degree-of-freedom pose estimation method for a space target according to claim 1, wherein: The step of inputting the current frame object image and the current frame image into a depth estimation network to obtain the current frame object depth information output by the depth estimation network includes: Encoding the current frame image through a depth estimation network to obtain a multi-scale feature map of the current frame; Performing continuous convolution on the current frame multi-scale feature map and the current frame object image through the attention gate feature selection module of the depth estimation network to obtain the current frame fusion feature; The current frame fusion feature is decoded by a depth decoder of the depth estimation network to obtain a current frame depth map, wherein the current frame depth map is used to represent the depth information of the current frame object.

6. The monocular six-degree-of-freedom pose estimation method for a space target according to claim 5, wherein: The step of inputting adjacent frame images of the current frame image and the depth information of the current frame object into a pre-trained space target six-degree-of-freedom pose estimation network to obtain relative space target monocular six-degree-of-freedom pose change information of the current frame and the adjacent frames output by the pre-trained space target six-degree-of-freedom pose estimation network includes: Encoding adjacent frame images of the current frame image through a pre-trained spatial target six-degree-of-freedom posture estimation network to obtain a multi-scale fusion image of adjacent frames; The multi-scale fusion map of the adjacent frames is fused with the multi-scale feature map of the current frame through the attention gate feature selection module of the pre-trained spatial target six-degree-of-freedom posture estimation network to obtain the fusion features of the current frame and the adjacent frames; The posture decoder of the pre-trained spatial target six-degree-of-freedom posture estimation network is used to decode the fusion features of the current frame and the adjacent frames to obtain the relative spatial target monocular six-degree-of-freedom posture change information between the current frame and the adjacent frames.

7. A monocular six-degree-of-freedom attitude estimation device for a space target, characterized in that: include: A preprocessing module is used to preprocess the input current frame image of the target spacecraft through the edge attention module to obtain an image carrying edge information; a first determination module, configured to input the image carrying edge information into a pre-trained mask segmentation network model, obtain binary mask information output by the mask segmentation network model, and determine the object image of the current frame based on the binary mask information; A first output module is configured to input the current frame object image and the current frame image into a depth estimation network to obtain the current frame object depth information output by the depth estimation network; A second output module is configured to input adjacent frame images of the current frame image and the depth information of the object of the current frame into a pre-trained spatial target six-degree-of-freedom attitude estimation network, and obtain relative spatial target monocular six-degree-of-freedom attitude change information of the current frame and the adjacent frames output by the pre-trained spatial target six-degree-of-freedom attitude estimation network; A reconstruction module is used to reconstruct an image based on the relative space target monocular six-degree-of-freedom posture change information and the current frame object depth information to obtain the absolute space target monocular six-degree-of-freedom posture change information of the adjacent frame; The second determination module is used to determine the absolute space target monocular six-degree-of-freedom posture change information of the current frame based on the relative space target monocular six-degree-of-freedom posture change information and the absolute space target monocular six-degree-of-freedom posture change information of the adjacent frame.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the monocular six-degree-of-freedom pose estimation method for a space target as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the monocular six-degree-of-freedom pose estimation method for a space target as claimed in any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the monocular six-degree-of-freedom pose estimation method for a space target as claimed in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Six-degree-of-freedom attitude estimation method and device and computer readable storage medium

    CN110119148A

  • Unmanned aerial vehicle ground target positioning method based on monocular depth estimation

    CN117455972A