Visual localization and attitude estimation method based on prior search in rocket recovery section

By using deep learning models and prior search methods, the problem of visual measurement under complex three-dimensional attitude changes during rocket recovery was solved, achieving high-precision and robust pose estimation that meets real-time requirements.

CN122089831APending Publication Date: 2026-05-26ORIENTAL SPACE TECH (SHANDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ORIENTAL SPACE TECH (SHANDONG) CO LTD
Filing Date
2026-02-02
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing rocket recovery visual measurement technology is insufficient in accuracy and has poor robustness when faced with complex three-dimensional attitude changes, and it fails to effectively integrate prior information to improve estimation accuracy and anti-interference ability.

Method used

A deep learning model is used to learn the mapping from images to low-dimensional feature vectors. A loss function is designed to keep the feature space consistent with the physical state space. Pose estimation is performed by combining prior search and efficient indexing algorithms.

Benefits of technology

It achieves high-precision and robust pose estimation under complex 3D attitude changes, meets real-time requirements, and improves measurement accuracy and system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122089831A_ABST
    Figure CN122089831A_ABST
Patent Text Reader

Abstract

The invention discloses a visual localization and attitude estimation method and device based on prior search in a rocket recovery section and a medium, and belongs to the technical field of spacecraft guidance, navigation and control, and the method comprises the steps: carrying out the systematic sampling of a rocket recovery tail end operation space containing a three-dimensional position and a three-dimensional camera attitude in advance; constructing a matching data set in which the rocket recovery state is matched with the image; learning mapping from an image to a low-dimensional feature vector by using a deep learning model, and enabling the structure of a feature space to be consistent with the structure of a physical state space through a designed loss function; and in the rocket recovery stage, pre-stored samples closest to real-time image features are retrieved, and state labels of the pre-stored samples are fused, so that pose estimation is realized. The method can adapt to complex three-dimensional attitude changes, is high in robustness, and can meet the real-time requirement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spacecraft guidance, navigation and control technology, and in particular to a method, device and medium for visual positioning and attitude estimation based on prior search in the rocket recovery stage. Background Technology

[0002] In the final stage of vertical rocket recovery, extremely high-precision, high-frequency measurements of its position and attitude relative to the predetermined landing point are required to support precise guidance and control. The combination of global satellite navigation systems and inertial navigation systems suffers from insufficient accuracy and limited update rate at the terminal stage. Therefore, vision-based measurement methods have become an important supplement or alternative.

[0003] Currently, existing visual measurement technologies for rocket recovery have the following main shortcomings: Most methods assume that the camera's optical axis is approximately perpendicular to the target plane (i.e., a small field of view) or only consider attitude changes in a single dimension. In actual recovery, the rocket may undergo significant lateral maneuvers and attitude adjustments, resulting in complex three-dimensional attitude changes in the camera's focal plane relative to the horizontal target plane, including pitch, yaw, and roll. Existing methods lack a systematic modeling and utilization of the mapping relationship between this complete three-dimensional attitude and image features.

[0004] Meanwhile, traditional methods heavily rely on the accurate detection, recognition, and geometric calculation of target patterns (such as concentric circles). This process is susceptible to interference from changes in lighting, image blurring, and local occlusion, exhibiting poor robustness and high computational complexity.

[0005] Furthermore, although the trajectory and attitude range of the rocket recovery stage are dynamic, they have clear physical boundaries, and integrated navigation can provide prior position and attitude information with clear error boundaries in real time. Existing methods have failed to systematically integrate this prior spatial and attitude constraint into the measurement model to improve the estimation accuracy and anti-interference capability.

[0006] Therefore, it is necessary to provide a new technical solution to solve the above problems. Summary of the Invention

[0007] To address at least one of the aforementioned technical problems, this application provides a visual localization and attitude estimation method for rocket recovery stage based on prior search, which can adapt to complex three-dimensional attitude changes, is robust, and meets real-time requirements.

[0008] A visual localization and attitude estimation method for rocket recovery stage based on prior search includes: Step S1: Systematically sample the rocket recovery end-of-life workspace containing 3D position and 3D camera attitude in advance to construct a paired dataset that matches the rocket recovery status with the image; Step S2: Use a deep learning model to learn the mapping from the image to the low-dimensional feature vector, and use a designed loss function to make the structure of the feature space consistent with the structure of the physical state space. Step S3, the rocket recovery stage, involves retrieving the pre-stored sample that most closely matches the features of the real-time image and fusing its state label to achieve pose estimation.

[0009] Optionally, step S1 includes: Define a six-dimensional state space for the rocket during the final recovery phase, and perform systematic, discretized sampling within this six-dimensional state space; For each sampled state point, an image of the predetermined target observed at that position and attitude is obtained through simulation or physical experiment, forming a paired dataset that matches the rocket recovery state with the image.

[0010] Optionally, the six-dimensional state space includes a three-dimensional position centered on the landing point and a three-dimensional attitude describing the camera focal plane relative to the horizontal target plane.

[0011] Optionally, step S2 includes: A deep convolutional neural network is used as an encoder to map an image into a fixed-dimensional feature vector. Deep convolutional neural networks are trained using the spatial contrast loss function; The trained encoding network processes all sample images to generate corresponding feature vectors, which are then associated with and stored with the original state labels to build a feature vector database, and an index structure is created for this database.

[0012] Optionally, when training a deep convolutional neural network using a spatial contrastive loss function, the mandatory requirements of the loss function include: Images that are spatially close should have a feature vector distance that is less than a set threshold. For images at the same location and with similar camera poses, the distance between their feature vectors should be less than a set threshold. For images with large differences in spatial location or camera pose, the feature vector distance should be greater than a set threshold.

[0013] Optionally, step S3 includes: During rocket recovery, the target images captured in real time by the airborne camera are preprocessed and then input into the trained encoding network to obtain real-time feature vectors. In the offline feature vector database, the approximate nearest neighbor search algorithm is used to quickly find the K nearest neighbor feature vectors that are most similar to the real-time feature vectors. Obtain the state labels corresponding to the K nearest neighbor feature vectors, and use a weighted average or interpolation algorithm to calculate the current three-dimensional position estimate and three-dimensional attitude estimate of the rocket based on the state labels of the K nearest neighbors and their similarity with the query vector.

[0014] Optionally, step S3 further includes: By utilizing the prior position and attitude information with clear error boundaries provided in real time by the rocket recovery stage integrated navigation, the search range in three-dimensional position and attitude space can be narrowed.

[0015] According to another aspect of this application, a computing device is also provided, comprising: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the aforementioned visual positioning and attitude estimation method for rocket recovery segment based on prior search.

[0016] According to another aspect of this application, a computer-readable storage medium is also provided, on which computer instructions are stored, which, when executed on a computer, cause the computer to perform the aforementioned visual positioning and attitude estimation method for the rocket recovery stage based on prior search.

[0017] Compared with the prior art, this application has at least the following beneficial effects: 1. This invention directly learns image feature representations under complex viewpoints and lighting conditions through a deep learning model, avoiding the fragile image feature extraction and geometric calculation steps in traditional methods. It has better tolerance for interference such as image blurring, uneven lighting, and local occlusion, thereby significantly improving the accuracy of pose estimation and the robustness of the system.

[0018] 2. This invention clearly and completely models the three-dimensional attitude of the camera relative to the target plane, and enables the network to distinguish image changes caused by different attitudes through spatial contrastive learning. Therefore, this method can effectively cope with large angle tilts and image rotations that may occur during rocket recovery, and its measurement accuracy in such scenarios far exceeds that of traditional methods based on simplified attitude models.

[0019] 3. The core operations in the online phase of this invention are feature extraction and vector retrieval, both of which are computationally efficient. Feature extraction is achieved through highly optimized neural network forward propagation, and vector retrieval can utilize efficient indexing algorithms (such as HNSW). This ensures that the entire pose estimation process can be completed in milliseconds, meeting the stringent requirement of a measurement update rate of over 100Hz for high-dynamic rocket control.

[0020] 4. The database constructed offline in this invention essentially encapsulates the prior constraints of the job space. The online retrieval process is equivalent to finding the optimal solution within these prior constraints, which naturally suppresses impossible estimation results and enhances the overall stability and reliability of the system. Attached Figure Description

[0021] The following sections will describe some specific embodiments of the invention in a detailed manner by way of example and not limitation, with reference to the accompanying drawings. The same reference numerals in the drawings denote the same or similar parts or portions. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings: Figure 1 This is a schematic diagram of the overall process of the present invention. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0023] like Figure 1 As shown, a visual localization and attitude estimation method for rocket recovery stage based on prior search includes: Step S1: Systematically sample the rocket recovery end-of-life workspace containing 3D position and 3D camera attitude in advance to construct a paired dataset that matches the rocket recovery status with the images.

[0024] Specifically, including: Step S11: Define the six-dimensional state space of the rocket in the final stage of recovery, and perform systematic and discretized sampling within this six-dimensional state space.

[0025] The six-dimensional state space includes a three-dimensional position centered on the landing point and a three-dimensional attitude describing the camera focal plane relative to the horizontal target plane.

[0026] The three-dimensional attitude of the camera's focal plane relative to the horizontal target plane can be described by an equivalent rotation matrix, or by Euler angles such as pitch, yaw, and roll.

[0027] Step S12: For each sampled state point, obtain the image of the predetermined target observed at that position and attitude through simulation or physical experiment to form a paired dataset that matches the rocket recovery state with the image.

[0028] Step S2: Use a deep learning model to learn the mapping from the image to the low-dimensional feature vector, and use a designed loss function to make the structure of the feature space consistent with the structure of the physical state space.

[0029] Specifically, step S2 includes: Step S21: Use a deep convolutional neural network as an encoder to map the image into a fixed-dimensional feature vector.

[0030] Step S22: Train the deep convolutional neural network using the spatial contrast loss function.

[0031] Step S23: Process all sample images with the trained encoding network to generate corresponding feature vectors, and store them in association with the original state labels to build a feature vector database and establish an index structure for the database.

[0032] In training deep convolutional neural networks using the spatial contrastive loss function, the mandatory requirements for the loss function are as follows: For images that are spatially close, the distance between their feature vectors should be less than a set threshold.

[0033] Images at the same location and with similar camera poses should have a feature vector distance that is less than a set threshold.

[0034] For images with large differences in spatial location or camera pose, the feature vector distance should be greater than a set threshold.

[0035] Step S3, the rocket recovery stage, involves retrieving the pre-stored sample that most closely matches the features of the real-time image and fusing its state label to achieve pose estimation.

[0036] Specifically, it includes: Step S31: During the rocket recovery process, the target images acquired in real time by the airborne camera are preprocessed and then input into the above-trained encoding network to obtain real-time feature vectors. Step S32: In the offline feature vector database, use the approximate nearest neighbor search algorithm to quickly find the K nearest neighbor feature vectors that are most similar to the real-time feature vectors; Step S33: Obtain the state labels corresponding to the K nearest neighbor feature vectors. Using a weighted average or interpolation algorithm, calculate the current three-dimensional position estimate and three-dimensional attitude estimate of the rocket based on the state labels of the K nearest neighbors and their similarity with the query vector.

[0037] Furthermore, after step S31 is completed and before step S32 is executed, the following steps are also included: By utilizing the prior position and attitude information with clear error boundaries provided in real time by the rocket recovery stage integrated navigation, the search range in three-dimensional position and attitude space can be narrowed.

[0038] Among them, in narrowing the search range in the three-dimensional position and attitude space, the search range is the pre-stored spatial sample encoding. Example

[0039] A Northeastern Sky (ENU) coordinate system is established with the center of the landing target as the origin O. The rocket's state space during the terminal guidance phase is a finite three-dimensional domain. Vertical altitude (Z): Along the celestial axis, ranging from the initial altitude of the guide segment to the landing point.

[0040] Horizontal position (X, Y): Within the horizontal plane, covering the expected lateral maneuver range. This space is discretized and sampled uniformly or non-uniformly in all three dimensions to form a three-dimensional position grid.

[0041] The camera is fixed to the arrow body, and its attitude is determined by the attitude of the arrow body. The target plane is defined as the horizontal plane (with the normal vector as the celestial axis). The attitude of the camera's focal plane (i.e., the imaging plane) is a three-dimensional attitude describing the rotational relationship between the camera coordinate system and the target coordinate system (or the horizontal plane coordinate system).

[0042] To fully and unambiguously describe this relative posture, the following two equivalent parameterization methods are adopted: Rotation matrix (R): A 3×3 orthogonal matrix that describes the rotation transformation from the target coordinate system to the camera coordinate system. It provides the most accurate mathematical description.

[0043] Euler angles (pitch θ, yaw ψ, roll φ): These are decomposed into three consecutive rotations about the principal axes according to a specific rotational sequence (e.g., ZYX). This description is more intuitive and easier to understand and analyze.

[0044] Pitch angle θ: affects the pitch of the camera's optical axis in the vertical plane.

[0045] Yaw angle ψ: affects the direction of the camera's optical axis in the horizontal plane.

[0046] Roll angle φ: affects the rotation of the camera around its optical axis, i.e. the degree of tilt of the horizon or target in the image.

[0047] Note: The attitude angles here describe the camera's attitude relative to the horizontal target plane, not the absolute attitude of the rocket body relative to the geodetic coordinate system. However, the two can be transformed using a matrix. This parameterization fully encompasses changes in optical axis pointing and in-plane rotation of the image, providing a complete geometric basis for the correlation between visual features and attitude.

[0048] System sampling is performed on the defined state space. For each discrete six-dimensional state point... In a simulation environment or ground experiment, images of the target observed by the camera at this position and attitude are acquired, either in a simulated environment or in actual practice. This allows for the construction of a large-scale paired dataset that covers the entire operational envelope. .

[0049] Design a deep coding network Input image Mapped to a low-dimensional feature vector This network typically uses a pre-trained convolutional neural network as its backbone.

[0050] The training objective is to make the feature space "isomorphic" to the physical state space: samples with similar states should have similar feature vectors in the feature space; samples with large state differences should have feature vectors that are far apart.

[0051] To address this, a spatial contrastive loss function is designed, with the following mandatory requirements: Images that are spatially close should have a feature vector distance that is less than a set threshold. For images at the same location and with similar camera poses, the distance between their feature vectors should be less than a set threshold. For images with large differences in spatial location or camera pose, the feature vector distance should be greater than a set threshold.

[0052] It should be noted that the threshold here can be determined based on expert experience.

[0053] Specifically, the spatial contrast loss function is designed as follows: Location proximity constraint: For sample pairs that are very close in spatial location (X, Y, Z), regardless of whether their poses are the same, their feature vectors are encouraged to be close. This forces the network to learn features that are sensitive to changes in location.

[0054] Pose proximity constraint: For sample pairs with similar camera poses (θ, ψ, φ) at the same location point, their feature vectors are encouraged to be similar. This forces the network to learn features that are sensitive to changes in viewpoint.

[0055] State dissimilarity exclusion: Since the rocket recovery section is highly sensitive to the vertical distance between itself and the target plane, for sample pairs that differ greatly in vertical distance, their feature vector distance is pushed further apart.

[0056] Through joint optimization, the network learns the feature vectors. It becomes a compact descriptor that simultaneously encodes information about the observer's position and perspective.

[0057] Online deployment consists of two phases: offline database building and online retrieval.

[0058] The offline feature database construction phase includes: using a pre-trained encoding network

[0059] Processing datasets Generate a feature vector set from all images. Each of them All of them are related to their true state tags Association. Construct an efficient vector index (e.g., using an HNSW graph).

[0060] The online real-time status retrieval phase includes: a. Encoding: Encoding the real-time acquired camera images Input the encoding network to obtain the query feature vector .

[0061] b. Retrieval: In the feature database Perform an approximate nearest neighbor search to find the nearest neighbor with... Most similar eigenvectors .

[0062] c. Solution: Obtain this The state labels corresponding to the nearest neighbors Due to the isomorphism between the feature space and the state space, these nearest neighbor states should also cluster near the real state in the physical space. This can be achieved by assigning state labels... By applying a weighted average (with weights related to feature similarity) or using a more refined local interpolation model, the six-dimensional state of the current image can be estimated. .

[0063] The method was validated in a simulation environment including high-fidelity visual rendering. By setting various flight trajectories (including large-scale horizontal maneuvers and attitude adjustments), the proposed method was compared with a visual pose measurement method based on traditional geometric calculations. Performance evaluation metrics included: Position estimation error: the root mean square error of the position in three-dimensional space.

[0064] Attitude estimation error: errors in pitch, yaw, and roll angles.

[0065] Real-time performance: The time required for single-frame image processing and state resolution.

[0066] The results show that the proposed method is significantly more accurate than traditional methods in challenging scenarios with large viewing angle tilt and image rotation, and has good prospects for engineering applications.

[0067] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0068] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in sequences other than those illustrated or described herein.

[0069] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A visual localization and attitude estimation method for rocket recovery stage based on prior search, characterized in that, include: Step S1: Systematically sample the rocket recovery end-of-life workspace containing 3D position and 3D camera attitude in advance to construct a paired dataset that matches the rocket recovery status with the image; Step S2: Use a deep learning model to learn the mapping from the image to the low-dimensional feature vector, and use a designed loss function to make the structure of the feature space consistent with the structure of the physical state space. Step S3, the rocket recovery stage, involves retrieving the pre-stored sample that most closely matches the features of the real-time image and fusing its state label to achieve pose estimation.

2. The visual localization and attitude estimation method for rocket recovery stage based on prior search as described in claim 1, characterized in that, Step S1 includes: Define a six-dimensional state space for the rocket during the final recovery phase, and perform systematic, discretized sampling within this six-dimensional state space; For each sampled state point, an image of the predetermined target observed at that position and attitude is obtained through simulation or physical experiment, forming a paired dataset that matches the rocket recovery state with the image.

3. The visual localization and attitude estimation method for rocket recovery stage based on prior search as described in claim 2, characterized in that, The six-dimensional state space includes a three-dimensional position centered on the landing point and a three-dimensional attitude describing the camera focal plane relative to the horizontal target plane.

4. The visual localization and attitude estimation method for rocket recovery stage based on prior search as described in claim 3, characterized in that, Step S2 includes: A deep convolutional neural network is used as an encoder to map an image into a fixed-dimensional feature vector. Deep convolutional neural networks are trained using the spatial contrast loss function; The trained encoding network processes all sample images to generate corresponding feature vectors, which are then associated with and stored with the original state labels to build a feature vector database, and an index structure is created for this database.

5. The visual localization and attitude estimation method for rocket recovery stage based on prior search as described in claim 4, characterized in that, When training deep convolutional neural networks using the spatial contrastive loss function, the mandatory requirements of the loss function include: Images that are spatially close should have a feature vector distance that is less than a set threshold. For images at the same location and with similar camera poses, the distance between their feature vectors should be less than a set threshold. For images with large differences in spatial location or camera pose, the feature vector distance should be greater than a set threshold.

6. The visual localization and attitude estimation method for rocket recovery stage based on prior search as described in claim 5, characterized in that, Step S3 includes: During rocket recovery, the target images captured in real time by the airborne camera are preprocessed and then input into the trained encoding network to obtain real-time feature vectors. In the offline feature vector database, the approximate nearest neighbor search algorithm is used to quickly find the K nearest neighbor feature vectors that are most similar to the real-time feature vectors. Obtain the state labels corresponding to the K nearest neighbor feature vectors, and use a weighted average or interpolation algorithm to calculate the current three-dimensional position estimate and three-dimensional attitude estimate of the rocket based on the state labels of the K nearest neighbors and their similarity with the query vector.

7. The visual localization and attitude estimation method for rocket recovery stage based on prior search as described in claim 6, characterized in that, Step S3 also includes: By utilizing the prior position and attitude information with clear error boundaries provided in real time by the rocket recovery stage integrated navigation, the search range in three-dimensional position and attitude space can be narrowed.

8. A computing device, characterized in that, include: The processor and the memory storing a computer program, which, when executed by the processor, performs the visual localization and attitude estimation method for the rocket recovery stage based on prior search as described in any one of claims 1 to 7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed on a computer, cause the computer to perform the visual localization and attitude estimation method for the rocket recovery stage based on prior search as described in any one of claims 1 to 7.