A method, system, device and medium for intelligent integrated navigation in deep space exploration
By combining deep Q-networks with unscented Kalman filtering, an intelligent integrated navigation method is used to optimize the state error and measurement error covariance by utilizing starlight angular distance and solar radial velocity information. This solves the problem of insufficient navigation accuracy in traditional methods and enables high-precision autonomous navigation for deep space probes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2026-04-03
Smart Images

Figure CN120521606B9_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep space exploration technology, and in particular to a method, system, device and medium for intelligent integrated navigation in deep space exploration. Background Technology
[0002] With the continuous development of space technology, humanity's footprint has gradually expanded from near-Earth space to distant deep space. Lunar exploration, Mars exploration, Jupiter exploration, and deep space exploration missions to the edge of the solar system are gradually being implemented. During long-distance interstellar travel, in order to accurately enter the influence sphere of a planet and conduct orbital and flyby explorations, the probe needs to gradually approach the planet and enter its capture region for capture, thus precisely entering an orbital trajectory. The capture phase is relatively short, requiring extremely high navigation accuracy; the accuracy of the capture directly affects the probe's execution of its planetary orbit or flyby exploration missions.
[0003] During the capture phase, the probe is very close to the planetary object, and is affected by perturbations from nearby planets and non-celestial gravitational forces such as solar radiation pressure. This affects the accurate determination of the noise statistical characteristics of the orbital dynamics state model and the observation model, ultimately making it difficult to obtain the optimal parameters for state estimation filtering. Traditional navigation methods use extended Kalman filtering or unscented Kalman filtering for noise statistics and information fusion processing of multiple measurement models. The relevant state noise and measurement noise covariance matrix parameters are predetermined and do not adjust in real time according to the probe's on-orbit operating environment. For deep space probes operating in orbit for extended periods, the orbital error gradually increases over time, and the position and velocity errors of the probe become increasingly larger, failing to meet the high-precision position and velocity requirements of deep space exploration probes. Summary of the Invention
[0004] The purpose of this invention is to provide a method, system, device and medium for intelligent integrated navigation of deep space exploration, which can solve the problem that the statistical analysis of noise and information fusion processing of multiple measurement models using extended Kalman filtering or unscented Kalman filtering cannot meet the high accuracy requirements of position and velocity of deep space exploration probes.
[0005] To address the aforementioned technical problems, embodiments of the present invention provide an intelligent integrated navigation method for deep space exploration, comprising the following steps:
[0006] Establish an orbital dynamics model for the probe during the terminal phase of its cruise when it explores planetary objects;
[0007] Based on the orbital dynamics model, a state model and a measurement model are established for the probe autonomously navigating a planetary object during the terminal phase of its cruise. The measurement model is based on the angular distance information between the probe and the planetary object and nearby satellite objects, as well as the radial velocity information of the sun.
[0008] The state error covariance of the state model and the measurement error covariance of the measurement model are obtained as the state input of the deep Q network. Multiple free action directions of the state error covariance and the measurement error covariance are used as the action output of the deep Q network to obtain the optimal state error covariance and the optimal measurement error covariance through the deep Q network. The action direction is used to indicate the adjustment direction of the state error covariance or the measurement error covariance.
[0009] Based on the optimal state error covariance and optimal measurement error covariance of the probe in real time at the end of the cruise phase, the probe's state is estimated to complete the probe's autonomous navigation of planetary objects at the end of the cruise phase.
[0010] Optionally, the orbital dynamics model of the detector is as follows:
[0011] ;
[0012] In the formula, r pj This is the position vector of the probe relative to the planetary object; r js This is the direction vector of a planetary object relative to the Sun; r ps Let be the position vector of the probe relative to the Sun; the first term on the right side of the above equation is the gravitational perturbation of the planetary body, and the second term is the gravitational perturbation of the Sun; Non-spherical gravitational perturbations of planetary bodies; This refers to minute perturbations of celestial satellites and slight disturbances in complex environments.
[0013] Optionally, the state model of the detector is established through the following steps:
[0014] Convert the orbital dynamics model to:
[0015] ;
[0016] Based on the transformed orbital dynamics model, the state model is obtained as follows:
[0017] ;
[0018] In the formula, This is the state vector at the end of the detector's cruise phase. r pj =[x pj , y pj , z pj ] T and v pj =[ v pjx , v pjy , v pjz ] T These represent the position and velocity vectors of the probe as it approaches the planetary object; ω pj To account for solar pressure and unmodeled perturbations of other planets.
[0019] Optionally, the measurement model of the detector includes a first measurement model based on the starlight angular distance information between the detector and planetary objects and satellite objects near the planetary objects, and a second measurement model based on the solar radial velocity information;
[0020] The first measurement model is as follows:
[0021] ;
[0022] In the formula, For the observation of the starlight angular distance of the navigation system in the terminal phase of cruise; h S A This is the measurement function for astronomical navigation using starlight angular distance; V SA To reduce the observation noise of the system that uses planetary sensors to measure starlight angular distance information;
[0023] The second measurement model is as follows:
[0024] ;
[0025] In the formula, h RV The navigation measurement function for the terminal phase of cruise is obtained using solar radial velocity information; V RV The observation noise introduced by the measurement of solar radial velocity information.
[0026] Optionally, the state error covariance of the state model includes the position error covariance and the velocity error covariance, and the measurement error covariance of the measurement model includes the starlight angular distance error covariance and the solar radial velocity error covariance.
[0027] The action output of the deep Q-network includes two action directions: forward and backward for the position error covariance, two action directions: forward and backward for the velocity error covariance, two action directions: forward and backward for the starlight angular distance error covariance, and two action directions: forward and backward for the solar radial velocity error covariance.
[0028] Optionally, the deep Q-network employs dual unscented Kalman filters. The primary unscented Kalman filter estimates the detector's state based on the position error covariance and velocity error covariance, while the secondary unscented Kalman filter estimates the detector's state based on the starlight angular distance error covariance and the solar radial velocity error covariance. Valuable position error covariance, velocity error covariance, starlight angular distance error covariance, or solar radial velocity error covariance are immediately rewarded based on the state estimation results of the primary and secondary unscented Kalman filters. This reward is used to update the network parameters of the deep Q-network to obtain the optimal state error covariance and the optimal measurement error covariance.
[0029] Embodiments of the present invention also provide a deep space exploration intelligent integrated navigation system, comprising:
[0030] The first model building module is used to build an orbital dynamics model of the probe during the terminal phase of its cruise when it explores planetary objects.
[0031] The second model building module is used to establish a state model and a measurement model of the probe autonomously navigating a planetary object during the terminal phase of its cruise, based on the orbital dynamics model. The measurement model is established based on the starlight angular distance information between the probe and the planetary object and satellite objects near the planetary object, as well as the solar radial velocity information.
[0032] The parameter optimization module is used to obtain the state error covariance of the state model and the measurement error covariance of the measurement model as the state input of the deep Q network, and to take multiple free action directions of the state error covariance and the measurement error covariance as the action output of the deep Q network, so as to obtain the optimal state error covariance and the optimal measurement error covariance through the deep Q network; wherein, the action direction is used to indicate the adjustment direction of the state error covariance or the measurement error covariance.
[0033] The detection and navigation module is used to estimate the state of the probe based on the optimal state error covariance and optimal measurement error covariance in real time at the end of the cruise phase, so as to complete the probe's autonomous navigation of planetary objects at the end of the cruise phase.
[0034] Embodiments of the present invention also provide a computer device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute the above-described intelligent integrated navigation method for deep space exploration.
[0035] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described intelligent integrated navigation method for deep space exploration.
[0036] The intelligent integrated navigation method for deep space exploration provided by this invention has at least the following beneficial effects:
[0037] By establishing an orbital dynamics model for the probe during the terminal phase (capture phase) of its exploration of planetary objects, a state model and a measurement model are created for the probe's autonomous navigation of the planetary object during this phase. The measurement model is constructed by combining the angular distance information between the probe and the planetary object and its nearby satellites, as well as the solar radial velocity information, to more accurately characterize the probe's current state. The combination of the state model and the measurement model participates in the subsequent intelligent navigation of the probe, enabling precise state estimation and improving navigation accuracy. Simultaneously, by using the state error covariance of the state model and the measurement error covariance of the measurement model to reflect the current state environment of the probe, a deep Q-network is used to optimize them, obtaining the optimal state parameters of the probe (i.e., state error covariance or measurement error covariance). This can compensate for errors formed by the probe over long-term operation, ensuring the accuracy of the probe's position and velocity in deep space exploration, thereby improving the probe's navigation accuracy. Attached Figure Description
[0038] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:
[0039] Figure 1 A schematic flowchart of an intelligent integrated navigation method for deep space exploration provided by the present invention;
[0040] Figure 2 A schematic diagram of starlight angle measurement between a navigation celestial body and a star during the terminal phase of cruise, provided by the present invention;
[0041] Figure 3 A schematic diagram illustrating the measurement of solar radial velocity information provided by this invention;
[0042] Figure 4 A schematic diagram of a deep neural network structure provided by the present invention;
[0043] Figure 5 A schematic diagram of the learning mechanism of a Q-network provided by this invention.
[0044] Figure 6 This is a schematic diagram illustrating the basic principle of the DQN algorithm provided by the present invention;
[0045] Figure 7 A schematic diagram of an experience playback mechanism provided by the present invention;
[0046] Figure 8 This invention provides a schematic diagram of the principle of unscented transformation Sigma sampling points;
[0047] Figure 9 This invention provides a depth-based Q A schematic diagram of the reinforcement learning mechanism of intelligent agents in a network;
[0048] Figure 10 A schematic diagram of motion space flow provided by the present invention;
[0049] Figure 11 This is a schematic diagram of a deep neural network structure for DQNUKF provided by the present invention. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0051] The technical solutions provided by the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0052] One embodiment of the present invention relates to an intelligent integrated navigation method for deep space exploration. The specific process of the intelligent integrated navigation method for deep space exploration in this embodiment can be as follows: Figure 1 As shown, it includes:
[0053] Step 101: Establish an orbital dynamics model of the probe during the terminal phase of its cruise when it explores planetary objects.
[0054] Step 102: Based on the orbital dynamics model, establish the state model and measurement model of the probe autonomously navigating the planetary object during the terminal phase of its cruise. The measurement model is established based on the starlight angular distance information between the probe and the planetary object and the satellite objects near the planetary object, as well as the solar radial velocity information.
[0055] Step 103: Obtain the state error covariance of the state model and the measurement error covariance of the measurement model as the state input of the deep Q network, and take multiple free action directions of the state error covariance and the measurement error covariance as the action output of the deep Q network, so as to obtain the optimal state error covariance and the optimal measurement error covariance through the deep Q network; wherein, the action direction is used to indicate the adjustment direction of the state error covariance or the measurement error covariance.
[0056] Step 104: Based on the optimal state error covariance and optimal measurement error covariance of the probe in real time at the end of the cruise phase, the probe's state is estimated to complete the probe's autonomous navigation of the planetary object at the end of the cruise phase.
[0057] The following is a detailed description of the implementation details of the deep space exploration intelligent integrated navigation method in this embodiment. The following content is only for the convenience of understanding and is not necessary for implementing this solution.
[0058] During long-distance, sustained interstellar cruises in the space environment, the position and velocity of space probes will deviate significantly over time. To provide real-time and accurate navigation of the probe's operational status during the terminal phase of its cruise, and to reduce the impact of environmental variability, perturbations, and accumulated errors on the probe's operational accuracy, it is necessary to analyze the principles of autonomous navigation during the terminal phase of the cruise, providing a theoretical basis for the probe to achieve autonomous navigation.
[0059] In step 101, the orbital dynamics model of the probe during the terminal phase of its cruise is first explained in detail:
[0060] Before a spacecraft approaching a planetary body and utilizing its gravitational pull for propulsion, it needs to enter the planet's gravitational influence zone. During this period, the spacecraft must perform orbital maneuvers to successfully enter a hyperbolic orbit for this purpose; this phase is known as the terminal phase of the spacecraft's cruise. Before entering the planet's gravitational influence sphere, the spacecraft will be subjected to gravitational forces from the planet, the Sun, perturbations from nearby satellites, and disturbances in a complex environment. To accurately calculate the spacecraft's operational state, all gravitational perturbations and disturbances must be considered, resulting in a very complex orbital dynamics model. For the terminal phase, since the spacecraft is relatively close to the planet, we only consider the gravitational perturbations from the Sun's center and the planet's center. Other gravitational disturbances have a relatively small impact and are not considered further. The terminal phase orbital dynamics model is simplified to a circular three-body model. This simplifies the model and allows for a relatively accurate depiction of the spacecraft's motion during this phase. Therefore, the circular restricted three-body orbit dynamics model will be used as the orbit dynamics model for the autonomous navigation of the Solar System Edge Probe during the terminal phase of its cruise.
[0061] Based on spacecraft orbital dynamics, a state model for the terminal phase of the solar system boundary probe's cruise is established. As the probe gradually approaches the planetary body, its operation is considered as a scenario centered on the planetary body. Its kinematic equations are referenced to the planetary body's central inertial coordinate system, and it is assumed that the planetary body revolves around the Sun in approximately uniform circular motion. Therefore, the orbital dynamics model for the terminal phase of the solar system boundary probe's cruise in the planetary body's central inertial coordinate system (J2000) is as follows:
[0062] ;
[0063] In the formula, r pj This is the position vector of the probe relative to the planetary object; r js This is the direction vector of a planetary object relative to the Sun; r ps Let be the position vector of the probe relative to the Sun; the first term on the right side of the above equation is the gravitational perturbation of the planetary body, and the second term is the gravitational perturbation of the Sun; For non-spherical gravitational perturbations of planetary objects, only the terminal phase of the cruise needs to be considered. J Two items are sufficient; Minor perturbations of celestial satellites and slight disturbances in complex environments can be considered as error terms.
[0064] In practical navigation measurement model calculations, vectors are typically converted to component forms. The orbital dynamics model of the probe's terminal cruise phase can be described as follows:
[0065] .
[0066] In step 102, based on the established orbital dynamics model of the terminal phase of the solar system boundary probe's cruise, the autonomous navigation state model of the celestial body during the terminal phase of the cruise can be obtained as follows:
[0067] ;
[0068] In the formula, This is the state vector at the end of the detector's cruise phase. r p j =[ x pj , y pj , z pj ] T and v p j =[ v pj x , v pjy , v pjz ] T These represent the position and velocity vectors of the probe as it approaches the planetary object; ω pj To account for solar pressure and unmodeled perturbations from other planets, these are considered as calculation errors in the system model during the terminal phase of cruise, and are approximated as having covariance. Q pj Zero-mean white noise.
[0069] Regarding the measurement models for the probe autonomously navigating planetary objects during the terminal phase of its cruise, there are two models: a first measurement model based on the angular distance information between the probe and the planetary object and its nearby satellite objects, and a second measurement model based on the solar radial velocity information.
[0070] First, let's discuss the navigation principle of star-angle distance information:
[0071] During the final stages of their cruise, the outer reaches of the solar system probe come close to planetary bodies. At this point, the distance to Earth is vast, making real-time state estimation impossible for ground-based navigation and measurement equipment. Therefore, a high degree of autonomy is required of the probe. As it approaches the planetary body, the probe enters the final stage of its cruise orbit, where it will attempt to capture the planet and utilize the gravitational slingshot effect for propulsion. This final stage places even higher demands on navigation accuracy and real-time performance.
[0072] For the final phase of the cruise mission aimed at exploring the edge of the solar system, the Earth-Jupiter cruise segment uses Jupiter as the target celestial body to ensure the probe can accurately enter the next stage of the mission. For example... Figure 2 The diagram shows the measurement of the starlight angle between the navigation object and the star during the terminal phase of the cruise. The starlight angular distance information of the planetary object, Jupiter I, and Europa is used as the observation information for the autonomous astronomical navigation during the terminal phase of the cruise.
[0073] The autonomous navigation during the final stage of the solar system boundary exploration cruise uses the angular distances between the probe and planetary objects, as well as between the probe and Io and Europa, as measurements. θ pj , θ p lo and θ peu These measurements, taken by the planetary sensors on the probe, represent the angles between the vectors pointing to the planets, Io, and Europa, and the vector pointing to the probe's star. The model for measuring starlight angular distance information at the end of the cruise phase is as follows:
[0074] ;
[0075] In the formula, r pj , r plo and r peu These are the position vectors of the probe relative to the planetary object, Io, and Europa, respectively. s 1. s 2 and s 3 represents the vector direction of starlight from planetary objects, Io, and Europa in the J2000.0 heliocentric inertial coordinate system, i.e., the starlight vector between the probe and planetary objects, Io, and Europa. The right ascension and declination of stars can be found from the star catalog using star identification. , and These are the measurement noise errors between the navigation star and planetary objects, and between Io and Europa, respectively. r lo and r eu These are the position vectors of Io and Europa relative to the Sun, respectively. Indicates the size of the vector.
[0076] The celestial navigation measurement equation (first measurement model) for the terminal phase of cruise using starlight angular distance information is as follows:
[0077] ;
[0078] In the formula, For the observation of the starlight angular distance of the navigation system in the terminal phase of cruise; h S A This is the measurement function for astronomical navigation using starlight angular distance; V SA This is to reduce the observation noise of the system used to measure starlight angular distance information using planetary sensors.
[0079] Next, the navigation principle based on solar radial velocity information will be explained:
[0080] Given the greater influence of gravitational perturbations from the Sun and the captured target object during the final stages of the solar system boundary cruise, solar radial velocity information is observed using a solar spectrometer. The frequency shift of the solar radial velocity can be obtained by utilizing the spectral shift of sunlight caused by the relative motion between the probe and the Sun. A schematic diagram of the solar radial velocity measurement is shown below. Figure 3 As shown, the measurement model for solar radial velocity information can be described as follows:
[0081] ;
[0082] In the formula, This is the position vector of the probe relative to the Sun's center of mass; Noise is used to measure radial velocity information.
[0083] The measurement equation (second measurement model) for the solar radial velocity information at the end of the probe's cruise is as follows:
[0084] ;
[0085] In the formula, h RV The navigation measurement function for the terminal phase of cruise is obtained using solar radial velocity information; V RV The observation noise introduced by the solar radial information measurement can be approximated as zero-mean Gaussian white noise during the modeling process, and its variance is given according to the measurement accuracy of the spectrometer.
[0086] In step 103, during long-distance interstellar travel, the Solar System Edge Probe needs to gradually approach the planetary body during its final cruise phase in order to accurately enter the influence sphere of the planetary body for leverage, ensuring the probe can continue flying with limited fuel, or to conduct orbital and flyby explorations of planets during its Solar System Edge mission. The final cruise phase is relatively short, demanding higher navigation accuracy, which directly affects the probe's effectiveness in leveraging planetary influences and its orbital or flyby explorations. Since the probe gradually approaches the planetary body requiring leverage during the final cruise phase, the planetary body and its nearby satellites are ideal navigation and observation objects. Furthermore, the probe is closer to the planetary body during the final cruise phase, and is affected by perturbations from nearby planets and non-celestial gravitational forces such as solar radiation pressure perturbations. This affects the accurate determination of the noise statistics of the orbital dynamics state model and observation model, ultimately making it difficult to obtain the optimal parameters for state estimation filtering. High-precision navigation of the probe during long-term cruise is achieved based on Q-learning. The learned parameters are stored in the form of tables. However, due to the limited scale of the stored parameters, it is not conducive to the trial and error optimization of the many and complex state parameters during the high-precision navigation of the probe in orbit. It will cause the curse of dimensionality of the learned parameters and the sparsity of the state parameters.
[0087] To address the complex and variable operating environment of the terminal phase of a solar system edge probe's cruise, this paper studies a high-precision intelligent integrated navigation method for the terminal phase. First, a Deep Q-Network (DQN) algorithm is introduced for the subsequent optimization strategy design of navigation state parameters. Then, an Unscented Kalman Filter (UKF) is introduced to avoid the truncation error caused by the Extended Kalman Filter (EKF) removing higher-order terms, thereby improving the accuracy of state estimation for nonlinear systems. Next, based on the terminal phase dynamics model and related measurement models, a DQNUKF intelligent integrated navigation method based on star distance / solar radial velocity information is proposed. The state error variance and the observation error variance of the two measurement systems are used as the state input of the DQN. A state-action space is defined to construct a deep convolutional neural network for optimizing state parameters. The reward mechanism of the DQN is designed through state estimation UKF and provisional UKF (i.e., main UKF and secondary UKF). The deep convolutional neural network is trained, and the filter parameters are tuned through periodic iteration using an empirical replay mechanism to verify the accuracy of the proposed intelligent integrated navigation system. Then, to ensure high accuracy and stability of intelligent navigation state estimation in the terminal phase, state parameters, action space, and reward mechanism are defined for the deep Q-learning network. A DQNUKF sub-filter and main filter based on star distance / dual-difference pulse arrival time and solar radial velocity are designed for information fusion. A high-precision intelligent integrated navigation method based on DQNUKF is proposed to meet the high-precision requirements of autonomous intelligent navigation in the terminal phase of solar system boundary exploration cruises.
[0088] Regarding the deep Q-network algorithm:
[0089] The Deep Q-Network (DQN) algorithm is a deep reinforcement learning model built upon Q-learning. It integrates deep convolutional neural networks (CNNs) with Q-learning, using CNNs to construct the evaluation Q-function. This allows for trial and error in a wider state space to obtain the optimal policy, overcoming the limitation of Q-learning, which relies solely on table-based representations and struggles with high-dimensional continuous states and action spaces. The DQN algorithm perfectly combines the powerful decision-making capabilities of reinforcement learning with the representational advantages of deep learning, giving it the robust learning ability of reinforcement learning while inheriting deep learning's capacity to represent high-dimensional continuous states in complex environments. The DQN algorithm replaces the finite Q-table (state-action-value) evaluation function with a neural network, fusing deep learning and reinforcement learning to achieve end-to-end learning from information fusion state perception to action selection between the detector agent and the observer.
[0090] The DQN algorithm uses a deep neural network to fit the Q-value evaluation function. Based on the input state observed by the detector agent, it outputs the Q-value for each state-action pair through multi-layer deep learning. Deep learning processes the data from the detector agent to perceive information about the surrounding environment and its own state, enabling the design of the reward mechanism and the modeling of the Q-value evaluation function. The core of deep learning-based DQN is the deep neural network. Using a deep neural network only requires some necessary preprocessing of the raw data; the relevant state parameters are directly input into the neural network, and the weights are automatically adjusted through training to learn the optimal Q-value—this is called end-to-end learning. In a trained deep neural network model, the output of the lower hidden layers is closer to the features of the original input state. As the number of neural network layers increases, more and more feature information is learned, and its feature representation ability becomes stronger. The more hidden layers in a deep neural network, the better the training effect; deep neural network learning with multiple hidden layers is called deep learning. In the DQN algorithm, a deep convolutional neural network (CNN) is used to construct the Q-learning network. CNN has a strong ability to represent the original input state data. The corresponding parameters observed by the detector agent are used as the state input network. After passing through the network convolutional layer, pooling layer, fully connected layer and classification layer, the discrete action value function is finally output. The agent selects the optimal action corresponding to the current state according to the action value function.
[0091] The structure of a deep neural network is as follows: Figure 4 As shown, the input layer of a convolutional neural network can process multidimensional data. The input layer typically has multiple channels, and the state data, as input features, needs to be pre-standardized to improve the network's learning efficiency. Convolutional layers, pooling layers, and fully connected layers are all hidden layers. The convolutional layer extracts features from the input state data. It contains many convolutional kernels. A single convolutional kernel can reduce the number of channels while maintaining the input feature data, thus reducing the computational cost of the convolutional layer. A multi-layer perceptron (MLP) CNN network with parameter sharing can be constructed using a single convolutional kernel. The activation functions within the convolutional layer are used to express complex features; commonly used activation functions include ReLU and the sigmoid function.
[0092] ;
[0093] In the formula, This is the output of the current layer of the network's convolutional layer; This is the output of the previous layer; To convolve the kernel of the current layer; M jSelected input feature parameters; The bias parameters for the current layer of convolution.
[0094] After feature extraction by the convolutional layer, the output feature data enters the pooling layer for sampling. The purpose of this sampling is to extract features and filter relevant information from the output data. Pooling networks have many different forms of non-linear pooling functions (including max pooling and average pooling), which act on all input feature parameters to pool their size. Pooling does not change the input features or the number of parameters; it only reduces the dimensionality of the features to decrease the input size and computational cost of the next layer. Pooling sampling specifically involves:
[0095] ;
[0096] In the formula, This is the sampling function for the pooling layer.
[0097] Furthermore, pooled convolutional layers cannot be directly connected to dense fully connected layers. The resulting data needs to be flattened and compressed before being fed into the fully connected layers. The fully connected layers, located after the hidden layers of the convolutional network, calculate the output value for each feature obtained from the preceding convolutional layers.
[0098] In the DQN algorithm, the deep neural network used for iterative learning to continuously approximate the Q-value function is the Q-network. Its input is the agent's state S, and its output is the Q-value of all possible actions corresponding to all states. The Q-network is well-suited for high-dimensional state space scenarios. The learning mechanism of the Q-network is as follows: Figure 5 As shown. In practical applications, the agent continuously interacts with the environment, and the efficiency and results of learning Q-values are generally good. However, as the Q-network learns, even with a reasonable step size factor, poor results can occur because unstable states can arise in the latter part of the learning process. Unlike the tabular approach in Q-learning, which updates the state-action pairs in real time, each update of a Q-network node significantly impacts the policy distribution, causing changes in the training parameter distribution. Furthermore, overfitting is inherent in neural network training, making it difficult to obtain interaction parameter values that accurately reflect global environmental information.
[0099] To address this, DQN employs a "dual-network" structure consisting of an online current prediction network and a target value network to optimize learning performance. The target value network addresses the issue of target changes in the DQN training parameter model. Its network structure is consistent with the prediction network, but its parameter updates are slower. This network is used to obtain the target Q-value during the input sample state training process. The dual-network structure reduces the difference between the predicted Q-value and the target Q-value caused by the continuous real-time updates of the prediction network parameters, ensuring stability throughout the training process. The experience replay mechanism addresses the data correlation problem of training parameters. This mechanism allows for uniform sampling, reducing the data correlation between experience samples used in training the neural network. The experience gained by the detector agent through continuous interaction with its environment is stored in a dynamic data buffer. During each neural network training process, a batch of datasets is randomly selected from this buffer to update the network parameters. This aims to eliminate temporal dependencies between sample data, increase the diversity of sample data, and improve the learning efficiency of the neural network by leveraging the learned experience.
[0100] The entire training network of the DQN algorithm mainly consists of an online current prediction network, a target value network, an error function, and a playback memory storage unit, etc. Figure 6 The diagram illustrates the network training principle of the DQN algorithm. DQN utilizes Bellman's formula combined with Q-learning, employing an approximate solution method to train a neural network containing network model parameters θ. During each forward computation, the current prediction network outputs the Q-estimates of all possible actions for that state, selecting the action corresponding to the largest Q-estimate as the optimal action. This process iteratively updates the Q-estimates of the state-action pairs, using the output approximate action Q-estimate function to continuously approximate the Q-estimate function of the actual action. During training, at each interval... t Step will break down experience fragments Stored in the playback memory unit's experience buffer pool During the parameter update process, the following will be performed: Sampling is performed. (Contains parameters) θ The value function fitted by the deep neural network can be expressed as:
[0101] ;
[0102] In the formula, Used to represent approximate action value functions It can also be written as Alternatively, one can use... Represents the approximate state value function It can also be written as .
[0103] based on Q -learning can update the action-value function of the DQN algorithm, resulting in:
[0104] ;
[0105] In the formula, State-action value for taking a specific action for the current state; The maximum state-action value obtained by taking an action for the next state.
[0106] DQN algorithm built Q The network loss function is as follows:
[0107] ;
[0108] In the formula, θ for Q Network parameters; For the goal Q value; For prediction Q value; The objective is time difference.
[0109] The goal of the above formula Q Values and Predictions Q The values all contain parameters. θ Then the loss function is solved using the semi-gradient method. The gradient is:
[0110] ;
[0111] In practice, the Mini-Batch Semi-Gradient Descent (MBSGD) method is typically used to update gradient parameters. During the update process, the gradient parameters are selected... M The semi-gradient can be obtained by sampling and evaluating each sample:
[0112] ;
[0113] By using randomly selected sample data, the parameters of the neural network model in the next time step can be updated using the stochastic gradient descent (SGD) algorithm. θ t +1 :
[0114] ;
[0115] From the above equation, the optimal policy function of the DQN network can be obtained as:
[0116] ;
[0117] The DQN algorithm uses a reinforcement learning strategy to estimate the optimal state-action pair function. This can be obtained from the following Bellman optimal equation:
[0118] ;
[0119] The above analysis shows that the current forecast Q Values and Objectives Q The network model and parameters used are the same, in During the process of growing larger, The error will also increase because the state sample data itself contains inherent instability, which can cause fluctuations during parameter learning and updates, leading to oscillations or even divergence in the neural network model. To address this, DQN introduces two CNN network models for parameter learning and updates: one for prediction and the other for dynamic programming. Q Network Model and target Q Network Model The former is used to evaluate the current state-action value, while the latter is used to obtain the target value. The loss function and neural network model parameters of the "dual-network" structure are calculated below. θ The semi-gradient is:
[0120] ;
[0121] ;
[0122] The DQN algorithm, after introducing the target network, can guarantee the target network for a considerable period of time. Q The value remains unchanged, which weakens the predictability and the target. Q The close correlation between the two values is also reduced, and the oscillations and divergences caused by the loss value of the neural network model during training may be reduced, thus improving the stability of the entire algorithm during the learning process.
[0123] like Figure 7 As shown, the experience playback mechanism of the playback memory storage unit is to store fragmented experience samples obtained by the detector agent through information interaction with the surrounding environment at various times. Cache to experience buffer pool Inside, after multiple training steps, batches are randomly selected. A set of sample parameters of a specific size is used as a discrete set of numbers in the prediction. Q Value network, and then use MBSGD to adjust the parameters of the neural network model. θ Update.
[0124] Regarding the unscented Kalman filter navigation algorithm:
[0125] For navigation filtering algorithms of deep-sea space probes, the Extended Kalman Filter (EKF) is widely used. It obtains a linearized system model by linearly transforming the nonlinear system's state model to perform real-time state estimation of the probe. However, the state system in the terminal phase of a probe's cruise is more nonlinear than that in the long-term cruise phase. The EKF inevitably introduces some linearization errors during the linearization process, and solving the Jacobian matrix corresponding to the state and observation equations is difficult. Ignoring second-order and higher-order terms in Taylor series expansions can lead to significant errors in the filtered state estimation model, greatly affecting the accuracy of the filtering estimate. Furthermore, the filtered state estimate is prone to divergence when the system's nonlinearity is severe. In addition, the EKF algorithm involves complex partial differential equation solutions for the state and measurement equations during approximate linearization. When the system model is complex, the computational order is high, and the operating environment is rapidly changing, solving the Jacobian matrix becomes extremely complex. The computation also places strict requirements on the nonlinear system, requiring it to be continuously differentiable. This is another reason for the limitations of the EKF algorithm.
[0126] The Unscented Kalman Filter (UKF) algorithm combines lossless transform (UT) with standard Kalman filtering. It selects sampling points from the sequence to approximate the posterior probability density of the state and avoids linearization by fitting nonlinear equations, thus further improving the filter's estimation accuracy for nonlinear systems. Specifically, UKF, based on lossless transform, trains the sequence sampling points using stochastic gradient descent to obtain a continuously updated model that can estimate the system's state space. In each update process, the selected samples undergo an update calculation using the unscented transform. By processing the sampled data, a state estimate related to the actual operating state is obtained. The transformed sampled data is then applied to the nonlinear state model of the cruise terminal trajectory dynamics, and recalculated in conjunction with the measurement model, resulting in higher state estimation accuracy.
[0127] (1) Unscented transformation:
[0128] The unscented transform is a method that uses a deterministic sampling strategy to approximate the probability density distribution of the state of a nonlinear system, and utilizes nonlinear transformations to approximate the statistical distribution of random variables. The specific principle of the unscented transform is as follows: Figure 8 As shown, the key to this transformation is... n The random variable of dimension is chosen to have its mean and covariance equal to the mean of the original state distribution. Covariance is Using identical special sample points, each selected Sigma point is substituted into a nonlinear state function for nonlinear transformation, resulting in a transformed set of sample points. Weighted calculations using these points yield the transformed posterior mean and covariance. Furthermore, unlike the EKF which uses a second-order Taylor expansion, the sample points in the unscented transformation are not linearized, thus avoiding the loss of higher-order terms. This is why the UKF's estimation accuracy is higher than that of the EKF.
[0129] In the sampling process of the UT transform, a set of sampling Sigma points needs to be selected according to specific rules. These sampling points can be calculated using the mean and covariance of the original distribution. When the sampling Sigma points are selected, the weighted sample mean is equal to the mean and covariance of the original distribution. Next, several Sigma points are selected from the sampling point set as new sample points. These points are then projected onto the original distribution and substituted into the nonlinear state function to obtain the point set of the corresponding nonlinear state function values. Then, the mean and covariance after the UT transform can be obtained from these point sets. For conventional nonlinear state systems, the parameters obtained after the UT transform are differentiable linear functions or power series. For nonlinear state systems of any form, the transformation of state variables following a Gaussian distribution, using a selected set of Sigma sampling points, can yield very accurate posterior mean and covariance of the third-order matrix.
[0130] For nonlinear transformation functions y = f ( x Considering the probabilistic statistical characteristics of noise and its application in the actual operating environment at the end of the detector's cruise phase, a mean input point for sampling needs to be added. Meanwhile, the weighting coefficients of the Sigma sampling points are adjusted to optimize the performance of UKF, as follows:
[0131] ;
[0132] In the formula, The above formula can also be expressed as:
[0133] , ;
[0134] The weighting coefficients corresponding to the Sigma sampling points are calculated as follows:
[0135] ;
[0136] In the formula, For the first matrix i List; For the UT transformation iOne Sigma sampling point; The size depends on the distribution of the state vector; The scaling parameter is used to determine the accuracy of the estimate, with the aim of reducing the prediction error of the UKF. To characterize the Sigma sampling points at the corresponding mean The surrounding distribution, through analysis Adjustments were made to change the Sigma sampling points relative to the mean. The distance, whose range is . The aim is to reduce the impact of nonlocality caused by strong nonlinear environments; As a scaling factor, it is adjusted to make It is a positive semi-definite matrix. n s The dimension of the state; As another scaling factor, changing its parameter value has a good effect on estimating the propagation accuracy of the output variance matrix; The weighting coefficients are the values of the mean. The weighting coefficients are the covariance values.
[0137] (2) UKF time updates:
[0138] For nonlinear process systems, depending on the specific navigation filter system, different moments within a specific time interval should be considered. k s If the process random variables of the state model X UKF The Gaussian white noise matrix is W ( k s ), the observed variables of the measurement model Z UKF The Gaussian white noise matrix is V ( k s The nonlinear state equations and measurement equations of the UKF are as follows:
[0139] ;
[0140] In the formula, This refers to the state transformation of a nonlinear equation. The observations are for the nonlinear measurement equation. The initial values for the filter are selected based on the state equation:
[0141] ;
[0142] Based on the above UT transformation, 2 was selected. n sWith +1 Sigma sampling points, combined with the above formula, the Sigma sampling points can be rewritten to obtain the UT transformation for state prediction:
[0143] ;
[0144] The above formula can be used to calculate 2. n s One-step prediction with a set of +1 Sigma sampling points:
[0145] ;
[0146] By weighted summing of the one-step predictions of these Sigma sampling point sets, we can calculate the one-step predicted states of the process variables in the nonlinear state equation and the corresponding one-step predicted covariance matrix:
[0147] ;
[0148] ;
[0149] In the formula, For process random variables X UKF Gaussian white noise W ( k s The covariance matrix of ).
[0150] (3) UKF measurement updates:
[0151] After completing the one-step state prediction calculation, applying the UT transform to the measurement equation to perform new Sigma sampling yields the set of Sigma sampling points for measurement prediction as follows:
[0152] ;
[0153] Substituting the Sigma sampling point set from the above equation into the observation equation yields the predicted measurement equation: ;
[0154] By weighted summing of the one-step predictions of the Sigma sampling point set of measurement predictions, we can calculate the one-step observation predictions of the measurement variables in the nonlinear observation equation and the corresponding prediction error variance matrix:
[0155] ;
[0156] ;
[0157] ;
[0158] In the formula, For the state vector of the nonlinear dynamics and the covariance matrix of the measured values of the observation equation; This is the measurement variance matrix of the observation equation; For observed variables Z UKF Gaussian white noise V ( k s The covariance matrix of ).
[0159] Calculate the nonlinear dynamic system at a certain moment k s The UKF filter gain is:
[0160] ;
[0161] The state vector and covariance matrix of the nonlinear dynamic system are updated in real time as follows:
[0162] ;
[0163] ;
[0164] Theoretical analysis reveals that UKF utilizes the UT transform to select a substantial set of Sigma sampling points and corresponding weighting coefficients. Compared to EKF, UKF avoids calculating the Jacobian matrix and a series of matrix multiplication operations, ensuring that the posterior mean and covariance are accurate to second order, and the covariance accuracy can even reach third order. Furthermore, it can be easily applied to state estimation of nonlinear systems without requiring linearization. Therefore, UKF offers higher computational speed and state estimation accuracy than EKF. Additionally, the system noise and measurement noise in the UKF algorithm are assumed to be Gaussian white noise; however, the noise parameters of the actual detector operating environment state model and measurement model cannot be guaranteed to be completely approximated as Gaussian white noise.
[0165] Therefore, we will combine deep reinforcement learning to improve the learning of noise parameters for navigation state estimation in the terminal phase of cruise.
[0166] The UKF intelligent integrated navigation method based on starlight angular distance / solar radial velocity information for the terminal phase of cruise is as follows:
[0167] The core of intelligent integrated navigation algorithm design for the terminal phase of cruise is to continuously improve and optimize the state estimation parameters, providing real-time, high-precision navigation for the detector during the terminal phase of cruise based on the operational state of the detector agent and the measurement parameters of the measurement model. Traditional filtered state estimation algorithms, despite continuous improvement and expansion, still have many shortcomings in the optimal tuning of system state estimation parameters, even with nonlinear filters (such as EKF, UKF, and adaptive filters). In fact, the parameter tuning process can be viewed as solving the problem of finding the optimal strategy. This can be combined with deep reinforcement learning to find the optimal value function. The current state of the agent is obtained from the environment in real time, and corresponding actions are selected based on the state. The Q-value function of the optimal action is obtained through fitting a deep neural network, and then the state-action pair value function guides the detector agent to explore the optimal strategy to find the optimal tuning parameters. The principle of agent deep reinforcement learning based on deep Q-networks is as follows: Figure 9 As shown.
[0168] When employing an optimal strategy based on deep reinforcement learning to estimate the detector's state error during the terminal phase of cruise, it is necessary to design the input state of the DQN algorithm, the actions taken during iterative learning, and the rewards obtained from executing those actions. The state represents the agent's real-time perception of environmental changes, statistically analyzing interference and noise from the environment, and also identifying the agent's own state in real time. For the terminal phase of solar edge exploration cruise, the detector agent's state includes the detector's operational status information, measurement information from the measurement sensors, and information about the highly dynamic and complex external environment. Actions are the optimal behaviors performed by the agent based on its perception of environmental changes and its own operational status, leading to a better operational state for the detector agent. Rewards provide evaluation metrics for the detector's chosen actions; positive rewards provide continuous positive incentives for the detector's operation, improving the detector agent's operational state.
[0169] (1) State space:
[0170] To enter the gravitational pull region of a planetary body's influence sphere, the probe agent needs to perform precise capture before the gravitational pull occurs after the terminal phase of its cruise. During the terminal phase, the probe's operational status is affected by highly dynamic and variable external environments and gravitational perturbations, causing its actual trajectory to deviate from the predetermined nominal orbit, with significant errors in both position and velocity. Furthermore, the probe's state measurement sensors are subject to measurement errors, and the accuracy of measurement information decreases due to environmental variability. The superposition of various interferences and errors causes the probe's operational accuracy to gradually decrease with environmental changes and increasing operational time. By fusing and processing the probe's state information during real-time operation, the probe's operation can maintain a minimal error from its designed trajectory, ensuring precise and stable on-orbit operation according to mission requirements.
[0171] Based on the operational status of the detector agent and the state input requirements of the DQN algorithm, the uncertainties in the variance of the state process error and the variance of the measurement error noise during the terminal phase of the detector's cruise have a significant impact on its operation. Therefore, the state input of the deep neural network should include the state process error. and measurement error noise covariance Specifically, it can be defined as follows:
[0172] ;
[0173] In the formula, These are the input state parameters of the network. i =1,2,…,4, Depth Q The learning network has four state input parameters; for k The process position error noise variance of the time-lapse detector; For a moment k The process velocity error noise variance of the detector; For a moment k The noise variance of the measurement error of starlight angular distance; for k The measurement error noise variance of the solar radial velocity information at any given time has a dimension of 3×3, forming a 4×3×3 state input.
[0174] (2) Action space:
[0175] Depending on the state space, different actions are required. The action space provides multiple free action directions for the autonomous optimization of state parameters. For example... Figure 10 The four different state input parameters shown undergo optimal parameter trial and error within their respective ranges. The position error noise variance parameter has two actions: forward and backward. and The velocity error noise variance parameter has two actions: backward and forward. and The measurement error noise variance parameter of starlight angular distance has two actions: backward and forward. and The measurement error noise variance parameter of solar radial velocity information has two actions: backward and forward. and In the neural network fitting and learning process, the action space is calibrated by mapping numerical values to state parameters, with each action being assigned specific parameters. The action space set can be represented as... This corresponds to the specific numerical form {0,1,2,3,4,5,6,7}.
[0176] During action execution, selecting a specific action will change the corresponding action's state parameters, while other state parameters remain consistent with the previous state. Each state estimation utilizes DQN to interact with the environment, returning an action and generating the next state accordingly, continuously optimizing the state parameters. The reward mechanism will be designed in conjunction with a specific navigation filter state estimator. Rewards will continuously evaluate the navigation system's performance and update the state and actions based on the detector agent's operating environment, enabling real-time parameter tuning autonomously.
[0177] The DQNUKF intelligent navigation method based on planetary star angular distance in this embodiment is as follows:
[0178] The key to the DQNUKF algorithm is to use the neural network in the DQN algorithm to guide action changes based on feedback from state inputs and accumulated rewards, thereby determining a suitable set of noise covariance matrix parameters and continuously optimizing the parameters of the navigation state estimation to improve the performance of nonlinear filtering estimation. The DQN algorithm uses a deep neural network to fit a set of output scalars based on the process and observation noise variance matrix state inputs and actions in the detector agent, representing the Q-value corresponding to the action taken in the current state. The specific action selection is based on the action space described in the previous section, alternately selecting state parameters such as position, velocity error variance, starlight angular distance, and solar radial velocity information observation error variance one by one, forward and backward. The selection is also based on the Q-value corresponding to each action, using the following strategy:
[0179] ;
[0180] In the formula, ε The probability of generation; m To divide by the largest Q Value corresponds to action a The number of actions other than * indicates that the probability of taking each of the remaining actions is equal.
[0181] In the process of trial and error optimization of state parameters, a deep convolutional neural network is used as an approximation to fit DQNUKF. Q Value parameter, which takes environmental state variables as input, action value Q ( s , a This is used as the output of a CNN. Since high-dimensional data can still maintain the original relationships between state parameters after convolution, this feature is suitable for the real-time state estimation scenario of a detector agent operating in orbit. While ensuring that the dimensionality of the state and measurement parameter data remains unchanged, such as... Figure 11The diagram shows a multi-layer deep convolutional neural network structure designed for real-time state parameter estimation of the detector agent. This CNN structure has 4 inputs, 5 convolutional layers and 5 normalization layers as hidden layers, and 1 fully connected output layer. The input parameters have a shape of (1, 4, 3, 3), a batch size of 1, and four input layers (number of channels). The matrix dimension is 3*3, representing the position and velocity error variance matrices, and the observation error variance matrices for starlight angular distance and solar radial velocity information, respectively. The first convolutional layer, Conv1, has 32 channels, with its height and width remaining constant at 3*3, resulting in a normalized shape of (32, 4, 3, 3). The second convolutional layer, Conv2, has 64 channels, with its height increased by 1 while its width remains unchanged, resulting in a normalized shape of (64, 5, 3, 3). The third convolutional layer, Conv3, has 128 channels, with both its height and width increased by 1, resulting in a normalized shape of (128, 6, 3, 3). The fourth convolutional layer, Conv4, has 128 channels, with both its height and width increased by 1, resulting in a normalized shape of (128, 6, 3, 3). The fifth convolutional layer, Conv5, has 128 channels, with its height and width both increased by 1, resulting in a normalized shape of (128, 8, 3, 3). The fully connected output layer has inputs of (1, 1152) and outputs corresponding to 8 action spaces. Q value.
[0182] The key to DQNUKF lies in using the DQN algorithm to select appropriate process and observation noise variance matrices in real time, thereby improving the state estimation performance of the UKF estimator. Deep neural networks are used for approximate fitting to obtain... Q function Q ( s k , a k Based on the above action space, different actions correspond to different... Q The value is used to guide the selection of state parameters and the tuning of filter state parameters in the next iteration. The initial process and observation noise variance matrices are constructed as follows: There exists a positive definite matrix in each state, and in each state, according to different... Q ( s k , a kThe agent can take specific actions to proceed to the next step, obtaining a state parameter, and the quality of the action is evaluated through a reward. During the training process of the DQN algorithm, the selection of state parameters is based on the initial conditions determined by the aforementioned parameters. Other values in the distribution before and after the initial state parameters, differing from the initial parameters, are obtained by selecting state parameters forward and backward based on the action behavior. To cover uncertain noise variance, the forward and backward ranges should be as large as possible to ensure the selection of state parameters with very high state estimation accuracy. The key to the entire DQN algorithm lies in the interaction between the environment and the agent. Through the constantly changing "environment" of state and measurement error parameters in a highly dynamic environment, the filtered estimation bias and reward are fed back to the deep neural network in real time. The training experience stored in the experience replay pool is used to adjust the state parameters forward and backward, continuously searching for suitable neural network state inputs near the initial values to adapt to environmental changes. The Q-value learned through the deep convolutional network is used to calculate the loss function value, providing parameters for back-updating the network optimization.
[0183] Based on the above description, the specific implementation steps of the DQNUKF algorithm are as follows:
[0184]
[0185] In the above algorithm, and These represent the system state, position, and velocity estimation parameters for the UKF master filter and the provisional UKF filter, respectively. and These represent the corresponding error estimation covariance matrices. and The variance matrices of process position and velocity noise, respectively. and These are the observation noise error variance matrices for starlight angular distance and solar radial velocity information, respectively. , , and Based on the current environmental status s k To determine. T The time period for estimating the navigation state is a positive integer. and These represent the performance metrics of the UKF main filter and the provisional UKF filter for state estimation, respectively, used for reward evaluation in the DQN algorithm. The two filters run in parallel, utilizing the process position and velocity noise variance matrices, respectively. and The state estimate is obtained by taking the observation noise variance error of starlight angular distance and solar radial velocity information. The state estimates from the two filters are then compared to evaluate the current state. s k Related provisional filter state estimates A valuable state , , and A state with a large immediate reward will be used to update the parameters of the current and target networks in the neural network; conversely, a state with a smaller value will receive a relatively small reward. This is achieved through multiple training and updates using the neural network. Q The function ensures that the experience pool of DQN has sufficient experience, enabling the detector agent to continuously utilize depth as the current environment changes. Q The network learns to obtain the optimal state parameters, selects appropriate process and observation noise error variance matrices as inputs to the provisional filter, and runs in parallel with the main state estimation filter to continuously fuse information and finally obtain the optimal state estimate.
[0186] The following algorithm illustrates the implementation process of the UKF algorithm, a subroutine of the above algorithm.
[0187]
[0188] As can be seen from the algorithm above, instant reward R ( s k , a k The QLUKF algorithm is derived from the basic performance metrics of UKF. This design generates a significant immediate reward for the high-performance provisional filter, indicating that the current process and observation noise variance matrices are suitable for practical on-orbit probe state estimation and navigation systems. The QLUKF algorithm uses multi-step performance metric calculation. R ( s k , a k This allows for more information to be included in the immediate reward. To avoid the influence of historical data on the evaluation of the current state, the evaluation will be performed on the current state in each training round. Q ( s k , a k Update and reset the process and observation noise variance matrix of the provisional filter.
[0189] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the protection scope of this invention. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, without changing the core design of the algorithm and process, are also within the protection scope of this invention.
[0190] Another embodiment of the present invention relates to an intelligent integrated navigation system for deep space exploration. The implementation details of this intelligent integrated navigation system for deep space exploration are described below. The following implementation details are provided for ease of understanding and are not essential for implementing this solution. The intelligent integrated navigation system for deep space exploration in this embodiment includes:
[0191] The first model building module is used to build an orbital dynamics model of the probe during the terminal phase of its cruise when it explores planetary objects.
[0192] The second model building module is used to establish a state model and a measurement model of the probe autonomously navigating a planetary object during the terminal phase of its cruise, based on the orbital dynamics model. The measurement model is established based on the starlight angular distance information between the probe and the planetary object and satellite objects near the planetary object, as well as the solar radial velocity information.
[0193] The parameter optimization module is used to obtain the state error covariance of the state model and the measurement error covariance of the measurement model as the state input of the deep Q network, and to take multiple free action directions of the state error covariance and the measurement error covariance as the action output of the deep Q network, so as to obtain the optimal state error covariance and the optimal measurement error covariance through the deep Q network; wherein, the action direction is used to indicate the adjustment direction of the state error covariance or the measurement error covariance.
[0194] The detection and navigation module is used to estimate the state of the probe based on the optimal state error covariance and optimal measurement error covariance in real time at the end of the cruise phase, so as to complete the probe's autonomous navigation of planetary objects at the end of the cruise phase.
[0195] It is not difficult to see that this embodiment is a system embodiment corresponding to the above method embodiments, and this embodiment can be implemented in conjunction with the above method embodiments. The relevant technical details and technical effects mentioned in the above embodiments are still valid in this embodiment, and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiments.
[0196] It is worth mentioning that all modules involved in this embodiment are logical modules. In practical applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. Furthermore, to highlight the innovative aspects of this invention, this embodiment does not introduce units that are not closely related to solving the technical problem proposed by this invention; however, this does not mean that other units are absent from this embodiment.
[0197] Another embodiment of the present invention relates to a computer device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the deep space exploration intelligent integrated navigation method of the above embodiments.
[0198] The memory and processor are connected via a bus, which can include any number of interconnecting buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits, such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and will not be described further herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by the processor is transmitted over the wireless medium via an antenna, which further receives data and transmits it to the processor.
[0199] The processor manages the bus and general processing, and also provides various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory is used to store data used by the processor during operation.
[0200] Another embodiment of the present invention relates to a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the method embodiments described above.
[0201] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0202] Those skilled in the art will understand that the above embodiments are specific embodiments for implementing the present invention, and in practical applications, various changes can be made to them in form and detail without departing from the spirit and scope of the present invention.
Claims
1. A deep space exploration intelligent integrated navigation method, characterized in that, The method comprises: establishing an orbit dynamics model of the probe when detecting the planet at the end of the cruise phase; establishing a state model and a measurement model of the probe when autonomously navigating the planet at the end of the cruise phase according to the orbit dynamics model, wherein the measurement model is established based on starlight angular distance information between the probe and the planet and a satellite near the planet and solar radial velocity information; obtaining state error covariance of the state model and measurement error covariance of the measurement model as state inputs of a deep Q network, and obtaining multiple free action directions of the state error covariance and the measurement error covariance respectively as action outputs of the deep Q network, so as to obtain optimal state error covariance and optimal measurement error covariance through the deep Q network; wherein the action directions are used to indicate adjustment directions of the state error covariance or the measurement error covariance; performing state estimation on the probe according to real-time optimal state error covariance and optimal measurement error covariance of the probe at the end of the cruise phase, so as to complete autonomous navigation of the probe on the planet at the end of the cruise phase; wherein the state error covariance of the state model comprises position error covariance and velocity error covariance, and the measurement error covariance of the measurement model comprises starlight angular distance error covariance and solar radial velocity error covariance; the action outputs of the deep Q network comprise two action directions of the position error covariance, two action directions of the velocity error covariance, two action directions of the starlight angular distance error covariance, and two action directions of the solar radial velocity error covariance; the deep Q network adopts double unscented Kalman filters, wherein a main unscented Kalman filter performs state estimation on the probe according to the position error covariance and the velocity error covariance, a vice unscented Kalman filter performs state estimation on the probe according to the starlight angular distance error covariance and the solar radial velocity error covariance, and valuable position error covariance, velocity error covariance, starlight angular distance error covariance or solar radial velocity error covariance are instantaneously rewarded through state estimation results of the main unscented Kalman filter and the vice unscented Kalman filter, so as to update network parameters of the deep Q network, so as to obtain optimal state error covariance and optimal measurement error covariance.
2. The intelligent integrated navigation method for deep space exploration according to claim 1, characterized in that, The orbit dynamics model of the probe is as follows: wherein r pj is the position vector of the probe relative to the planet; r js is the direction vector of the planet relative to the sun; r ps is the position vector of the probe relative to the sun; 3. The intelligent integrated navigation method for deep space exploration according to claim 2, characterized in that, The state model of the probe is established by the following steps: convert the orbit dynamics model into: obtain the state model according to the converted orbit dynamics model as follows: wherein, 4. The intelligent integrated navigation method for deep space exploration according to claim 3, characterized in that, the measurement model of the probe comprises a first measurement model established based on starlight angular distance information between the probe and the planet and a satellite near the planet and a second measurement model established based on solar radial velocity information; the first measurement model is as follows: wherein, the second measurement model is as follows: wherein h RV is the cruise phase navigation measurement function using the radial velocity information of the sun; V RV is the observation noise due to the radial velocity information measurement of the sun.
5. An intelligent integrated navigation system for deep space exploration, characterized in that, The system comprises: a first model establishing module, configured to establish an orbit dynamics model of the probe when detecting the planet at the end of the cruise phase; The second model establishing module is configured to establish a state model and a measurement model of the probe during autonomous navigation of the planet object in the end of the cruise according to an orbit dynamics model, wherein the measurement model is established based on starlight angular distance information between the probe and the planet object and a satellite object near the planet object, and solar radial velocity information; The parameter optimization module is configured to obtain a state error covariance of the state model and a measurement error covariance of the measurement model as state inputs of the deep Q network, and obtain a plurality of free action directions of the state error covariance and the measurement error covariance respectively as action outputs of the deep Q network, so as to obtain an optimal state error covariance and an optimal measurement error covariance through the deep Q network, wherein the action directions are used to indicate adjustment directions of the state error covariance or the measurement error covariance. The probe navigation module is configured to perform state estimation on the probe according to the optimal state error covariance and the optimal measurement error covariance of the probe in real time in the end of the cruise, so as to complete autonomous navigation of the probe on the planet object in the end of the cruise. The state error covariance of the state model includes a position error covariance and a velocity error covariance, and the measurement error covariance of the measurement model includes a starlight angular distance error covariance and a solar radial velocity error covariance. The action outputs of the deep Q network include two action directions of the position error covariance, two action directions of the velocity error covariance, two action directions of the starlight angular distance error covariance, and two action directions of the solar radial velocity error covariance. The deep Q network adopts double unscented Kalman filters, wherein a main unscented Kalman filter performs state estimation on the probe according to the position error covariance and the velocity error covariance, and a vice unscented Kalman filter performs state estimation on the probe according to the starlight angular distance error covariance and the solar radial velocity error covariance. The parameter optimization module is specifically configured to perform instant rewards on valuable position error covariances, velocity error covariances, starlight angular distance error covariances, or solar radial velocity error covariances through state estimation results of the main unscented Kalman filter and the vice unscented Kalman filter, so as to update network parameters of the deep Q network, so as to obtain the optimal state error covariance and the optimal measurement error covariance.
6. A computer device, comprising: comprise: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the deep space probe intelligent integrated navigation method according to any one of claims 1 to 4.
7. A computer readable storage medium storing a computer program, characterized in that, The computer program is executed by the processor to implement the deep space probe intelligent integrated navigation method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Deep space explorer acquisition phase celestial navigation method based on target object ephemeris correction
CN105203101A
Unmanned ship positioning method based on Kalman filtering
CN117268400A