Lunar surface positioning enhancement method based on lunar navigation constellation

Through deep neural networks and reinforcement learning methods, the noise interference law in complex environments in the lunar polar regions is learned, and the problem of low positioning accuracy in the existing technology is solved, and accurate lunar-based positioning is achieved in complex dynamic environments.

CN119991783APending Publication Date: 2025-05-13GUANGDONG UNIV OF TECH +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510073045.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing lunar positioning enhancement methods are difficult to accurately model complex dynamic random noise interference in complex dynamic environments in the lunar polar regions, resulting in low positioning accuracy.

Method used

Deep neural network combined with reinforcement learning methods are used to learn the law of the impact of dynamic changes in random noise in complex unknown environments in the lunar polar region on location, avoid strict prior assumptions of the noise model, and achieve accurate lunar basis positioning in complex dynamic environments in the lunar polar region.

Benefits of technology

Through deep reinforcement learning methods, agents can effectively learn optimal position correction strategies, improve the positioning accuracy of the lunar rover in the lunar pole region, and overcome the problem of limited accuracy of traditional methods in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991783A_ABST
    Figure CN119991783A_ABST
Patent Text Reader

Abstract

The invention discloses a lunar surface positioning enhancement method based on a lunar navigation constellation. The method comprises the following steps: constructing a deep learning environment; wherein the deep learning environment comprises an action space, an observation space, a potential state and a reward mechanism; the method comprises the following steps of: constructing a deep reinforcement learning model of an Actor-Critic Net (Actor-Critic Net) by taking a double-layer deep neural network as an infrastructure; performing dynamic interaction on the deep reinforcement learning model and the deep learning environment, minimizing a preset loss function by using a stochastic gradient descent method, and updating parameters of the deep reinforcement learning model; and deploying the trained positioning correction model to a server cloud, and performing positioning enhancement according to the observed quantity input in real time. According to the invention, accurate satellite lunar surface positioning under the complex dynamic environment of the lunar surface polar region is realized. The method can be widely applied to the field of moon positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of lunar positioning, and in particular to a lunar surface positioning enhancement method based on a lunar navigation constellation. Background Art

[0002] In the field of international lunar exploration, the current positioning research around the lunar navigation constellation mainly uses model-driven methods such as Kalman filtering and its variants. However, in the special environment of the moon, especially the lunar polar regions, the deep space environment has severe space radiation, drastic temperature changes, craters are everywhere, and the terrain is complex and undulating. As a result, the satellite signals received by the lunar rover positioning equipment are subject to complex dynamic environmental interference. The model-based method relies on strict prior assumptions and data statistics such as the propagation characteristics of lunar-based satellite signals and model parameters, and it is difficult to model and correct the dynamic interference of random noise on the lunar surface, resulting in the positioning performance of the rover in the complex and unknown environment on the lunar surface being severely limited.

[0003] Therefore, how to achieve high-precision lunar-based positioning in the complex and unknown dynamic environment of the lunar polar regions is an important challenge that needs to be urgently solved in the current lunar exploration project. Summary of the invention

[0004] In view of this, in order to solve the technical problem that the existing lunar surface positioning enhancement methods are interfered by complex dynamic random noise, making it difficult to achieve accurate modeling, which in turn leads to low positioning accuracy.

[0005] Before implementing this method, the problem is modeled:

[0006] That is: analyze the lunar rover positioning and rewrite the position output of the Kalman filter;

[0007] Assuming that the lunar rover is moving in the complex environment of the lunar South Pole, based on the satellite signal of the lunar navigation system, the extended Kalman filter algorithm is used to realize the initial solution of its own position. The position output result of the Kalman filter at the tth moment is:

[0008] u t|t =u t|t-1 +K t (ρ t -ρ 0,t )#(1)

[0009] In the formula, u t|t-1 =F t u t-1|t-1 Represents the prior estimated position, which is represented by the state transfer matrix F t And the position output result u at the previous moment t-1|t-1 Joint decision; K t represents the Kalman gain, ρ t represents the pseudorange measurement value, ρ 0,tRepresents the geometric distance between the lunar satellite and the prior estimated position. The position output result at this moment can be seen as consisting of the following parts:

[0010]

[0011]

[0012] In the formula, ω t represents the process noise caused by dynamic modeling and is assumed to follow a Gaussian distribution ω t ~N(0,Q t ), Q t represents the process noise covariance matrix; v t / v 0,t represents the measurement noise caused by the measurement, which is also assumed to follow a Gaussian distribution v t / v 0,t ~N(0,R t ), R t represents the measurement noise covariance matrix; P t-1|t-1 represents the error covariance matrix of the previous moment, H t Denotes the geometric matrix. Let ε t is the total noise term, that is:

[0013]

[0014] It can be seen that the total noise ε that causes the lunar rover positioning deviation is t It is mainly caused by two factors: First, the extreme physical environment of the deep space of the moon causes the satellite signals received by the rover to be subject to complex dynamic random noise. t , v t and v 0,t It no longer conforms to the Gaussian distribution, and it is difficult to accurately model this random noise interference, which greatly limits the accuracy of the satellite positioning method based on the model. Second, for the random noise model Q that has been modeled t and R t Generally, it is simplified to assume time-invariant models Q and R. The selection of model parameters depends on empirical knowledge and data statistics based on the orbit determination accuracy of lunar-based satellites, signal propagation characteristics, and receiver capture and tracking performance. It is a strict a priori assumption. However, in the complex and unknown dynamic environment of the moon, the time-invariant model cannot accurately describe the time-varying noise. Therefore, the a priori assumption is severely damaged, which also leads to a decrease in the accuracy of the model-based satellite positioning method.

[0015] To this end, the present invention proposes a lunar positioning enhancement method based on the lunar navigation constellation, which uses the highly nonlinear fitting advantage of deep neural networks to learn the influence of the dynamic change interference of random noise in the complex and unknown environment of the lunar polar region on positioning. At the same time, the characteristics of reinforcement learning and dynamic interactive exploration of the environment are used to help the intelligent agent effectively learn the optimal position correction strategy, thereby avoiding strict a priori assumptions about the noise model and achieving accurate lunar-based positioning in the complex dynamic environment of the lunar polar region. For the convenience of subsequent expression, the position output result of the Kalman filter at the tth moment represented by formula (2) is rewritten as:

[0016] POS ekf,t =pos true,t +ε t #(4)

[0017] The proposed deep reinforcement learning method intends to learn the action output Δpos obtained by the optimal correction strategy through interaction with the lunar environment. DRL,t , eliminating the total noise term ε t ,Right now:

[0018] POS DRL,t =pos ekf,t +Δpos DRL,t

[0019] =pos true,t +ε t +Δpos DRL,t

[0020] ≈pos ture,t #(5)

[0021] Thereby ultimately achieving the purpose of positioning enhancement.

[0022] The method comprises the following steps:

[0023] Build a deep learning environment;

[0024] Among them, the deep learning environment includes action space, observation space, potential state and reward mechanism;

[0025] Based on a two-layer deep neural network as the basic architecture, a deep reinforcement learning model of the Actor-Critic Net is constructed;

[0026] Dynamically interacting the deep reinforcement learning model with the deep learning environment, minimizing a preset loss function using a stochastic gradient descent method, and updating parameters of the deep reinforcement learning model;

[0027] The trained positioning correction model is deployed to the server cloud, and positioning enhancement is performed based on the real-time input observations.

[0028] Based on the above scheme, the present invention provides a lunar positioning enhancement method based on the lunar navigation constellation, constructs a positioning correction environment for the lunar polar region, designs a reward function with maximizing the positioning correction effect as an indicator, and adopts the proximal policy gradient optimization algorithm (PPO) to optimize the learning of the optimal strategy. By learning the law of the influence of the dynamic change interference of random noise in the complex and unknown environment of the lunar polar region on positioning, the intelligent agent is helped to effectively learn the optimal position correction strategy, avoid strict prior assumptions on the noise model, and realize accurate satellite lunar surface positioning in the complex dynamic environment of the lunar polar region. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a flow chart of the steps of a method for enhancing lunar surface positioning based on a lunar navigation constellation according to the present invention;

[0030] Figure 2 It is a data flow diagram of a lunar surface positioning enhancement method based on a lunar navigation constellation according to the present invention. DETAILED DESCRIPTION

[0031] Existing technical solutions mainly use model-driven methods such as Kalman filtering and its variants. However, these model-based positioning methods are all simulation experiments conducted under ideal external environments with high satellite orbit determination accuracy. In the actual complex and unknown environment of the moon, they will face the problems of noise models being difficult to accurately model and strict a priori assumptions being destroyed in complex and unknown environments, which greatly limits positioning accuracy.

[0032] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0033] It should be noted that, for the convenience of description, only the parts related to the relevant invention are shown in the drawings. In the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other.

[0034] It should be understood that the "system" and / or "module" used in this application is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the word can be replaced by other expressions.

[0035] As shown in this application and claims, unless the context clearly indicates an exception, the words "a", "an", "a kind" and / or "the" do not refer to the singular, but also include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements. The elements defined by the sentence "includes a..." do not exclude the existence of other identical elements in the process, method, commodity or device that includes the elements.

[0036] In the description of the embodiments of the present application, "plurality" means two or more than two. The following terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features.

[0037] In addition, flow charts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed accurately in order. On the contrary, the various steps may be processed in reverse order or simultaneously. At the same time, other operations may also be added to these processes, or a certain step or several steps of operations may be removed from these processes.

[0038] Reference Figure 1 , is a flow chart of an optional example of a lunar surface positioning enhancement method based on a lunar communication constellation proposed in the present invention. The method can be applied to a computer device. The positioning method proposed in this embodiment may include but is not limited to the following steps:

[0039] Step S1, building a deep learning environment;

[0040] The learning environment includes an action space and an observation space based on the rover trajectory context and the ELFO satellite feature enhanced context.

[0041] Step S2, building a deep reinforcement learning model;

[0042] Step S3: dynamically interact the proposed model with the environment to minimize the loss function and complete the training of the deep reinforcement learning model;

[0043] In some feasible embodiments, the step S1 specifically includes:

[0044] Action space: This embodiment defines the action of the deep reinforcement learning agent as the correction of the result of the lunar rover using the extended Kalman filter algorithm, denoted as: Δpos DRL,t =[Δx t ,Δy t,Δz t ] T , then the deep reinforcement learning lunar orbit satellite positioning enhancement result can be finally expressed as: pos DRL,t =pos ekf,t +Δpos DRL,t =[x DRL,t ,y DRL,t ,z DRL,t ] T Here, the agent’s actions are determined by the actor network Given, and to avoid the complexity of discrete action space caused by a large number of actions, this patent chooses to output continuous actions. Specifically, the participant network According to the given potential state Output the mean-variance pairs of the three continuous action components respectively Then, the corresponding Gaussian distribution is constructed based on the output mean-variance pair, and the continuous action value Δx is obtained by sampling from the Gaussian distribution. t / Δy t / Δ z t.

[0045] Observation space and potential state. The lunar rover receiver outputs positioning results at a certain frequency. These positioning results indirectly reflect the motion law of the rover. Therefore, in order to make full use of the rover trajectory context information to capture the motion state of the rover, the trajectory observation value of the agent is not limited to the positioning output result at the current moment, but a stacked vector, which contains the Kalman filter position output result at the current moment and the reinforcement learning position correction output result of the previous k moments, that is: pos,t :=[pos DRL,t-k ,pos DRL,t-2 ,…,pos DRL,t-1 ,pos ekf,t ] T At the same time, in order to fully capture the motion state of the lunar rover in different lunar geographical locations and environmental conditions, the ELFO satellite feature context is used as observation enhancement to form a multi-perspective deep reinforcement learning environment observation, namely: The superscript i represents the i-th satellite. represents the pseudo-range measurement between the i-th ELFO satellite and the lunar rover at time t, represents the geometric distance between the i-th ELFO satellite and the lunar rover position obtained by EKF at time t. Based on this, the observation space based on the rover trajectory context and the ELFO satellite feature enhancement context can be defined as: t = {O pos,t ,O elfo,t}.

[0046] Facing the observation space O tThe task is to identify the most critical features for policy learning and construct accurate latent state representations based on these features. To this end, the long short-term memory (LSTM) network module and the attention mechanism module are used to process the observations from each observation perspective respectively, so as to effectively extract effective features from different observation perspectives and fuse them into latent states. Specifically, suppose that the two LSTM modules are composed of parameter sets θ pos and θ elfo Composition, given O pos,t and O elfo,t , then through the forward propagation of the LSTM module, the output gate can obtain two environmental observation feature outputs containing historical information and At the same time, in order to make the feature fusion of each perspective input fully reflect the difference information and consistency information of multi-perspective input, the attention mechanism is further used to mine the attention weights representing the relationship between different perspectives to effectively perform multi-perspective feature fusion. The output of the attention module for each perspective feature is defined as w pos,t and w elfo,t , the hidden representations in the two partial observations are fused to form the latent state of the environment, namely:

[0047]

[0048] Reward mechanism: To ensure that the reward signal can effectively guide strategy learning, the corrected advantage error is used as the reward signal. The corrected advantage error at time t is defined as:

[0049] r t =||pos ekf,t -pos true,t || 2 -||pos rl,t -pos true,t || 2 #(7)

[0050] where pos true,t =[x true,t ,y true,t ,z true,t ] T Represents the true position of the lunar rover at time t. Compared with simply using the negative value of the Euclidean distance between the corrected position and the true position as a reward signal, the corrected advantage error is more inclined to encourage the agent's good correction behavior rather than just punishment, which can better indicate whether the positive behavior based on good correction is effective and speed up the model convergence.

[0051] In some feasible embodiments, the step S2 specifically includes:

[0052] This embodiment uses a two-layer deep neural network to construct participant networks and reviewer network

[0053] In some feasible embodiments, the step S3 specifically includes:

[0054] The PPO algorithm is used to train the participant-evaluator network, ensuring that the update of the trust region follows the advantage clipping strategy and guarantees its effectiveness in the continuous action space. The goal is to update the participant network parameters so that the action output value is constantly close to the optimal strategy, and the loss function of the participant network is defined by clipping the objective function of the proximal strategy optimization algorithm:

[0055]

[0056] In the formula, θ a are the participant network parameters, is the generalized advantage estimate; clip(·) is the clipping function, which modifies the objective function of the policy network by clipping the probability ratio, eliminating p i (θ a ) is outside the interval [1-∈, 1+∈]; p i (θ a ) is the probability ratio between the new and old strategy networks, and its expression is as follows:

[0057]

[0058] in For the old policy network, For the new strategy network.

[0059] For the reviewer network In order to achieve accurate value estimation and long-term precise positioning, the evaluator loss function It is defined as the Mean-Squared Return Error (MSRE) taking into account the cumulative expected reward:

[0060]

[0061] In the formula, φ c is the evaluator network parameter, is the cumulative discounted reward considering the T-step lunar rover trajectory, and γ is the discount factor.

[0062] Finally, the above loss functions (8) and (10) are minimized by the gradient descent method to update the parameters of the deep reinforcement learning lunar orbit satellite positioning enhancement model.

[0063] Specifically:

[0064] S3.1. Collect the positioning data of the lunar rover in the real lunar scene based on the lunar navigation system satellite to achieve single-point positioning, which is used to construct a data set for subsequent model training. Assuming that the lunar rover receives the navigation message from n lunar navigation system ELFO satellites at time t, a set of pseudo-range observation data can be obtained and satellite location data At the same time, based on the navigation message and EKF algorithm, the initial position pos of the lunar rover at that moment can be obtained. ekf,t Based on the high-resolution lunar topographic map, the real position pos at that moment can be obtained true,t . Collecting N time data in total, you can get the positioning data: and

[0065] S3.2. Construct a deep reinforcement learning environment and a deep reinforcement learning model. Use the pseudo-range observation data in the positioning data, the satellite position data, and the initial position data of the lunar rover to construct a reinforcement learning environment observation model. t = {O pos,t ,O elfo,t}, where the probe vehicle trajectory context O pos,t and ELFO satellite feature enhancement context O elfo,t For the specific expression, see the "Observation Space and Potential State" section. At the same time, the correction action Δpos of the agent is defined using the explanation in the "Action Space" section DRL,t and the corresponding environmental reward r t , using the participant-evaluator network to build a reinforcement learning model, where the evaluator network Used to estimate the value function of the current observation state, the participant network Output action for initial position pos ekf,t In order to achieve accurate position correction of the lunar rover. In addition, the observation inputs from two perspectives are fused using the LSTM module and the attention module to form a latent state as shown in formula (6), thereby identifying and extracting effective features that can produce accurate latent state representation.

[0066] S3.3. Training the deep reinforcement learning lunar surface positioning enhancement model. The proposed model dynamically interacts with the environment, and the stochastic gradient descent method is used to minimize the loss functions of equations (8) and (10) to update the network parameters of the participant network and the evaluator network.

[0067] Based on the above scheme, the data flow based on this model refers to Figure 2 .

[0068] In some feasible embodiments, it also includes:

[0069] After the model obtains the optimal positioning correction strategy, the proposed model will be deployed to the lunar base station server cloud, and tested and applied in the real complex and unknown environment of the moon with the help of the cloud. By inputting the observations calculated by the satellite measurements of the lunar communication and navigation system received in real time, the participant network outputs actions to the user terminal end-to-end for precise position correction of the lunar rover.

[0070] Based on the above scheme, the effects to be achieved by the present invention are as follows: learning the law of the dynamic change interference of random noise in the complex unknown environment of the lunar polar region on positioning through the highly nonlinear fitting advantage of deep neural networks, and helping the intelligent agent to effectively learn the optimal position correction strategy by means of the characteristics of reinforcement learning and dynamic interactive exploration of the environment, solving the problem that the traditional model-based method is limited by the difficulty of accurate modeling of noise models and the destruction of strict prior assumptions in complex unknown environments; in order to ensure that the reinforcement learning training scene is more in line with the real lunar environment, the present invention specifically constructs a lunar orbit satellite positioning reinforcement learning environment, in which the environmental observation system uses the estimated position sequence, the line of sight vector and the pseudo-range residual as the feature observation input of the ELFO satellite, so that the environmental observation can more accurately characterize the current state information of the lunar rover intelligent agent; in order to ensure the effective training of the reinforcement learning intelligent agent, the present invention adopts the participant-evaluator network to construct the reinforcement learning model, and uses the long short-term memory module and the attention module to extract and fuse the multi-view environmental observation input to form the potential state of the environment, thereby identifying and extracting effective features that can produce accurate potential state representation. At the same time, the proximal policy gradient optimization algorithm is used to optimize the learning optimal strategy, ensure that the trust domain is updated with the advantage clipping strategy, and is effective in the continuous action environment.

[0071] A lunar surface positioning enhancement system based on a lunar navigation constellation, comprising:

[0072] The environment construction unit uses the ELFO satellite feature context as observation enhancement to form multi-perspective deep reinforcement learning environment observations and build a deep learning environment;

[0073] The model building unit builds a participant network and an evaluator network based on a two-layer deep neural network to obtain a deep reinforcement learning model;

[0074] A training unit dynamically interacts the deep reinforcement learning model with the deep learning environment, trains the deep reinforcement learning model, and obtains a trained positioning correction model;

[0075] The application unit performs positioning enhancement based on the trained positioning correction model to obtain a positioning enhancement result.

[0076] The contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0077] A lunar surface positioning enhancement device based on the lunar navigation constellation:

[0078] at least one processor;

[0079] at least one memory for storing at least one program;

[0080] When the at least one program is executed by the at least one processor, the at least one processor implements the lunar surface positioning enhancement method based on the lunar guidance constellation as described above.

[0081] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0082] A storage medium stores processor-executable instructions, which are used to implement a lunar surface positioning enhancement method based on a lunar navigation constellation when executed by a processor.

[0083] The contents of the above method embodiments are all applicable to the present storage medium embodiments. The functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0084] The above is a specific description of the preferred implementation of the present invention, but the invention is not limited to the embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A lunar surface positioning enhancement method based on a lunar navigation constellation, characterized in that: The following steps are involved: The ELFO satellite feature context is used as observation enhancement to form multi-perspective deep reinforcement learning environment observation and build a deep learning environment; Based on a two-layer deep neural network, a participant network and an evaluator network are constructed to obtain a deep reinforcement learning model. Dynamically interacting the deep reinforcement learning model with the deep learning environment, training the deep reinforcement learning model, and obtaining a trained positioning correction model; Positioning enhancement is performed based on the trained positioning correction model to obtain a positioning enhancement result.

2. According to claim 1, a method for enhancing lunar surface positioning based on a lunar navigation constellation is characterized in that: Also includes: The trained positioning correction model is deployed to the lunar base station server cloud.

3. According to claim 2, a method for enhancing lunar surface positioning based on a lunar navigation constellation is characterized in that: The result of the positioning enhancement is expressed as: pos DRL,t =[x DRL,t ,y DRL,t ,z DRL,t ] T =pos ekf,t +Δpos DRL,t Among them, pos DRL,t represents the result of positioning enhancement at time t, x DRL,t ,y DRL,t ,z DRL,t Respectively represent pos DRL,t The three components of pos ekf,t represents the initial position of the lunar rover at time t, Δpos DRL,t Indicates the amount of positioning enhancement.

4. According to claim 2, a method for enhancing lunar surface positioning based on a lunar navigation constellation is characterized in that: The observation space in a deep learning setting is represented as follows: The t ={O pos,t ,O elfo,t } THE pos,t :=[pos DRL,t-k ,pos DRL,t-2 ,…,pos DRL,t-1 ,pos ekf,t ] T Among them, O t represents the observation space, O pos,t represents the observation space based on the context of the probe vehicle trajectory, represents the multi-view deep reinforcement learning environment observation, the superscript i represents the i-th satellite, represents the pseudo-range measurement between the i-th ELFO satellite and the lunar rover at time t, represents the geometric distance between the i-th ELFO satellite and the lunar rover position obtained by EKF at time t, k represents the time length of the rover trajectory context, They represent the three directional components of the i-th satellite at time t, x ekf,t ,y ekf,t ,z ekf,t They respectively represent the three components of the position output result of the Kalman filter at the tth moment.

5. The method for enhancing lunar surface positioning based on the lunar navigation constellation according to claim 4, characterized in that: In the observation space: Based on the LSTM network module and the attention mechanism module, the observation values ​​under each observation perspective are processed respectively, and the effective features are extracted and fused to obtain the potential state.

6. The method for enhancing lunar surface positioning based on the lunar navigation constellation according to claim 5, characterized in that: During training, the corrected advantage error is used as a reward, which is expressed as follows: Among them, pos ekf,t Represents the position of the Kalman filter at time t; pos true,t Indicates the actual position at time t; pos DRL,t Represents the positioning enhancement result at time t.

7. A lunar surface positioning enhancement system based on a lunar navigation constellation, characterized in that: include: The environment construction unit uses the ELFO satellite feature context as observation enhancement to form multi-perspective deep reinforcement learning environment observations and build a deep learning environment; The model building unit builds a participant network and an evaluator network based on a two-layer deep neural network to obtain a deep reinforcement learning model; A training unit dynamically interacts the deep reinforcement learning model with the deep learning environment, trains the deep reinforcement learning model, and obtains a trained positioning correction model; The application unit performs positioning enhancement based on the trained positioning correction model to obtain a positioning enhancement result.

8. A lunar surface positioning enhancement device based on a lunar navigation constellation, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the lunar surface positioning enhancement method based on the lunar guidance constellation as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Moon area navigation enhancement method based on translation point navigation constellation

    CN114578395A

  • Beidou positioning method based on adaptive noise covariance learning and denoising

    CN118981031A

  • Multi-agent federated reinforcement learning-based vehicle-road collaborative control system and method under complex intersection

    WO2024016386A1

  • System for vision-based self-decisive planetary hazard free landing of a space vehicle

    WO2024052928A1