A frequency domain feature perception graph reinforcement learning single-beacon satellite positioning enhancement method
By employing a graph reinforcement learning method based on frequency domain feature perception, and utilizing discrete-time Fourier transform and graph feature processor, the problem of insufficient accuracy of single BeiDou satellite positioning in complex environments was solved, achieving high-precision positioning correction.
Patent Information
- Application Number
- CN202411982517.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-12-31
AI Technical Summary
Existing single-BeiDou satellite positioning methods lack sufficient positioning accuracy in complex urban environments and cannot effectively capture the potential changing patterns of satellite observation characteristics, resulting in low positioning accuracy.
A graph reinforcement learning method with frequency domain feature perception is adopted. The potential regularity information of satellite state sequence is extracted through discrete-time Fourier transform. The positioning is corrected by combining graph feature processor and cyclic actor-judge network. The reinforcement learning environment is constructed to improve positioning accuracy.
It improves the accuracy of satellite positioning, enabling it to meet the accuracy requirements of lane-level navigation in complex environments, and enhances the accuracy and reliability of positioning.
Smart Images

Figure CN119902245B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of satellite positioning, in particular to a graph reinforcement learning single Beidou satellite positioning enhancement method based on frequency domain feature perception. BACKGROUND
[0002] Nowadays, the global navigation satellite system (GNSS) is the most widely used method in various positioning technologies, and various industries are applying global satellite navigation positioning systems in complex multipath environments, so the demand for high-precision positioning of satellite navigation systems has increased dramatically. In the "3+16" key areas, the application of single Beidou is particularly prominent, and these application fields include housing planning, transportation logistics, etc., which are important directions for national development. Single Beidou allows devices to completely independently receive signals from the Beidou satellite navigation system, supports pure Beidou satellite navigation system signal resolution, and performs positioning and timing. The unique feature of this technology is that it does not rely on other satellite navigation systems, but is completely based on the Beidou satellite navigation system, thereby ensuring the accuracy and reliability of positioning. However, in complex urban environments, including urban canyons, under urban overpasses, etc., the multipath effect caused by the refraction and reflection of GNSS signals by obstacles, signal weakening caused by buildings and trees, a decrease in the number of visible satellites, and the deterioration of geometric distribution, etc. cause satellite positioning to deviate by tens of meters, which cannot meet the accuracy requirements of lane-level navigation. In recent years, with the rapid development of artificial intelligence technology, applying artificial intelligence methods to positioning correction in urban environments has become a mainstream solution, but the trained models have poor generalization performance and low positioning accuracy, and cannot be widely applied in practice. One of the key technical problems is that satellite observation features have time correlation and certain potential change rules within a period of time, such as periodic errors from satellite orbits, clock drift, atmospheric interference, etc., which produce certain regular fluctuations in satellite observations. However, the potential rules of this data cannot be clearly reflected in the time domain, and artificial intelligence models that only input satellite features in the time domain often lack the ability to capture the long-term trends and short-term disturbances of satellite signal features. SUMMARY
[0003] In view of this, in order to solve the technical problem that the existing positioning method can only rely on the satellite feature input in time sequence, cannot capture the potential change of data characteristics, and thus leads to low positioning accuracy, the present application provides a graph reinforcement learning single-beidou satellite positioning enhancement method based on frequency domain feature perception. Discrete time Fourier transform is used to extract potential regular information of state sequence signals and assist in learning the best strategy to improve positioning accuracy. First, a partially observable Markov decision process (POMDP) is introduced to model the environment and build a reinforcement learning environment for single-beidou satellite measurement. Then, a positioning correction model based on frequency domain state sequence prediction is constructed, which consists of three modules: 1) a frequency domain state sequence prediction model is constructed to extract and utilize the long-term trend and underlying structure of the state sequence. 2) a graph feature processor is constructed to extract the topological structure features of Beidou satellites and fully describe the state of the vehicle agent. 3) a recurrent actor-critic network is constructed to correct the positioning based on historical information and frequency domain information.
[0004] The graph reinforcement learning single-beidou satellite positioning enhancement method based on frequency domain feature perception specifically includes the following steps:
[0005] Building a reinforcement learning environment based on a single-beidou multi-frequency satellite graph;
[0006] Constructing a graph feature processor, a frequency domain state sequence prediction module, and an actor-critic network and integrating them to obtain a positioning correction model;
[0007] Using the large amount of state sequence data generated during the interaction between the reinforcement learning agent and the environment to train the positioning correction model;
[0008] Using the trained positioning correction model to correct the positioning.
[0009] Based on the above scheme, the application provides a frequency domain feature perception graph reinforcement learning single Beidou satellite positioning enhancement method, which converts the state sequence of state reinforcement learning to the frequency domain using discrete time Fourier transform (DTFT), so as to extract the potential long-term trend and short-term disturbance information to participate in the learning of the optimal strategy. First, in order to extract the spatial topological information between satellites, a satellite graph observation based on single Beidou multi-frequency is constructed, taking the features of each satellite as nodes, and constructing edges between nodes according to the correlation of satellite measurement, so as to construct the spatial relationship between satellites. Secondly, in order to capture the potential regularity change of satellite features, the input state sequence is converted to the frequency domain through DTFT, and then an auxiliary task model is constructed, which performs representation learning by predicting the frequency domain state sequence, and trains the feature extractor of reinforcement learning, so as to capture the potential change rule of the state sequence. Finally, a reinforcement learning model with graph neural network structure and frequency domain state representation learning is constructed, which combines the satellite spatial information extracted by the graph neural network and the satellite frequency domain information encoded by the auxiliary task model, can effectively model the satellite measurement error with potential regularity, and realize satellite dynamic positioning correction. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1 is a step flow chart of a frequency domain feature perception graph reinforcement learning single Beidou satellite positioning enhancement method of the application;
[0011] Figure 2 is a data flow direction schematic diagram of a frequency domain feature perception graph reinforcement learning single Beidou satellite positioning enhancement method of the application;
[0012] Figure 3 is a structure block diagram of a frequency domain state sequence prediction module of a specific embodiment of the application;
[0013] Figure 4 is a structure block diagram of an online encoder and a target encoder of a specific embodiment of the application. DETAILED DESCRIPTION
[0014] In addition to the method of using artificial intelligence for positioning correction mentioned in the background art, there is also a positioning correction method based on hardware devices, but this method relies too much on strict prior assumptions of sensor or model parameters, and the cost is high, which is difficult to popularize on a large scale.
[0015] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0016] It should be noted that only parts related to the present application are shown in the drawings for the convenience of description. The embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0017] It should be understood that the "system", "device", "unit" and / or "module" used in the present application is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the words can be replaced by other expressions.
[0018] As shown in the present application and claims, unless the context clearly indicates otherwise, the words "one", "a", "an" and / or "the" do not refer to the singular, but can also include the plural. Generally speaking, the terms "comprise" and "include" only indicate the inclusion of the steps and elements explicitly identified, and these steps and elements do not constitute an exclusive list, and the method or device can also include other steps or elements. The element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, product or device comprising the element.
[0019] In addition, flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or subsequent operations are not necessarily performed in sequence. On the contrary, each step can be processed in reverse order or simultaneously. At the same time, other operations can be added to these processes, or one or more steps of operation can be removed from these processes.
[0020] Reference Figure 1 The flowchart of an optional example of the frequency domain feature-aware graph reinforcement learning-based Beidou satellite positioning enhancement method proposed in the present application can be applied to a computer device. The positioning enhancement method proposed in the present embodiment can include but is not limited to the following steps:
[0021] Step S1, building a reinforcement learning environment based on a single Beidou multi-frequency satellite graph;
[0022] Step S2, constructing a positioning correction model, the positioning correction model including a graph feature processor, a frequency domain state sequence prediction module and an actor-critic network;
[0023] Step S3, training the positioning correction model based on the reinforcement learning environment to obtain a trained positioning correction model;
[0024] Step S4, acquiring a current Beidou satellite measurement signal;
[0025] Step S5, generating an initial position according to the current Beidou satellite measurement signal;
[0026] Step S6, input the current Beidou satellite measurement signal to the trained positioning correction model, and output correction data;
[0027] Step S7, correct the initial position according to the correction data to obtain a corrected position.
[0028] Satellite observation features have certain potential change rules in a period of time, such as periodic errors from satellite orbits, clock drift, atmospheric interference, etc. The errors of satellite observations generated thereby present certain regular fluctuations, but the potential rules are often difficult to reflect in the time domain.
[0029] In the frequency domain state sequence prediction module, the input observation features of reinforcement learning are converted to the frequency domain through discrete Fourier transform, so as to capture the potential regular disturbance of satellite measurement data. Disturbances of various error sources will appear as fluctuations or peaks at specific frequencies in the frequency domain. In the frequency domain, low-frequency components usually correspond to long-term change trends, and components close to zero correspond to the average value or slow change trend of the signal; high-frequency parts can reveal the interference from short-term changes or noise in the signal, and the high-frequency part helps to distinguish the rapid change interference in the signal. Therefore, it can help the reinforcement learning model to better capture the potential regular change of the signal, thereby improving the performance of the model.
[0030] The overall data flow of the method refers to Figure 2 .
[0031] In some feasible embodiments, the step S1 specifically comprises:
[0032] Action space a t is set as:
[0033] To avoid the processing of a large number of discrete actions leading to learning difficulties, the action space is set as a continuous action space in this paper. The action is defined as a correction operation on the initial position First, sampling is performed from a Gaussian distribution with mean μ and variance σ 2 , and the positioning correction action is defined as “:=” represents definition. Among them, is the mean of the Gaussian distribution in the x direction at time t, is the variance of the x direction, and similarly, are the mean and variance of the y direction in the Gaussian distribution, respectively, are the mean and variance of the z direction, respectively. Next, the position correction values Δx t , Δy t , and Δz t are sampled from the Gaussian distributions in the three directions, respectively, and the initial position is corrected to output the corrected position
[0034]
[0035] where u denotes a scale factor of the positioning correction action.
[0036] The reward function r t is set as:
[0037] To guide the reinforcement learning agent to learn the optimal positioning correction strategy, the initial position and the corrected position are calculated as the reward function of the environment:
[0038]
[0039] The coordinates of the reference position can be obtained by a high-precision but expensive integrated navigation system.
[0040] In some possible embodiments, the step S2 specifically comprises:
[0041] To make full use of the large amount of state sequence data generated in the process of interaction between the reinforcement learning agent and the environment, a model for predicting the long-term state sequence in the frequency domain is constructed to assist the representation learning, so as to extract the long-term trend and short-term disturbance information of the state sequence to assist the search for the optimal strategy. Finally, a recurrent actor-critic network is used to learn the optimal positioning correction strategy.
[0042] 1) Constructing a frequency-domain state sequence prediction model:
[0043] The satellite observation features have certain potential change rules in a period of time, such as periodic errors from satellite orbits, clock drifts, atmospheric interference, etc. The errors in satellite observations generated thereby present certain regular fluctuations, but the potential rules are often difficult to reflect in the time domain. By using discrete Fourier transform to convert the state sequence to the frequency domain, the potential regular disturbance in the satellite measurement data can be captured. In the frequency domain, the low-frequency component usually corresponds to the long-term change trend, and the component with frequency close to zero corresponds to the average value or slow change trend of the signal; the high-frequency part can reveal the interference from short-term changes or noise in the signal, and the disturbance of various error sources will appear as fluctuations or peaks at specific frequencies in the frequency domain, and the high-frequency part helps to distinguish the rapid change interference in the signal. Therefore, the frequency-domain state sequence information can help the reinforcement learning model to better capture the potential regular change of the signal, thereby improving the performance of the model. The reinforcement learning model based on frequency-domain state sequence prediction proposed in the present patent is shown in FIG. Figure 3 .
[0044] First, an auxiliary self-supervised task F Re (s t , at ), F Im (s t ,a t ) are the real and imaginary parts of the Fourier transform of the predicted state sequence, s t is the state of the reinforcement learning, which can be replaced by the observation o t . This state sequence prediction model consists of an online network and a target network, both of which use the same structure of encoder, predictor and projector. The parameters are updated by computing the loss function and backpropagating in the direction of the dashed line. In this structure, the online encoder and the target encoder have the same structure, which is shown in Fig. Figure 4
[0045] The output of the encoder φ is The input is the state s t at time t. (s t ,a t ,s t+1 ,a t+1 ) is a four-tuple sampled in the experience replay area. In this paper, the loss of reinforcement learning is not used to update the encoder, but the loss gradient of the state sequence Fourier transform prediction model is used to update the real-time encoder, but not the target encoder.
[0046] The method of the present application performs an auxiliary self-supervised task by training a predictor F(·) to predict the discrete-time Fourier transform (DTFT) of an infinite-step state sequence to capture the structural information in the state sequence for representation learning. Because of the conjugate symmetry property of the Fourier transform of the real state signal, only half a period is predicted in this paper, and then the transform of one period is obtained by symmetry to save storage space.
[0047] Modeling steps:
[0048] Given s t ,a t , the expectation of the future state sequence set on the infinite horizon of the nth step state is defined as
[0049]
[0050] γ∈[0,1) is a discount factor, E represents the expectation operation, and the next step is to perform the discrete Fourier transform on
[0051]
[0052] where ω is the frequency variable. Because the state sequence is a discrete signal, the corresponding discrete-time Fourier transform has a 2π periodicity in frequency, so we take L equidistant samples in the interval [0, π] to predict a matrix of size L*D, where D is the dimension of the state space. In continuous time, the DTFTs of different time instants are related to each other in a recursive form, and have the following relationship:
[0053]
[0054] where: and Γ are defined as
[0055]
[0056] Because the prediction result is a complex value, we divide the last layer of the predictor F into two parts, the real part F Re and the imaginary part F Im . Define the loss function of the prediction model as:
[0057]
[0058] where π represents the actor network, and represent the observation of the corresponding time. In training, we train the encoder φ, the predictor F, and the projector ψ to minimize the auxiliary loss L pred (φ, F, ψ), and then use the trained encoder φ to alternately update the actor-critic model of reinforcement learning.
[0059] 2) Graph feature processor:
[0060] This embodiment uses the pseudo-range residual RES and the LOS vector calculated by the GNSS measurement M t as the positioning correction features, and then defines the observation of the reinforcement learning environment as the node features of the graph neural network.
[0061]
[0062] and are the pseudo-range residual and the LOS vector of the jth satellite.
[0063] Next, the edges between the nodes in the Beidou satellite graph are constructed. The similarity of two satellites in signal frequency and pseudo-range residual is judged to determine whether to establish an edge between the nodes. In summary, the observation of the reinforcement learning environment in this paper is defined as a matrix containing the Beidou satellite node features and the adjacency matrix :
[0064]
[0065] For graph structure processing:
[0066] Let V be the set of n nodes at time t. t Define F t Let A be the feature matrix of the nodes, and assume that each node has m-dimensional features. The adjacency matrix of the graph is A. t ∈D n×n The dimension is n×n. Next, we define the j-th satellite as a node at time t. Represented as:
[0067]
[0068] Measurements of a reinforcement learning environment include GNSS satellite characteristics and adjacency matrices. The graph G is defined at time t. t This represents the information between satellites at the current moment. It is defined as:
[0069]
[0070] It is an adjacency matrix. These are node features. The core formula of the graph convolutional network algorithm is:
[0071]
[0072] in and These are the hidden features at time t, (l+1) and the l-th layer, respectively. The network input is the observations of the vehicle agent, i.e. The feature propagation rule for the l-th layer of GNN is:
[0073]
[0074] in I is the identity matrix. yes The degree matrix of the diagonal nodes, These are trainable weights, with n (0) =4. The activation function σ in this method uses the ELU function. β is a parameter in the ELU that controls the range of saturation values in the negative part. We use an L-layer GNN to compute G. t The final node embedding, GNN parameter set θ g ={W (1) ,…,W (L) To avoid satellite information loss, this embodiment uses pooling operations to aggregate features from multiple satellites and connects the node outputs of different satellites to form a flattened node embedding. tAs the topological belief state:
[0075]
[0076] For In other words, the jth satellite in the feature row of the Lth layer in the graph neural network.
[0077] 3) Recurrent actor-critic network:
[0078] The strategy output model recursive actor-critic network is constructed, and the model is trained through the proximal policy optimization algorithm (PPO). Define The value function of the critic, the corresponding parameter is θ c The estimated value of the output value of the critic network can be defined as:
[0079]
[0080] The loss function of the recurrent actor network obtained from the objective function of the PPO algorithm:
[0081]
[0082] The actor network updated parameters and pre-update parameters are respectively, E(.) represents the calculation of the expectation, ∈ is a hyperparameter related to the clipping range, is the probability ratio, the change range is [1-∈,1+∈], by doing so to limit the distance between the new policy and the old policy. Among them is the generalized advantage estimation (GAE), which combines the discount-return method and can balance the variables and bias in reinforcement learning. γ∈[0,1] is a discount factor, used to adjust the policy estimate.
[0083] In order to help the critic network to estimate the value more accurately, we define the TD error as:
[0084]
[0085] The goal of the reinforcement learning agent is to maximize the cumulative reward, so the critic network should accurately estimate the value of the belief state, so that the estimate is closer to the true cumulative reward, so the mean squared return error (MSRE) is used to define the loss function of the critic:
[0086]
[0087] In order to ensure that the agent interacts with the environment sufficiently, we define the entropy loss function:
[0088]
[0089] Then define the total target loss function:
[0090]
[0091] Where β1, β2 are coefficients.
[0092] In some possible embodiments, the step S4 specifically comprises:
[0093] This embodiment uses multi-frequency Beidou satellite measurement information as observation O t That is, using multiple frequency points of satellite signals of four systems of GPS, BDS, GLONASS and Galileo. We set the GNSS measurement set related to the jth Beidou satellite at time t as:
[0094]
[0095] Where is expressed as N is the total number of currently visible satellites, is a set of satellite positions, is a set of pseudoranges received by the receiver, and j is the number of the jth satellite.
[0096] In some possible embodiments, the step S5 specifically comprises:
[0097] Through a model-based method Such as Weight Least Squares (WLS), Kalman Filtering (KF), taking M t as input to obtain the initial position of the rough estimation.
[0098]
[0099] In some possible embodiments, the training process of the positioning correction model is specifically as follows:
[0100] Step 1: Use the receiver to collect positioning data in a real urban scene, for example, at time t, a set of pseudorange measurements is obtained by receiving measurement signals from multiple satellites The GNSS measurement set related to the Beidou satellite is set as: Next, through a model-based method, such as Weight Least Squares (WLS), Kalman Filtering (KF), taking M tAs input, we obtain the estimated initial position.
[0101] Step 2: Initialize an experience replay pool; the agent continuously interacts with the environment to collect experience. The observation at time t... t Next, action a is obtained through the policy network. t Discount rate γ, number of steps to accumulate rewards T, learning rate lr.
[0102] Step 3: Next, define the initial parameters, including the step size N, batch size t, and the parameter set θ of the graph neural network. g The parameter set θ of the actor network a The parameter set θ of the judge network c The parameters of the online encoder and target encoder of the state sequence prediction model are defined as follows: the parameters of the online encoder include the parameters of the corresponding predictor F and projector ψ, respectively, and are defined as θ. aux The parameters of the target encoder include the corresponding predictor. Projector The parameters are defined as follows: Define the smoothing coefficient τ and update interval K of the target network.
[0103] Step 4: Train the reinforcement learning model using the experience in the experience replay pool via the PPO algorithm, and update the parameters (θ) of the graph network, actor network, and judge network using stochastic gradient descent. g ,θ a ,θ c The loss function is: Simultaneously update the online network of the state sequence prediction model: This indicates the gradient calculation operation; and the updating of the parameters of the target network of the state sequence prediction model after every k steps:
[0104] Step 5: After reaching the maximum number of training iterations, fix the parameters of the trained reinforcement learning model, deploy the multi-agent policy network into the satellite positioning chip, and test its application in a real environment.
[0105] The overall training process is not significantly different from that of reinforcement learning.
[0106] A frequency domain feature-aware graph reinforcement learning-based single-BeiDou satellite positioning augmentation device:
[0107] At least one processor;
[0108] At least one memory for storing at least one program;
[0109] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned frequency domain feature-aware graph reinforcement learning single Beidou satellite positioning enhancement method.
[0110] The contents in the above method embodiments are applicable to the device embodiments, the device embodiments specifically implement the same functions as the above method embodiments, and achieve the same beneficial effects as the above method embodiments.
[0111] A storage medium, wherein the storage medium stores processor-executable instructions, and the processor-executable instructions, when executed by a processor, are used to implement the above-mentioned frequency domain feature-aware graph reinforcement learning single Beidou satellite positioning enhancement method.
[0112] The contents in the above method embodiments are applicable to the storage medium embodiments, the storage medium embodiments specifically implement the same functions as the above method embodiments, and achieve the same beneficial effects as the above method embodiments.
[0113] The above is a specific description of the preferred embodiments of the application, but the application is not limited to the above-mentioned embodiments, and those skilled in the art can make various equivalent modifications or replacements without departing from the spirit of the application, and these equivalent modifications or replacements are all included in the scope defined by the claims of the present application.
Claims
1. A graph reinforcement learning method for BeiDou satellite positioning enhancement with frequency domain feature perception, characterized in that, The method comprises the following steps: An environment for reinforcement learning is built based on a single-Beidou multi-frequency satellite graph; A positioning correction model is constructed, which comprises a graph feature processor, a frequency domain state sequence prediction module, and an actor-critic network; The positioning correction model is trained based on the environment for reinforcement learning to obtain a trained positioning correction model; A current Beidou satellite measurement signal is acquired; An initial position is generated according to the current Beidou satellite measurement signal; The current Beidou satellite measurement signal is input into the trained positioning correction model; The current Beidou satellite measurement signal is converted into a satellite graph by the graph feature processor; The state sequence is converted to the frequency domain by the frequency domain state sequence prediction module, and regular disturbances in the Beidou satellite measurement signal are captured to obtain predicted frequency domain information; The corrected data is generated by the actor-critic network in combination with the predicted frequency domain information and historical information; The initial position is corrected according to the corrected data to obtain a corrected position. 2.The method of claim 1, wherein, The environment for reinforcement learning specifically comprises: The action space is set as a continuous action space, and the correction action is defined as: wherein, is the mean of the Gaussian distribution in the x direction at time t, is the variance of the Gaussian distribution in the x direction at time t, and similarly, are the mean and variance of the Gaussian distribution in the y direction at time t, respectively, are the mean and variance of the Gaussian distribution in the z direction at time t, respectively. The reward function is set as follows: wherein, represents an initial position, represents a corrected position, represents a reference position. 3.The method of claim 2, wherein, The satellite graph is constructed as follows: The observation is defined as a node of the Beidou satellite graph, and is represented as follows: wherein, and is the jth BeiDou satellite pseudo-range residual and LOS vector; According to the similarity of two satellite measurements in signal frequency and pseudorange residual, it is determined whether to establish an edge between nodes. 4.The method of claim 2, wherein, The frequency domain state sequence prediction module comprises an online network and a target network, wherein: The online network comprises an online encoder, a real-time predictor, and a real-time projector; The target network comprises a target encoder, a target predictor, and a target projector; The real-time predictor is updated by a loss gradient; The real-time predictor and the target predictor are used to predict the discrete-time Fourier transform of an infinite-step state sequence to perform an auxiliary self-supervised task.
5. The method of claim 4, wherein the graph reinforcement learning is performed in a frequency domain. dividing the last layer of the real-time predictor into a real part F Re and an imaginary part F Im defining a loss function L pred for the real-time prediction period as follows: where φ denotes the real-time encoder, F the real-time predictor, and denotes the representation of the observation at the corresponding time instant, π denotes the actor network, denotes the Fourier transform of the observation vector, d denotes the cosine similarity. 6.The method of claim 5, wherein, The total target loss function of the actor-critic network is represented as follows: where β1, β2 represent the corresponding coefficients, E(.) represents the calculation of expectation, represents the loss function of the actor network, represents the loss function of the critic network, represents the policy entropy of the actor network, are the updated parameters and the parameters before updating of the actor network respectively; ∈ is a hyperparameter related to the clipping range; is the probability ratio, is the generalized advantage estimation, and γ represents the discount factor.
7. A frequency domain feature-aware graph reinforcement learning single-beacon GNSS augmentation device, characterized in that, comprises: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method for single-Beidou satellite positioning enhancement based on graph reinforcement learning with frequency domain feature perception according to any one of claims 1-6.
Citation Information
Patent Citations
Multi-agent autonomous collaborative obstacle avoidance navigation method based on deep reinforcement learning
CN118089734A
Beidou satellite high-precision positioning method and system in urban complex environment
CN118519179A