Satellite Selection Method and System for Power Enhancement of Navigation Satellites Based on Deep Reinforcement Learning
By modeling the navigation satellite power-enhanced star selection mission as a Markov decision-making process, and using deep reinforcement learning models to generate the task participation of navigation satellites within the mission time window, the problem of high complexity of star selection calculation and difficulty in adapting to the dynamic interference environment in the existing technology is solved, and dynamic optimization and real-time improvement of navigation satellite signal power-enhanced star selection is achieved.
Patent Information
- Application Number
- CN202510432117.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing navigation satellite signal power enhancement star selection methods have problems such as high computational complexity, insufficient real-time performance, and difficulty in adapting to dynamically changing interfering environments.
The navigation satellite power-enhanced star selection method based on deep reinforcement learning is adopted. By modeling the navigation satellite power-enhanced star selection task as a Markov decision-making process, state space, action space and reward functions are defined, and the pre-trained power-enhanced star selection model is used to generate task participation situations for each navigation satellite at different moments in the task time window, realizing dynamic optimization of the star selection results.
The influence of subjective factors is reduced, the optimization effect of star selection results is improved, and the selected navigation enhanced satellite combination is optimal under the constraints of multiple factors, adapting to the star selection needs in different complex scenarios.
Smart Images

Figure CN119959982B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of satellite navigation and artificial intelligence, and particularly relates to a satellite selection method and system for power enhancement of navigation satellites based on deep reinforcement learning. Background Art
[0002] The satellite orbits of the Global Navigation Satellite System (GNSS) are at a relatively high altitude and far from users, resulting in a low signal level reaching the ground, being vulnerable to interference or shielding, thus affecting the continuity and reliability of navigation services. To improve the signal anti-interference ability and ensure transmission stability, enhancing the navigation signal power is regarded as an effective means. By increasing the satellite signal power, the downlink signal strength can be significantly enhanced, the signal quality can be improved, and the stability and anti-interference performance of navigation services in complex environments can be enhanced. However, how to scientifically select the satellite combination participating in power enhancement to maximize the enhancement effect while avoiding resource waste is still a difficult problem to be solved urgently. Currently, the satellite selection method based on the traversal of the Geometric Dilution of Precision (GDOP) has significant limitations: one is the problem of combinatorial explosion, resulting in high computational complexity and insufficient real-time performance; the other is that static optimization is difficult to adapt to the dynamically changing interference environment. In recent years, deep reinforcement learning has shown unique advantages in the field of dynamic decision-making, but its application in satellite selection for navigation signal power enhancement is still blank, and relevant research urgently needs to break through. Summary of the Invention
[0003] To solve the above technical problems, the present invention provides a satellite selection method and system for power enhancement of navigation satellites based on deep reinforcement learning, which reduces the influence of subjective factors and enables the selected navigation enhancement satellite combination to reach the optimal under the constraints of various factors.
[0004] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0005] A satellite selection method for power enhancement of navigation satellites based on deep reinforcement learning, the method comprising:
[0006] Step 1, modeling the satellite selection task for power enhancement of navigation satellites as a Markov decision process, and defining a state space, an action space, and a reward function;
[0007] Step 2, inputting satellite data and enhancement task data of the navigation satellite system, and constructing a state vector of the navigation satellite system based on the Markov decision process;
[0008] Step 3, inputting the state vector of the navigation satellite system into a pre-trained power enhancement satellite selection model corresponding to the enhancement task data, and the power enhancement satellite selection model generates the task participation situation of each navigation satellite at different moments within the task time window according to the start time, end time of the task, and the time step of policy update;
[0009] Step 4: Based on the task participation of each navigation satellite at different times within the task time window, generate the overall satellite selection result during the task time.
[0010] On the other hand, the present invention provides a navigation satellite power enhancement satellite selection system based on deep reinforcement learning, including a task requirement input module, a satellite parameter database, a deep reinforcement learning decision module, and a situation visualization module. Among them,
[0011] The task requirement input module is used to receive satellite data and enhancement task data of the navigation satellite system, and construct a state vector of the navigation satellite system based on the Markov decision process;
[0012] The satellite parameter database module is used to store and manage relevant parameters of navigation satellites and enhancement task service data;
[0013] The deep reinforcement learning decision module is used to input the constructed vector of the navigation satellite system into a pre-trained power enhancement satellite selection model. The power enhancement satellite selection model generates the task participation of each navigation satellite at different times within the task time window according to the task start and end times and the policy update time step, and generates the overall satellite selection result during the task time based on the task participation of each navigation satellite at different times within the task time window;
[0014] The situation visualization module is used to visually display the satellite selection result and the navigation enhancement scheme based on the digital earth platform.
[0015] In the third aspect, the present invention provides an electronic device, including: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the foregoing method for selecting navigation satellites with enhanced power based on deep reinforcement learning.
[0016] In the fourth aspect, the present invention provides a computer-readable storage medium, on which executable instructions are stored, and when the instructions are executed by a processor, the processor is enabled to implement the foregoing method for selecting navigation satellites with enhanced power based on deep reinforcement learning.
[0017] The beneficial effects of the present invention are as follows:
[0018] An innovative state space representation method including relative geometric features such as satellite azimuth, elevation, and distance is constructed, and an action space decomposition strategy at the satellite granularity is developed to break through the traditional combinatorial optimization dimension disaster problem; a multi-objective reward function integrating GDOP, handover continuity, and enhanced energy consumption and its dynamic weight adjustment mechanism are designed to meet the satellite selection requirements in different complex scenarios; the number of selected satellites and the satellite number threshold can be dynamically adjusted according to the task scenario, and the visualization display of the satellite power enhancement effect can be performed. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a schematic flow chart of the method for selecting navigation satellites with enhanced power based on deep reinforcement learning according to the present invention;
[0020] Figure 2 is a schematic diagram of the composition structure of the navigation satellite power enhancement satellite selection system based on deep reinforcement learning according to the present invention;
[0021] Figure 3 is a graph of the satellite selection timing result based on the method of the present invention;
[0022] Figure 4 is a graph of the implementation visualization result based on the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] To make the above objects and features of the present application more clearly understood, the technical solutions in this embodiment will be described below in conjunction with the accompanying drawings in the embodiments. The embodiments described below are exemplary and are intended to explain the present application, and should not be construed as a limitation to the present application.
[0024] Referring to Figure 1 as shown, the present invention discloses a step flow chart of a method for selecting navigation satellites with enhanced power based on DQN. The method specifically includes the following steps:
[0025] Step 1: Model the navigation satellite power enhancement satellite selection task as a Markov decision process, define the state space S, action space A, and reward function R to provide a mathematical framework for the interaction between the agent and the environment;
[0026] In the embodiment of the present application, the state space S is used to describe the real-time state characteristics of the navigation satellite system, and the real-time state of the navigation satellite system is composed of the states of individual navigation satellites in the system.
[0027] The state of an individual navigation satellite is composed of the visibility of the satellite relative to the task target point or the center point of the task area, the relative position characteristics, and the task working state of the satellite at the previous time. As the task time progresses, the state of the navigation satellite system will also change accordingly.
[0028] Specifically, the visibility refers to whether the satellite is within the line of sight of the target point or the center point of the area. The relative position characteristics refer to the azimuth, elevation angle, and distance of the satellite relative to the task point. The task working state of the satellite at the previous time refers to whether the satellite participated in or did not participate in the navigation power enhancement task at the previous time point. For example, if the number of satellites in the navigation satellite system is N, the state of the navigation satellite system is a vector of 5 N, and the state space of the navigation satellite system at time T Defined as:
[0029] ,
[0030] wherein, represents the visibility of the -th satellite, 0 means invisible, 1 means visible; respectively represent the azimuth angle, elevation angle and distance of the -th satellite relative to the mission point; represents the working state of the -th satellite, 0 means not participating, 1 means participating in the mission.
[0031] In the embodiment of the present application, the action space A is used to describe the satellite selection decision made by the navigation satellite system for the mission at time T, and the decision continues until time T + 1. The satellite selection decision of the system is composed of the single-satellite mission participation decision; the single-satellite mission participation decision refers to whether the satellite participates or does not participate in the navigation power enhancement mission between time T and time T + 1; the time interval between time T and time T + 1 is the time update step. . Exemplarily, at time T, the action space of the satellite selection decision of the navigation satellite system containing N satellites is
[0032] ,
[0033] wherein, the action of each satellite is defined as , 0 means not participating in the mission, 1 means participating in the mission;
[0034] In the embodiment of the present application, the reward function R is used to evaluate the quality of the satellite selection decision made by the navigation satellite system. The reward function R comprehensively considers GDOP, service continuity and system task enhancement cost, and adopts the form of normalized weighted sum. Specifically, GDOP refers to the Geometric Dilution of Precision, which is used to describe the influence of the geometric distribution of the satellites participating in the enhancement mission on the positioning error of the mission target point or the center of the mission area; service continuity refers to the ability of the navigation satellite system to provide continuous and stable navigation services for the mission target point or the center receiver of the mission at different times; the service continuity is quantified by the change amount of the single-satellite mission participation status between adjacent satellite selection task decisions. The enhancement cost refers to the resource consumption involved in the process of the navigation satellite system completing the power enhancement mission, and is represented by the sum of the distances between the satellites participating in the mission and the mission point or the mission target.
[0035] Exemplarily, at time T, the reward function of the satellite selection decision for the power enhancement mission of the navigation satellite system is defined as:
[0036] ,
[0037] Among them, are the weight coefficients corresponding to GDOP, service continuity, and enhanced cost, satisfying . , representing the geometric accuracy reward, where GDOP is the geometric dilution of precision factor, is the historical maximum value of GDOP;
[0038] represents the decision continuity reward, N represents the total number of satellites in the system, represents the th satellite's task participation status at T - 1 and T moments, ⊕ represents the exclusive - OR operation, and the smaller the difference in the participation status of adjacent - time satellites, the higher the reward;
[0039] is the sum of the distances between the satellites participating in the task and the task points, is the historical maximum value, and the smaller the sum of the distances, the higher the reward.
[0040] Step 2: Input the satellite data and enhanced task data of the navigation satellite system, and construct the state vector of the navigation satellite system based on the Markov decision process; the satellite data refers to the two - line element data (TLE) used to calculate the spatial coordinates of satellites at different times; the enhanced task data specifically includes the longitude and latitude coordinates of the enhanced target point or the center of the enhanced area, the task start time and task end time, the policy update time step, and the power enhancement task mode. The construction of the state vector of the navigation satellite system means constructing it according to the state space established in Step 1;
[0041] The policy update time step refers to the time interval for updating or re - evaluating the satellite selection strategy, which is used to update the state change of the satellite's spatial position; the power enhancement task mode can be divided into the balanced mode, high - precision mode, stable - connection mode, and energy - saving mode. Specifically, the balanced mode means adopting a balanced weighting strategy for each index of GDOP, service continuity, and enhanced cost, such as the weight distribution is ; the high - precision mode means giving priority to optimizing the GDOP index to meet the high positioning accuracy requirements of the navigation power enhancement task, such as the weight distribution is ; the stable - connection mode means giving priority to optimizing service continuity to meet the consistency requirements of the navigation power enhancement task, such as the weight distribution is ; the energy - saving mode means giving priority to optimizing the enhanced cost to reduce resource consumption, such as the weight distribution is .
[0042] Step 3: Input the state vector of the navigation satellite system into the pre-trained power enhancement satellite selection model corresponding to the enhancement task data. The power enhancement satellite selection model generates the task participation status of each navigation satellite at different moments within the task time window according to the task start time, end time, and policy update time step; the corresponding pre-trained power enhancement satellite selection model is a power enhancement satellite selection model trained separately with different weight distributions under different power enhancement task modes.
[0043] In the embodiment of the present application, the power enhancement satellite selection model has pre-learned the satellite selection actions for power enhancement of the navigation satellite system according to the generation strategy. The power enhancement satellite selection model can generate the Q value of each navigation satellite participating or not participating in the enhancement task. The Q value refers to the total expected future reward obtained after taking a certain action in a given state. Specifically, if there are N satellites in the navigation satellite system, the power enhancement satellite selection model outputs a Q value vector of size 2 N. Each satellite corresponds to two Q values for participating and not participating. When making a decision, the maximum Q value criterion is adopted to obtain whether each navigation satellite participates in the enhancement task at time T.
[0044] Among them, the power enhancement satellite selection model is based on randomly generated enhancement task data. The power enhancement satellite selection model is used to generate a strategy, and training samples are constructed according to the generated strategy. The power enhancement satellite selection model is iteratively trained. After sufficient number of enhancement task trainings, the trained power enhancement satellite selection model is obtained.
[0045] In an optional embodiment, the power enhancement satellite selection model is trained according to the following steps:
[0046] Step 3.1: Construct a power enhancement satellite selection model, and use the power enhancement satellite selection model to perform satellite selection to obtain a satellite selection strategy. In the example of the present application, the power enhancement satellite selection model adopts the Deep Q-Network (DQN) algorithm to optimize the satellite selection decision-making process of the navigation satellite through the reinforcement learning method;
[0047] In the embodiment of the present application, when using the initialized DQN model to perform the navigation satellite power enhancement satellite selection task, the input is a state vector with a dimension of 5N. Specifically, the network structure of the DQN adopts 5 fully connected layers, and the dimensions of the hidden layers are 256, 512, 256, and 128 in sequence. Batch normalization is equipped after each layer to improve the training stability. The number of nodes output by the fifth fully connected layer is 2 N, which respectively represent the Q values of each satellite choosing not to participate in the task and participating in the task. The output is a vector used to characterize the value estimation of each satellite in the current state, where N represents the total number of navigation satellites.
[0048] Step 3.2: Calculate the reward value of the satellite selection strategy, and store the satellite selection experience in the experience pool, where represents the state of the navigation satellite system at time T; represents the satellite selection action at time T, indicating the decision options of the satellite; represents the reward function at time T, measuring the quality of the satellite selection strategy; represents the state of the navigation satellite system at time T+1, obtained from the current action and time step;
[0049] The capacity of the experience pool is set to a fixed value (e.g., 10,000). When the capacity is full, the first-in-first-out strategy is adopted to overwrite the previous experience. When the number of stored strategies in the experience pool reaches the training threshold (e.g., 500), a small batch of samples (e.g., 32) are randomly drawn from the experience pool for training the DQN model;
[0050] Step 3.3: Conduct the main network training. This application adopts the target network technology to stabilize the training process by introducing an independent network structure. The target network has the same structure as the main network, and its parameters are consistent with the main network in the initial stage. The target network parameters are updated every time the step size G is reached.
[0051] The main network training is to calculate the estimated value of the target network function and the value predicted by the network. Further, the weights and biases and other parameters of the target network are updated by the stochastic gradient descent method. When the step size G is reached, the target network parameters are updated . The error value is calculated by the error function as follows:
[0052] ,
[0053] where E represents the expected value, indicating the average error of a small batch of randomly drawn training samples; is the reward at the j-th time step, representing the value feedback of the current satellite selection decision; is the discount factor; represents the target network taking the action in the next state with the maximum Q value, and the network parameters are ; represents the Q value function of the main network Q in the current state and the action with the parameter θ;
[0054] Step 3.4, termination condition check. If the current number of selected satellites is less than 4, or the task end time is reached, start the next iteration; otherwise, maintain the current policy and jump to the next moment, and return to Step 3.2;
[0055] Step 3.5, end of training. After completing the training of all random tasks, the algorithm ends, and a trained power enhancement satellite selection model is obtained.
[0056] Step 4: Generate the overall satellite selection result during the task time based on the task participation of each navigation satellite at different times within the task time window.
[0057] In the embodiment of the present application, based on the task participation of each navigation satellite at different time slices output by the power enhancement satellite selection model in Step 3, the overall satellite selection result from the task start time to the task end time is generated. The specific implementation method is as follows:
[0058] Step 4.1, input task participation. Receive the task participation decision of each navigation satellite at each moment T . The = 1 indicates that the i-th satellite participates in the power enhancement task at moment T, = 0 indicates that the i-th satellite does not participate in the power enhancement task at time slice T;
[0059] Step 4.2, integrate the task participation of each navigation satellite at each moment within the task time range (from the start time to the end time, with a time step of to divide time slices), generate a time-sequential task participation data set, and obtain the task satellite selection result for navigation power enhancement. For example, if there are N satellites in the navigation satellite system and the task lasts for 24 hours, = 10 minutes, then there are 144 moments, and the result data set is:
[0060] .
[0061] Referring to Figure 2 shown, the present invention also provides a structural schematic diagram of a power enhancement satellite selection system based on reinforcement learning. Each module is connected by a data bus, and the data flow is smooth, ensuring the efficient operation of the system. The core of the present invention lies in constructing a satellite power enhancement dynamic decision system based on deep reinforcement learning. The power enhancement satellite selection system based on deep reinforcement learning includes:
[0062] (1)Task Requirement Input Module: It is used to receive satellite data and enhanced task data of the navigation satellite system, construct the state vector of the navigation satellite system based on the Markov decision process, and provide input for the satellite selection decision of the system. It includes but is not limited to: the longitude and latitude coordinates of the center point of the target area, the start and end times of the task, the time step (such as 10 minutes), and the task mode (such as balanced mode, high-precision mode, stable connection mode, or energy-saving mode). The input data is transmitted to the satellite parameter database module as one of the basic data for satellite selection decision. For example, the user can input that the target point is (100°E, 40°N), the task start time is 2025.01.01 00:00:00, the task end time is 2025.01.01 00:00:00, the task time step is 10 minutes, and the mode is high-precision mode.
[0063] (2)Satellite Parameter Database Module: This module is used to store and manage relevant parameters of navigation satellites and enhanced task service data to support real-time calculation of satellite selection decision. The satellite parameter database specifically includes the orbital parameters of the satellite (such as two-line element set (TLE) data), position parameters (space coordinates calculated by the SGP4 model) and working status. For example, the database stores the TLE data of GPS satellites and calculates the position information of each satellite within the task time window in real time through the SGP4 model.
[0064] (3)Deep Reinforcement Learning Decision Module: This module inputs the vector of the constructed navigation satellite system into a pre-trained power enhancement satellite selection model. The power enhancement satellite selection model generates the task participation status of each navigation satellite at different moments within the task time window according to the start and end times of the task and the policy update time step. Based on the task participation status of each navigation satellite at different moments within the task time window, it generates the overall satellite selection result during the task time. The deep reinforcement learning decision module receives the data and service data of the satellite parameter database, constructs the current environmental state, inputs the state into the pre-trained power enhancement satellite selection model, and the model outputs the optimal satellite combination in time series. The result is as Figure 3 shown.
[0065] (4)Situation Visualization Module: This module is based on the digital earth platform to visually display the satellite selection result and the navigation enhancement plan. The result visualization module specifically includes a basic situation visualization module and an enhanced situation visualization module. The basic situation visualization module is used for visualizing the surface environment, space environment, and satellite models; the enhanced situation visualization module is used for dynamic visualization of the satellite trajectories participating in the task, the satellite enhancement task combinations under task time slices, and the switching situations. The visualization result is as Figure 4 shown.
[0066] In a third aspect, the present invention provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the foregoing method for selecting navigation satellites for power enhancement based on deep reinforcement learning.
[0067] In a fourth aspect, the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, enable the processor to implement the foregoing method for selecting navigation satellites for power enhancement based on deep reinforcement learning.
[0068] The specific embodiments described above further elaborate on the object, technical solution, and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A navigation satellite power enhancement selection method based on deep reinforcement learning, characterized in that: The method comprises: Step 1: Model the navigation satellite power enhancement and selection task as a Markov decision process, and define the state space, action space, and reward function; Step 2: inputting satellite data and augmented mission data of the navigation satellite system, and constructing a state vector of the navigation satellite system based on the Markov decision process; Step 3: input the state vector of the navigation satellite system into a pre-trained power enhancement satellite selection model corresponding to the enhanced mission data, wherein the power enhancement satellite selection model generates the mission participation status of each navigation satellite at different times within the mission time window according to the mission start and end time and the strategy update time step; Step 4: Based on the mission participation of each navigation satellite at different times within the mission time window, generate the overall satellite selection result within the mission time.
2. According to claim 1, a navigation satellite power enhancement star selection method based on deep reinforcement learning is characterized in that: The state space is used to describe the real-time state characteristics of the navigation satellite system. The real-time state of the navigation satellite system is composed of the state of a single navigation satellite in the system. The state of the single navigation satellite is composed of the visibility and relative position characteristics of the satellite relative to the mission target point or the center point of the mission area and the mission working state of the satellite at the previous time; The action space is used to describe the satellite selection decision taken by the navigation satellite system for the mission at time T, where the action of each satellite is defined as , 0 means not participating in the task, 1 means participating in the task; The reward function is used to evaluate the quality of the satellite selection decision taken by the navigation satellite system, including three factors: GDOP, service continuity and enhanced cost evaluation.
3. According to claim 2, a navigation satellite power enhancement star selection method based on deep reinforcement learning is characterized in that: The visibility refers to whether the satellite is within the line of sight of the target point or the center point of the area; the relative position characteristics refer to the azimuth, altitude angle and distance of the satellite relative to the mission point; the mission working status of the satellite at the last time refers to whether the satellite participated in or did not participate in the navigation power enhancement mission at the last time.
4. According to claim 2, a navigation satellite power enhancement star selection method based on deep reinforcement learning is characterized in that: The reward function Defined as: , , , , in, is the weight coefficient corresponding to GDOP, service continuity and enhancement cost, satisfying , represents the geometric precision bonus, GDOP is the geometric precision factor, This is the historical maximum value of GDOP; represents the decision continuity reward, N represents the total number of satellites in the system, Respectively represent The mission participation status of the satellites at time T-1 and T. ⊕ represents an XOR operation. The smaller the difference in the satellite participation status at adjacent times, the higher the reward. represents the enhanced cost reward, is the sum of the distances between the satellites participating in the mission and the mission point, for The historical maximum value. The smaller the sum of distances, the higher the reward.
5. According to the method of claim 1, wherein: In step 2, the enhanced task data includes the longitude and latitude coordinates of the enhanced target point or the center of the enhanced area, the task start time and the task end time, the strategy update time step and the power enhancement task mode, and the enhanced task mode includes a balanced mode, a high-precision mode, a stable connection mode and an energy-saving mode.
6. The method for selecting navigation satellites for power enhancement based on deep reinforcement learning according to claim 5, characterized in that: In step 3, the pre-trained power enhancement star selection model corresponding to the enhancement task data is a power enhancement star selection model trained respectively for different weight allocations under different power enhancement task modes.
7. The method for selecting navigation satellites for power enhancement based on deep reinforcement learning according to claim 6, characterized in that: The trained power enhancement star selection model is trained according to the following steps: Step 3.1, construct a power enhancement star selection model, randomly generate tasks and use the power enhancement star selection model to perform star selection to obtain a star selection strategy; Step 3.2, calculate the reward value of the star selection strategy, store the star selection experience in the experience pool, set the capacity of the experience pool to a fixed value, and adopt a first-in-first-out coverage strategy. When the number of stored strategies reaches the training threshold, randomly extract small batches of samples from the experience pool for training the model; Step 3.3: Train the main network using the target network technology. The target network has the same structure as the main network, and its parameters are consistent with those of the main network in the initial stage. Every time the step length G is reached, the target network parameters are updated. Step 3.4, check the termination condition. If the current number of selected stars is less than 4, or the task end time is reached, start the next round of iteration. Otherwise, keep the current strategy running and jump to the next moment, returning to step 3.2; Step 3.5, end the training. After completing the training of all random tasks, the algorithm ends and a trained power enhancement star selection model is obtained.
8. The method for selecting navigation satellites for power enhancement based on deep reinforcement learning according to claim 1, characterized in that: The step 4 comprises: Step 4.1, inputting the mission participation status, wherein the mission participation status is the mission participation decision state of each navigation satellite at each time T; Step 4.2: Integrate the mission participation status of each navigation satellite at each time within the mission time range, generate a time-series mission participation data set, and obtain the satellite selection result for the navigation power enhancement mission.
9. A navigation satellite power enhancement and star selection system based on deep reinforcement learning, characterized in that: It includes mission requirement input module, satellite parameter database, deep reinforcement learning decision module and situation visualization module, among which: The mission requirement input module is used to receive satellite data and enhanced mission data of the navigation satellite system, and construct the state vector of the navigation satellite system based on the Markov decision process; Satellite parameter database module, used to store and manage relevant parameters of navigation satellites and enhanced mission business data; A deep reinforcement learning decision module is used to input the vector of the constructed navigation satellite system into a pre-trained power enhancement satellite selection model, and the power enhancement satellite selection model generates the mission participation status of each navigation satellite at different times within the mission time window according to the mission start and end time and the strategy update time step, and generates the overall satellite selection result within the mission time based on the mission participation status of each navigation satellite at different times within the mission time window; The situation visualization module is used to visualize the satellite selection results and navigation enhancement plans based on the digital earth platform.
10. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; Among them, when one or more programs are executed by the one or more processors, the one or more processors implement the navigation satellite power enhancement selection method based on deep reinforcement learning as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that: Executable instructions are stored thereon, which, when executed by a processor, enable the processor to implement a navigation satellite power enhancement and star selection method based on deep reinforcement learning as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Beidou navigation satellite selection method under tabu search artificial bee colony algorithm
CN112578414A
Communication satellite multi-beam resource management and control method based on deep reinforcement learning
CN119727863A