Metareinforcement learning scheduling method and device for earth observation satellite task planning
By introducing a dynamic task reward mechanism and meta-learning layer, an adaptive dynamic Markov decision-making environment and deep reinforcement learning algorithm are built, which solves the problem of insufficient processing capabilities of traditional satellite mission planning methods under large-scale dynamic tasks and complex constraints, and realizes efficient and flexible satellite mission planning.
Patent Information
- Application Number
- CN202510847129.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When traditional satellite dynamic task planning methods face large-scale dynamic tasks and complex constraints, they lack processing capabilities and are difficult to respond to sudden tasks quickly. In addition, the existing reinforcement learning algorithms have insufficient generalization capabilities and computing efficiency in dynamic environments, making it impossible to achieve global optimal task planning.
A dynamic task reward mechanism is introduced, and a meta-learning layer is used to capture the commonalities of the task. Through the meta-reinforcement learning scheduling method, an adaptive dynamic Markov decision-making environment and deep reinforcement learning algorithm are constructed to realize cross-scene strategy migration and improve the efficiency and flexibility of satellite mission planning.
The efficiency and stability of satellite mission planning strategies have been improved, and the agents quickly adapt and make efficient decisions in a changing environment, improving the utilization rate of satellite resources and the real-time nature of mission planning.
Smart Images

Figure CN120355195A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of satellite remote sensing, and in particular to a meta-reinforcement learning scheduling method and device for earth observation satellite mission planning. Background Art
[0002] Remote sensing satellite mission planning abstracts mission requirements into point targets, and based on the analysis of mission requirements and satellite resource capacity constraints, establishes and solves a mathematical model for satellite mission planning, aiming to reasonably arrange satellite observation tasks for ground targets within a specific time to maximize observation efficiency and scientific value. With the continuous increase in the number of satellites and the increasing diversification of observation requirements, the sources of remote sensing satellite observation requirements are extensive and show extremely high dynamics in the time series. The traditional static task planning method oriented to tasks has been difficult to meet the actual needs.
[0003] Against the background of the sharp increase in the number of requirements, the dynamic changes in requirements have become more and more significant. If the traditional "offline" planning method is still used, it is obviously unable to adapt to the application scenarios of future large-scale task growth. At present, the satellite mission planning methods considering dynamic tasks are mainly complete replanning modeling and incremental replanning methods. The complete replanning abandons the existing observation plan. After the dynamic task is added, a new observation plan is completely re-formulated from scratch. This method takes a long time for mission planning calculation. When the number of dynamic tasks to be observed is large, it greatly increases the method solving time and reduces the satellite resource utilization rate. The incremental replanning modeling method makes appropriate adjustments on the basis of the existing observation plan, such as adopting a modeling method with the ability to increment dynamic tasks or setting a sliding window, and re-planning the tasks within the window as the sliding window moves. [1] However, the results of such task planning methods are easily affected by the perturbation of the added tasks to the original model. The increase in the number of task replannings will increase the model size, resulting in a rapid increase in the subsequent solution dimension and increasing the solution difficulty. When sudden tasks occur, the existing methods often lack the ability of rapid response and flexible adjustment. Sudden tasks may have extremely high priorities and urgent time requirements. When dealing with such tasks, the existing methods are unable to reasonably reconstruct the original task plan in time due to the complex decision-making process and slow response speed, which easily leads to the aggravation of task conflicts and affects the stability and reliability of the entire task planning system.
[0004] Regarding the dynamic scheduling algorithm for satellites, existing research has proposed deterministic algorithms and (meta) heuristic algorithms as the main methods. This type of method has strong description ability for application scenarios and relatively mature solving algorithms. However, it still has limitations. On the one hand, this type of method needs to customize rules for different scheduling scenarios, resulting in poor generalization ability and difficulty in adapting to the modeling requirements of large-scale task scheduling. Reinforcement learning is a machine learning method suitable for dynamic task planning problems, aiming to maximize the cumulative reward, with advantages such as strong generalization ability and good scalability. In satellite task planning, the current application research of reinforcement learning mainly focuses on two aspects: one is to combine the reinforcement learning algorithm with the optimization algorithm to jointly solve the satellite task planning model; the other is to solve it through the reinforcement learning method on the basis of constructing a scheduling model based on the Markov process to obtain the optimal task scheduling scheme. [2] . However, most of the current reinforcement learning algorithms are designed and trained for specific satellite task scenarios and task types, and have not considered dynamic task scenarios, and have strong dependence on the training environment and task distribution. Once the task environment changes, such as the emergence of new task types, changes in target distribution, or different resource constraint conditions, the generalization ability of the algorithm is significantly insufficient, it is difficult to quickly adapt to the new environment, and it is unable to effectively generate reasonable task planning strategies. Especially when dealing with large-scale tasks and complex constraint conditions, the computational complexity of the algorithm increases exponentially, resulting in too long calculation time and difficulty in meeting the real-time requirements of satellite task planning. In addition, the design of the existing reinforcement learning reward mechanism is relatively simple, and it is difficult to comprehensively and accurately reflect the complex constraints and objectives in satellite task planning. In a dynamic task environment, the reward mechanism often cannot give the intelligent agent accurate feedback in a timely and effective manner according to the dynamic changes of the task (such as priority, time window, resource competition, etc.), resulting in the intelligent agent being prone to falling into local optimal solutions during the learning process and unable to achieve the global optimal task planning.
[0005] The relevant reference documents are as follows: [1] NIU X, TANG H, WU L. Satellite scheduling of large areal tasks for rapid response to natural disaster using a multi-objective genetic algorithm[J]. International Journal of Disaster Risk Reduction, 2018, 28: 813-825.
[0006] [2] CHEN Y, SHEN X, ZHANG G, etc. Multi-Objective Multi-Satellite Imaging Mission Planning Algorithm for Regional Mapping Based on Deep Reinforcement Learning[J]. Remote Sensing, 2023, 15(16). Summary of the Invention
[0007] To solve the problem that the traditional satellite dynamic task planning method has obvious insufficient processing ability in the face of large-scale dynamic tasks and complex constraints, the present invention provides a meta-reinforcement learning scheduling method for earth observation satellite task planning. By introducing a task dynamic reward mechanism and using the meta-learning layer to capture task commonalities to achieve cross-scenario policy transfer, the efficiency of solving the satellite task planning strategy is improved.
[0008] According to one aspect of the specification of the present invention, there is provided a meta-reinforcement learning scheduling method for earth observation satellite task planning, including: Obtain the visible time window of the satellite for the ground target, including the observation start time and the end time; Input task data and initialize the environment. At each time step, read the policy of the trained reinforcement learning model and input the observation start time and the end time of the feasible ground target, and output the imaging task planning policy; wherein, the training of the reinforcement learning model includes: Model the satellite imaging task execution process as a dynamic Markov decision process with adaptive rewards, define the state space, action space, reward, state transition probability and discount factor, where the action is represented as selecting an observable task; after the dynamic task is added, the dimension of the action space is expanded in real time; initialize the built Markov decision environment and initialize the PPO algorithm, and establish a policy network and a value network; Build a deep reinforcement learning algorithm based on meta-learning, introduce memory for the agent using a recurrent neural network, transform the learning between tasks into a sequence modeling problem, introduce a meta-learning layer to train multiple tasks and learn how to quickly adapt to new tasks; Update the policy network parameters and the value network according to the objective function of the PPO algorithm, minimize the mean square error of the value function estimation, continuously iterate the training, and obtain the optimal scheduling policy in the Markov decision environment when the training converges; Set different learning rates, discount factors, clipping coefficients, value function coefficients, entropy coefficients, and at the same time set different training batch sizes and training epochs, start the training of the deep reinforcement learning algorithm based on meta-learning, and store the algorithm training results when the training converges.
[0009] As a further technical solution, building a Markov decision environment further includes: Building an adaptive dynamic reward space, where the dynamic reward function includes task priority reward, task time reward, and task dynamic reward; Building an adaptive dynamic action space, defining the maximum action space dimension, covering all static tasks and defining null actions for dynamic tasks, setting reserved positions for the number of dynamic tasks. At each time step, according to the current satellite state and task constraint conditions, calculate the effectiveness of each action.
[0010] As a further technical solution, building a Markov decision environment further includes: In the task selection strategy, give priority to dynamic tasks; Use the imaging task priority to construct the reward in the Markov process; Segment the time scheduling period according to the target demand refresh period. The satellite state and imaging task observation situation of each time slice are used as state information. Different rewards are given to the states feedback by different actions in the environment, and the satellite observation sequence is obtained with the goal of maximizing the reward; Transform the multi-target imaging task scheduling problem of the earth observation satellite into an optimal strategy solving problem of an intelligent agent, and obtain the optimal action strategy in time series.
[0011] As a further technical solution, building a deep reinforcement learning algorithm based on meta-learning further includes: Building a meta-learning network architecture and a reinforcement learning framework, and constructing an input gate; Meta-training task sampling and initialization, sampling several meta-training tasks from the task distribution, where each task corresponds to different satellite orbit parameters, resource constraints, or dynamic task combinations; for each meta-training task, perform multiple rounds of interaction, and the hidden state is passed between rounds, and the final hidden state is retained at the end of the round as the initial state of the next round; According to the recursive neural network unit state update rule, combine the forget gate, the unit state at the previous moment, the input gate, and the candidate unit state to obtain the updated unit state in the meta-learning layer; The meta-learner updates the parameters of the base learner using the calculated input gate, forget gate, and updated unit state according to its own parameters.
[0012] As a further technical solution, building a deep reinforcement learning algorithm based on meta-learning further includes: Randomly select tasks from the task set and generate relevant task statuses to start training. In each round of training, the agent interacts with the environment according to the current policy, collects state, action, and reward information, and uses this information to update the policy network parameters and value network parameters using the PPO algorithm. The objective functions are the clipped policy loss and the mean squared error of the value function, gradually optimizing the policy until the predetermined number of training rounds is reached or the algorithm converges.
[0013] As a further technical solution, the training of the reinforcement learning model further includes: Step A, initialize the initial state information of all satellites and targets; for different imaging task sets, load the saved models. Step B, in each training time step, the agent interacts with the environment to collect trajectory data during the training process. Each trajectory consists of state, action, reward, and the next state; advance the time step according to the scheduling period to obtain the current satellite state; input the current satellite state and the hidden state into the policy network to output the action probability distribution; select the task with the highest probability in the discrete action space of the satellite according to the current state using the policy, and mask the inexecutable actions through constraint judgment. Step C, execute the action after the constraint judgment, consume resources and update the task queue, and the environment returns the next state and reward. Step D, when a dynamic task is inserted, pause the current training process, expand the action space, the output layer of the policy network dynamically increases the corresponding action dimensions, and recalculate the policy and reward function based on the updated action space; record the tasks executed at each time step, including time, remote sensing satellite, target, and the execution result of the imaging task. Step E, repeat Steps A to D until the visible time window of all targets ends or the predetermined imaging task time range is reached. After the end, obtain the remote sensing satellite for the imaging task execution of each point target, the start observation time, and the end observation time, forming an imaging task planning scheme for the remote sensing satellite for different point targets.
[0014] According to one aspect of the specification of the present invention, a meta-reinforcement learning scheduling device for earth observation satellite task planning is provided, including: The first main module is used to obtain the visible time window of the satellite for the ground target, including the observation start time and the end time. The second main module is used to input task data and initialize the environment, read the policy of the trained reinforcement learning model at each time step and input the observation start time and the end time of the feasible ground target, and output the imaging task planning policy; wherein, the training of the reinforcement learning model includes: Model the satellite imaging task execution process as a dynamic Markov decision process with adaptive rewards, define the state space, action space, rewards, state transition probabilities, and discount factor, where the action is represented as selecting observable tasks; after adding dynamic tasks, the dimension of the action space is expanded in real time; initialize the constructed Markov decision environment and initialize the PPO algorithm to establish a policy network and a value network; Build a deep reinforcement learning algorithm based on meta-learning, use a recurrent neural network to introduce memory for the agent, transform the learning between tasks into a sequence modeling problem, and introduce a meta-learning layer to train multiple tasks and learn how to quickly adapt to new tasks; Update the policy network parameters and value network according to the objective function of the PPO algorithm, minimize the mean square error estimated by the value function, and continuously iterate the training to obtain the optimal scheduling strategy in the Markov decision environment when the training converges; Set different learning rates, discount factors, clipping coefficients, value function coefficients, entropy coefficients, and at the same time set different training batch sizes and training epochs, start the training of the deep reinforcement learning algorithm based on meta-learning, and store the training results of the algorithm when the training converges.
[0015] According to one aspect of the specification of the present invention, there is provided a meta-reinforcement learning scheduling device for earth observation satellite task planning, including one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement a meta-reinforcement learning scheduling method for earth observation satellite task planning as described in the above technical solution.
[0016] According to one aspect of the specification of the present invention, there is provided a non-transitory computer-readable storage medium storing computer instructions that cause the computer to execute the method.
[0017] Traditional satellite dynamic task planning methods are solved by optimization methods such as partial replanning or full replanning, and face large-scale dynamic tasks and complex constraints, with obvious insufficient processing capabilities. Compared with the prior art, the beneficial effects of the present invention are as follows: The method of the present invention proposes a meta-reinforcement learning method suitable for dynamic satellite tasks, introduces a task dynamic reward mechanism, and enables the satellite to automatically adjust the strategy according to task changes. The meta-learning layer is used to capture task commonalities to achieve cross-scenario policy transfer and improve generalization performance. Through multiple rounds of training and interaction with the environment to update the network, the agent can quickly adapt and make efficient decisions in a changing environment. Compared with other reinforcement learning algorithms, it shows higher stability, convergence speed, and robustness in performance, and improves the efficiency of solving the satellite task planning strategy. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the attached drawings used in the description of the embodiments or the prior art. Obviously, the attached drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other attached drawings can also be obtained based on these attached drawings.
[0019] Figure 1 It is a schematic flowchart of the method provided by the embodiment of the present invention.
[0020] Figure 2 It is a schematic diagram of the meta-reinforcement learning algorithm provided by the embodiment of the present invention.
[0021] Figure 3 It is a schematic diagram of the target distribution provided by the embodiment of the present invention. Detailed implementation manners
[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the attached drawings in the embodiments of the present invention. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention. In addition, the technical features in each embodiment or a single embodiment provided by the present invention can be combined arbitrarily with each other to form a new technical solution. This combination is not restricted by the order of steps and / or the mode of structural composition, but must be based on the fact that those of ordinary skill in the art can implement it. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0023] The following will introduce Figures 1 to 3 A meta-reinforcement learning scheduling method for earth observation satellite mission planning provided by the detailed implementation manners of the present invention, including the following steps: Step 1, satellite visibility window preprocessing: Obtain the visible time window of satellite i for ground target j , and respectively represent the start time and the end time.
[0024] Step 2, Construction of the Adaptive Dynamic Reward Task Markov Decision Environment: It includes modeling the satellite imaging task execution process as a multi-agent Markov decision process, defining the state space, action space, reward, state transition probability, and discount factor, where the action is represented as selecting an observable task and is constrainedly expressed in each training step; after adding dynamic tasks, the dimension of the action space is expanded in real time; initialize the Markov decision environment and initialize the PPO algorithm to establish a policy network and a value network.
[0025] Step 3, Construction of the MetaPPO, a Deep Reinforcement Learning Algorithm Based on Meta-Learning: By using a recurrent neural network (LSTM) to introduce memory for the agent, the learning between tasks is transformed into a sequence modeling problem, and a meta-learning layer is introduced to train multiple tasks and learn how to quickly adapt to new tasks; update the policy network parameters according to the objective function of the PPO algorithm, and at the same time update the value network to minimize the mean square error of the value function estimation. Through continuous iterative training, the optimal scheduling policy is obtained when the training converges. Set different learning rates, discount factors, clipping coefficients, value function coefficients, and entropy coefficients, and at the same time set different training batch sizes and training epochs, start training, and store the algorithm training results when the training converges.
[0026] Step 4, Load the MetaPPO algorithm, input the test data and initialize the environment, read the policy at each time step and input the start time and end time of the observable target of the feasible target, and output the imaging task planning policy.
[0027] The specific implementation of Step 1 includes the following sub-steps: Step 1.1, the preprocessing of the satellite visible window includes: predicting the target visible window according to the orbit and the maximum field of view angle of the remote sensing satellite; assuming that the remote sensing satellite moves in a two-body model, predicting the visible window of the satellite for each task point target through the six orbital elements of the satellite, and obtaining the visible time window of satellite i for ground target j respectively 。
[0028] Step 1.2, initialize the remote sensing satellite information. The known six orbital elements of the satellite are the semi-major axis a, eccentricity e, inclination , right ascension of the ascending node , argument of perigee , mean anomaly , the epoch moment is , and the field of view angle of the remote sensing satellite is 。
[0029] Step 1.3, calculate the mean anomaly of the satellite:
[0030] where \(t\) is the current time, is the epoch moment; \(n\) is the average motion speed of the satellite, and the calculation formula is as follows:
[0031] where \(G\) is the gravitational constant, \(M\) is the mass of the central celestial body, i.e., the mass of the Earth, \(k\) is the Kepler constant, is the semi-major axis of the satellite orbit; Solve the Kepler equation to obtain the eccentric anomaly \(E\):
[0032] \(E\) is the eccentric anomaly, \(e\) is the orbital eccentricity, \(M\) is the mean anomaly, and \(t\) is the current time; Calculate the true anomaly :
[0033] where \(E\) is the eccentric anomaly and \(e\) is the orbital eccentricity; Calculate the geocentric coordinates of the satellite:
[0034] Calculate the coordinates of the remote sensing satellite in the geocentric inertial coordinate system according to the orbital parameters :
[0035] Step 1.4, calculate the geocentric coordinates of the ground target , \(R\) is the radius of the Earth;
[0036] where \(R\) is the radius of the Earth and \(h\) is the height of the ground target, represents the longitude of the ground target, represents the latitude of the ground target.
[0037] Step 1.5, calculate the distance between the satellite and the geocenter:
[0038] Step 1.6, calculate the distance between the target and the geocenter:
[0039] Step 1.7, determine whether the line connecting the satellite and the target passes through the Earth;
[0040] where is the angle between the satellite and the target, and the calculation formula is:
[0041] If the above equation holds, the target is not blocked by the Earth; otherwise, the target is blocked by the Earth.
[0042] Step 1.8, calculate the angle between the remote sensing satellite and the ground target, and compare it with the field of view angle of the remote sensing satellite for judgment. When then the ground target is within the field of view of the remote sensing satellite.
[0043]
[0044] Step 1.9, starting from the epoch time, every Repeat steps 1.1 - 1.7 to calculate the visibility of the mission target once, and obtain the visible time window of each remote sensing satellite i for each ground target j .
[0045] The specific implementation of step 2 includes the following sub - steps: Step 2.1, model the process of the remote sensing satellite performing multiple ground point - target imaging tasks as the construction of an adaptive - reward dynamic Markov decision - making environment, which is determined by the five - tuple . They are the state space, action space, reward, state - transition probability, and discount factor respectively. Abstract the remote sensing satellite as an agent. A remote sensing satellite has N imaging tasks to be completed. Regard the imaging tasks of the remote sensing satellite as point targets. Each target i has a fixed geographical location , has multiple visible observation time windows , and each visible time window contains the start time and end time of visibility ; The state describes the joint state of the imaging task target state and the remote sensing satellite resources, includes the observation state of the imaging task target, describes the satellite state, including the list of imaging tasks completed by the remote sensing satellite and the current working state. Among them, A represents the action space. The action of the agent is to select observable tasks, and the number of tasks is num , then the action space is .
[0046] Step 2.2, calculate the attitude conversion time of the imaging task target in the time window for observing the target at each time slice within the visible duration ; Let be the position vector of the target in the geocentric inertial system at the specified moment, be the target The geocentric inertial system position vector at a specified moment, and the target direction vector is calculated through the position vectors of the remote sensing satellite and the target , ; Normalize to obtain the unit direction vector ; Construct the target direction cosine matrix DCM, and use the orthogonal basis matrix of the direction vector as the target direction vector , assuming that the initial attitude matrix points to the target , calculate the attitude error matrix , refers to the attitude of the satellite at the initial moment; use the trace of the rotation matrix to calculate the maximum values of the pitch angle and roll angle rotation angles ; For the attitude conversion duration between two target windows, calculate the attitude conversion duration every , where is the angular velocity of attitude conversion, and obtain the attitude conversion duration sequence within this time window , the start time and end time of satellite i for ground target j; Due to the differences in maneuvering conditions of remote sensing satellites, imaging task objectives or or do not have visible time windows, do not meet the maneuvering attitude constraints, and the moment after the attitude conversion ends is not within the time windows of both, etc., making it impossible to switch from task to , and a penalty is given at this time
[0047] Step 2.3, when a dynamic task is inserted, pause generating the task scheduling strategy and expand the action space. When d dynamic tasks are added .
[0048] Step 2.4, build an adaptive dynamic reward space. The dynamic reward function consists of the following three parts: task priority reward, task time reward, and task dynamic reward . The task priority reward represents the task priority. The task completion time reward represents a non-linear reward to prevent the reward for tasks completed too late from approaching zero, while ensuring that the reward for tasks completed earlier is higher: , and a positive reward can be obtained if the task is completed before the deadline, and the earlier it is completed, the higher the reward. The dynamic task reward is calculated based on time decay, and the calculation method is , represents whether the task is a dynamic task (take 1 for dynamic tasks and 0 for ordinary tasks). Exponential decay mechanism: The closer the task is to the deadline, the higher the reward, ensuring that urgent tasks are more likely to be selected
[0049] Step 2.5, building an adaptive dynamic action space. Define the maximum action space dimension , , covering all static tasks, defining certain null actions for dynamic tasks, and setting a reserved position for the number of dynamic tasks; at each time step t, according to the current satellite state and task constraint conditions (such as visibility window, attitude maneuverability, etc.), calculate the validity of each action . If an action is not executable in the current state (such as the target has been executed 、 the target is not visible or the attitude conversion time is insufficient), then mark it as an invalid action and select ; to generate a binary mask vector , where indicates that the action is valid, indicates that the action is invalid. Apply the mask vector to the output layer of the policy network, and mask the probability distribution of invalid actions through element-wise multiplication , where is the original output of the policy network ; When a dynamic task is added, update the mask vector to ensure the stability of the training and testing processes.
[0050] Step 2.6, in the task selection strategy, dynamic tasks are given priority. When a dynamic task is added, during the training process, select tasks according to the following logic: repeat steps 2.1 - 2.2.
[0051] Step 2.7, use the imaging task priority to construct the reward in the Markov process: each imaging task i has a priority , indicating the urgency and importance of the task. The reward r is determined by the current state and the action , describing the benefit obtained by taking the action in the current state. The action rewards of different remote sensing satellites have different calculation methods, and the total reward is the accumulation of the rewards of all remote sensing satellites . When the reward value is the largest and no longer changes, it converges to obtain the optimal policy.
[0052] Step 2.8, segment the time into scheduling periods according to the target demand refresh period. The status of the remote sensing satellite and the imaging task observation situation in each time slice are used as status information. Different rewards are assigned to the statuses feedback by different actions in the environment, and the remote sensing satellite observation sequence is obtained with the goal of maximizing the reward. After the remote sensing satellite completes the imaging task of a point target, a certain task switching time is required to observe the next target. The task switching time depends on the satellite attitude maneuverability, and the attitude conversion needs to be performed according to the attitude requirements of two imaging tasks. Read the attitude switching time calculated in Step 2. When the value is less than 0, the current target is regarded as... At the window cannot of the target perform attitude conversion and observe the target. When the value is greater than 0, the current target at the window can of the target perform attitude conversion and observe the target ; If the maneuvering conditions for attitude conversion are not met, the next target cannot be observed, and the next target that meets the attitude conversion constraints is reselected.
[0053] Step 2.9, transform the multi-target imaging task scheduling problem of the remote sensing satellite into an optimal strategy solving problem for the intelligent agent: The reward state transition probability function describes the probability that after the remote sensing satellite executes action a respectively under the current imaging task and the remote sensing satellite transfer state s, the imaging task and the remote sensing satellite transfer to the new state . The policy is the probability distribution of the remote sensing satellite taking action a in state s. is the stochastic policy; The neural network fits the policy function to obtain the probability distribution of action a in state s. The goal is to determine the action sequences of each remote sensing satellite and find the current optimal policy to maximize the expected report of the overall constellation. The objective function involved is defined as:
[0054] In the above formula, is the parameter of the policy , represents the expected value of the initial state , is the value function, which represents the expected total return executed according to the policy in the initial state ; Specifically, the value function is defined as starting from the state and following the policy in future time steps. The expected value of the cumulative discounted return that can be obtained from an action; Denotes the expectation of sampling under the policy ; is the discount factor. The closer the value is to 1, the greater the impact of future rewards on the current decision; the closer the value is to 0, the more attention is paid to recent rewards; Denotes that at time step t, the state and action ; Denotes the expectation of the value function under all possible initial states starting from the initial state and executed according to the policy ; Denotes that when executed according to the policy , an immediate reward is obtained at each time step t, and the future rewards are weighted and summed according to the discount factor ; when solving for the policy to maximize the objective function , find , thus obtaining the optimal action policy for the time series.
[0055] In step 3, a meta-learning layer is constructed using a recurrent neural network (LSTM) to transform the learning between tasks into a sequence modeling problem. Through the special structure of the LSTM, the satellite agent can record the information of historical tasks, so that it can refer to the training experience of the meta-learning layer when facing new tasks. The specific implementation of step 3 includes the following sub-steps: Step 3.1, constructing the meta-learning network architecture. Construct a meta-learning layer with a shared LSTM, and initialize the LSTM hidden state and the initial state of the environment .
[0056]
[0057] Step 3.2, constructing the reinforcement learning framework. A policy network (Actor) and a value network (Critic) are defined, with the input vector being , and the input includes the current state , the previous action , the previous reward and the termination flag , and the output hidden state ; the output layer of the policy network is a fully connected layer, generating an action probability distribution ; the output layer of the value network is a fully connected layer, generating a state value estimation scalar .
[0058] Step 3.3, constructing an input gate, including the gradient of the loss function at the current parameter value , the current loss function value , the parameter values of the previous iteration and the value of the input gate at the previous time step , through a weight matrix and a bias vector to perform a linear transformation, and through the non-linear mapping of the Sigmoid function, the value of the input gate is obtained . The value of the input gate is equivalent to the learning rate, which is dynamically adjusted according to this input information and determines the influence degree of new information on the state update of the meta-learner
[0059] Step 3.4, Meta-training task sampling and initialization. Assume there is a set of MDP sets M and the distribution on it. In each training, a specific MDP interaction of n episodes is sampled from it . The goal is to maximize the cumulative expected total discounted reward during the entire trial through an end-to-end optimization process, rather than the reward of a single episode. Using a recurrent neural network (LSTM), each time step accepts the output of the embedding function as input
[0060] Step 3.5, Meta-training task sampling and initialization. Sample N meta-training tasks from the task distribution , each task corresponding to different satellite orbit parameters, resource constraints or dynamic task combinations; for each meta-training task , perform K rounds of interactions, and the hidden state is passed between rounds
[0061] Step 3.6, Cross-round experience transfer and interaction. Perform K training rounds for each task. At each time step t within each round, the input of the LSTM is , and the output action is . According to the policy , select the action , and after execution, obtain the reward and the next state , until the round terminates. When the round terminates, retain the final hidden state of the LSTM , as the initial state of the next round
[0062] Step 3.7, Meta-learning layer unit state update, the candidate unit state is directly determined by the gradient of the loss function , that is . The unit state at the previous moment is equal to the parameter value updated after the previous iteration . According to the LSTM unit state update rule Combine the forget gate with the previous unit state and the input gate to obtain an updated unit state in the meta - learning layer .
[0063] Step 3.8, Parameter update and task interaction: Randomly sample T batches of data . For each batch of data, the base learner (including trainable parameters) calculates the loss function value and the loss function gradient value , and then inputs this information into the meta - learner (including trainable parameters ). The meta - learner, based on the input gate , forget gate and the updated unit state obtained from the above calculations, uses its own parameters to update the parameters of the base learner. After processing all batches of data for the t - th task, use the validation set of this task to calculate the loss function value and the loss function gradient value on the validation set, and update the parameters of the meta - learner based on this information.
[0064] Step 3.9, PPO optimization and meta - gradient update: Randomly select a task from the task set and generate the relevant task state to start training. In each round of training, the agent interacts with the environment according to the current policy, collects information such as states, actions, rewards, etc., and uses this information to update the policy network and value network. Collect trajectory data , calculate the generalized advantage estimate .
[0065] By clipping the objective function , the calculation method is , minimize the value network through the mean - squared error , use the Adam optimizer to back - propagate and update the parameters and , and retain the LSTM gradient chain transfer. Repeat the above steps to update the policy network parameters and the value network parameters , with the objective functions being the clipped policy loss and the mean - squared error of the value function, so that the policy is gradually optimized until the predetermined number of training rounds or the algorithm converges.
[0066] The specific implementation of Step 4 includes the following sub-steps: Step 4.1, model loading and initialization. Initialize the initial state of the system, including the initial state information of all remote sensing satellites and targets; for different imaging task sets, load the saved model. During the training process, the reinforcement learning model learns a policy , which selects the optimal action a according to the current state s. The policy is represented by a neural network, and its output is the probability distribution of taking different actions in each state.
[0067] Step 4.2, real-time task decision-making process. Start testing. In each training time step, the agent interacts with the environment to collect trajectory data during the training process , and each trajectory consists of state, action, reward, and next state; advance the time step t according to the scheduling period minutes to obtain the current satellite state ; input and to the policy network, and output the action probability distribution ; at time step t, according to the current state use the policy to select the task with the highest probability in the discrete action space of the satellite , and mask the inexecutable actions through constraint judgment.
[0068] Step 4.3, execute the executable actions after constraint judgment (such as photographing the target), consume resources and update the task queue, and the environment returns the next state and reward ; input to the LSTM meta-learning layer to update the state and generate ; when the episode terminates ( ), reset the LSTM hidden state . After executing the action , update the system state , and the update includes reducing the remaining visible time window of the corresponding target and updating the state of the remote sensing satellite.
[0069] Step 4.4, when a dynamic task is inserted, pause the current test process, update the action space, and recalculate the policy and reward function based on the updated action space; record the executed tasks at each time step, including time, remote sensing satellite, target, and imaging task execution results.
[0070] Step 4.5, repeat Steps 4.1 to 4.4 until the visible time windows of all targets end or the predetermined imaging task time range is reached. After completion, obtain the remote sensing satellite for each point target imaging task, the start observation time, and the end observation time, and obtain the imaging task planning scheme of the remote sensing satellite for different point targets.
[0071] For the convenience of implementation reference, an example is presented. The present invention uses 1 remote sensing satellite and 1000 static ground target tasks and 100 dynamic ground targets for algorithm training, and another task set of the same scale is selected for testing to illustrate the specific implementation of the present invention. The experiment is developed based on the Python environment, with the learning rate set to 0.1, the discount factor to 0.99, the clipping coefficient to 0.2, the entropy coefficient to 0.01, the value function coefficient to 0.1, the number of environment interaction steps to 2048, and the training batch size to 100000. The orbital parameters of the remote sensing satellite are shown in Table 1, and the geographical location visible windows and priorities of the point targets are shown in Table 2.
[0072] Table 1 Orbital parameters of the satellite
[0073] Table 2 Task target information
[0074] After modeling the dynamic task scenario of the satellite through Steps 1, 2, 3, and 4, the experiment adopts and solves it through the method of the present invention. The final satellite task scheduling scheme is shown in Table 3, and the target distribution is as Figure 3 shown.
[0075] Table 3 Satellite task scheduling scheme
[0076] The implementation basis of each embodiment of the present invention is achieved through programmed processing by a device with processor functions. Therefore, in engineering practice, the technical solutions and functions of each embodiment of the present invention are encapsulated into various modules. Based on this actual situation, on the basis of the above embodiments, an embodiment of the present invention provides a meta-reinforcement learning scheduling device for earth observation satellite task planning, and this device is used to execute a meta-reinforcement learning scheduling method for earth observation satellite task planning in the above method embodiments.
[0077] The device includes: a first main module for obtaining the visible time window of the satellite for a ground target, including the start time and end time of observation; a second main module for inputting task data and initializing the environment, reading the policy of the trained reinforcement learning model at each time step and inputting the start time and end time of observation of the feasible ground target, and outputting an imaging task planning policy; wherein, the training of the reinforcement learning model includes: modeling the satellite imaging task execution process as a dynamic Markov decision process with adaptive rewards, defining the state space, action space, rewards, state transition probabilities, and discount factor, where the action is represented as selecting an observable task; after the dynamic task is added, the dimension of the action space is expanded in real time; initializing the constructed Markov decision environment and initializing the PPO algorithm, and establishing a policy network and a value network; constructing a deep reinforcement learning algorithm based on meta-learning, introducing memory for the agent using a recurrent neural network, transforming the learning between tasks into a sequence modeling problem, introducing a meta-learning layer to train multiple tasks and learning how to quickly adapt to new tasks; updating the parameters of the policy network and the value network according to the objective function of the PPO algorithm, minimizing the mean square error of the value function estimation, continuously iterating the training, and obtaining the optimal scheduling policy in the Markov decision environment when the training converges; setting different learning rates, discount factors, clipping coefficients, value function coefficients, and entropy coefficients, and at the same time setting different training batch sizes and training epochs, starting the training of the deep reinforcement learning algorithm based on meta-learning, and storing the training results of the algorithm when the training converges.
[0078] The meta-reinforcement learning scheduling device for earth observation satellite task planning provided by the embodiment of the present invention aims at the problem that the traditional satellite dynamic task planning method has obvious insufficient processing ability in the face of large-scale dynamic tasks and complex constraints. By using the foregoing several modules, through introducing a task dynamic reward mechanism and using the meta-learning layer to capture task commonalities to achieve cross-scenario policy transfer, the efficiency of solving the satellite task planning policy is improved.
[0079] It should be noted that the device embodiment provided by the present invention, in addition to being used to implement the method in the above method embodiment, is also used to implement the methods in other method embodiments provided by the present invention. The difference is only in setting the corresponding functional modules. Its principle is basically the same as the principle of the above device embodiment provided by the present invention. As long as those skilled in the art, on the basis of the above device embodiment, refer to the specific technical solutions in other method embodiments, obtain the corresponding technical means by combining technical features, and the technical solutions composed of these technical means, and on the premise of ensuring the practicability of the technical solutions, improve the device in the above device embodiment to obtain the corresponding device type embodiment for implementing the methods in other method type embodiments.
[0080] On the other hand, the embodiment of the present invention also provides a meta-reinforcement learning scheduling device for earth observation satellite task planning, including: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement a meta-reinforcement learning scheduling method for earth observation satellite mission planning as described in the above technical solution.
[0081] On the other hand, an embodiment of the present invention further provides a non-transitory computer-readable storage medium storing computer instructions that cause the computer to execute the method.
[0082] In summary of the above embodiments, the present invention discloses a meta-reinforcement learning scheduling method for earth observation satellite mission planning. By predicting the target visible window for all visible time windows of the task objective, the method preprocesses the satellite visible window, designs a dynamic reward mechanism and constructs an adaptive dynamic reward task Markov decision environment, then constructs a reinforcement learning algorithm based on meta-learning, and finally loads the algorithm to input test data and outputs an imaging task planning strategy. By introducing a task dynamic reward mechanism and using the meta-learning layer to capture task commonalities, after multiple rounds of training, the intelligent agent can quickly adapt and make efficient decisions in a changing environment to obtain the satellite task planning result. The present invention can be effectively applied to the mission planning of remote sensing satellites with dynamic task scenarios.
[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the technical solutions of the embodiments of the present invention.
Claims
1. A meta-reinforcement learning scheduling method for earth observation satellite mission planning, characterized in that, Including: Obtain the visible time window of the satellite for ground targets, including the start time and end time of observation; Input task data and initialize the environment. At each time step, read the policy of the trained reinforcement learning model and input the start time and end time of observation of feasible ground targets, and output the imaging task planning policy. Among them, the training of the reinforcement learning model includes: Model the satellite imaging task execution process as a dynamic Markov decision process with adaptive rewards, define the state space, action space, rewards, state transition probabilities, and discount factor, where the action is represented as selecting observable tasks; after the dynamic task is added, the dimension of the action space is expanded in real time; initialize the built Markov decision environment and initialize the PPO algorithm, and establish a policy network and a value network; Build a deep reinforcement learning algorithm based on meta-learning, introduce memory for the agent using a recurrent neural network, transform the learning between tasks into a sequence modeling problem, and introduce a meta-learning layer to train multiple tasks and learn how to quickly adapt to new tasks; Update the policy network parameters and value network according to the objective function of the PPO algorithm, minimize the mean square error of the value function estimation, continuously iterate the training, and obtain the optimal scheduling policy in the Markov decision environment when the training converges; Set different learning rates, discount factors, clipping coefficients, value function coefficients, entropy coefficients, and at the same time set different training batch sizes and training epochs, start the training of the deep reinforcement learning algorithm based on meta-learning, and store the algorithm training results when the training converges.
2. The meta-reinforcement learning scheduling method for the mission planning of an earth observation satellite according to claim 1, wherein Building a Markov decision environment also includes: Build an adaptive dynamic reward space, and the dynamic reward function includes task priority rewards, task time rewards, and task dynamic rewards; Build an adaptive dynamic action space, define the maximum action space dimension, cover all static tasks and define empty actions for dynamic tasks, set the reserved position for the number of dynamic tasks, and at each time step, calculate the effectiveness of each action according to the current satellite state and task constraint conditions.
3. The meta-reinforcement learning scheduling method for the mission planning of an earth observation satellite according to claim 2, wherein, Building a Markov decision environment also includes: In the task selection strategy, give priority to dynamic tasks; Use the imaging task priority to construct the rewards in the Markov process; Segment the time into scheduling periods according to the target demand refresh period. The satellite state and imaging task observation situation of each time slice are used as state information, and different rewards are given to the states feedback by different actions in the environment, and the satellite observation sequence is obtained with the maximization of rewards as the goal; Transform the multi-target imaging task scheduling problem of the earth observation satellite into an optimal policy solving problem of the agent, and obtain the optimal action policy in time series.
4. The meta-reinforcement learning scheduling method for the mission planning of an earth observation satellite according to claim 1, characterized in that, Building a deep reinforcement learning algorithm based on meta-learning also includes: Build a meta-learning network architecture and a reinforcement learning framework, and construct an input gate; Meta-training task sampling and initialization, sample several meta-training tasks from the task distribution, and each task corresponds to different satellite orbit parameters, resource constraints, or dynamic task combinations; for each meta-training task, perform multiple rounds of interaction, and the hidden state is passed between rounds, and the final hidden state is retained at the end of the round as the initial state of the next round; According to the recursive neural network unit state update rule, the forget gate, the unit state at the previous moment, the input gate and the candidate unit state are combined to obtain the updated unit state in the meta-learning layer; The meta-learner updates the parameters of the base learner using its own parameters according to the calculated input gate, forget gate, and updated unit state.
5. The meta-reinforcement learning scheduling method for the mission planning of an earth observation satellite according to claim 4, characterized in that, Building a deep reinforcement learning algorithm based on meta-learning also includes: Randomly select tasks from the task set and generate related task states to start training. In each round of training, the agent interacts with the environment according to the current strategy, collects state, action, and reward information, and uses this information to update the policy network parameters and value network parameters using the PPO algorithm. The objective function is the clipping policy loss and the mean square error of the value function, so that the strategy is gradually optimized until the predetermined number of training rounds is reached or the algorithm converges.
6. The meta-reinforcement learning scheduling method for the mission planning of an earth observation satellite according to claim 1, wherein The training of the reinforcement learning model further includes: Step A: Initialize the initial state information of all satellites and targets; for different imaging task sets, load the saved models; Step B: In each training time step, the agent interacts with the environment and collects trajectory data during the training process. Each trajectory consists of state, action, reward and next state. The time step is advanced according to the scheduling cycle to obtain the current satellite state. The current satellite state and hidden state are input to the policy network, and the action probability distribution is output. According to the current state, the strategy is used to select the task with the highest probability in the discrete action space of the satellite, and the inaction action is blocked through constraint judgment. Step C, execute the action after executing the constraint judgment, consume resources and update the task queue, and the environment returns the next state and reward; Step D: When a dynamic task is inserted, the current training process is paused, the action space is expanded, the policy network output layer dynamically increases the corresponding action dimension, and the policy and reward function are recalculated based on the updated action space; the tasks executed are recorded at each time step, including time, remote sensing satellite, target, and imaging task execution results; Step E, repeat steps A to D until the visible time window of all targets ends or reaches the predetermined imaging task time range. After the end, the remote sensing satellite executing the imaging task of each point target, the start observation time, and the end observation time are obtained to form an imaging task planning scheme for different point targets by the remote sensing satellite.
7. A meta-reinforcement learning scheduling device for earth observation satellite mission planning, characterized in that, include: The first main module is used to obtain the satellite's visible time window for ground targets, including the observation start time and end time; The second main module is used to input task data and initialize the environment, read the strategy of the trained reinforcement learning model at each time step, input the observation start time and end time of feasible ground targets, and output the imaging task planning strategy; wherein the training of the reinforcement learning model includes: The satellite imaging task execution process is modeled as a dynamic Markov decision process with adaptive rewards. The state space, action space, reward, state transition probability and discount factor are defined, where the action is represented by the selection of observable tasks. After the dynamic task is added, the action space dimension is expanded in real time. The Markov decision environment is initialized, and the PPO algorithm is initialized to establish the policy network and value network. Build a deep reinforcement learning algorithm based on meta-learning, introduce memory for the agent using a recurrent neural network, transform the learning between tasks into a sequence modeling problem, introduce a meta-learning layer to train multiple tasks and learn how to quickly adapt to new tasks; Update the policy network parameters and value network according to the objective function of the PPO algorithm, minimize the mean square error of the value function estimation, continuously iterate the training, and obtain the optimal scheduling policy in the Markov decision environment when the training converges; Set different learning rates, discount factors, clipping coefficients, value function coefficients, entropy coefficients, and at the same time set different training batch sizes and training epochs, start the training of the deep reinforcement learning algorithm based on meta-learning, and store the algorithm training results when the training converges.
8. A meta-reinforcement learning scheduling device for earth observation satellite mission planning, characterized in that Comprising one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium stores computer instructions that cause the computer to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Imaging satellite autonomous task planning method based on machine learning
CN109767128A
GEO on-orbit service task planning method based on meta reinforcement learning
CN117382920A
Deep reinforcement learning scheduling method and device for satellite multi-point target imaging
CN118709748A
Method and apparatus for generating multi-drone network cooperative operation plan based on reinforcement learning
US20230297859A1
Efficient adaption of robot control policy for new task using META-learning based on META-imitation learning and META-reinforcement learning
WO2020154542A1
Cited By
Deep reinforcement learning-driven drilling parameter intelligent real-time optimization method
CN120597735A
A Deep Reinforcement Learning-Driven Intelligent Real-Time Optimization Method for Drilling Parameters
CN120597735B
Satellite orbit control method based on continuous time near-end strategy optimization reinforcement learning algorithm
CN120722768A
Planning method for agile observation satellite
CN120746211A
Scheduling method and system for improving execution efficiency based on time and task state
CN120762862A