SAR imaging satellite task planning method based on deep reinforcement learning
Through a method based on deep reinforcement learning, the observation characteristics of Earth observation satellites are analyzed and mathematical models are established, which solves the problems of satellite mission scheduling efficiency and quality, and achieves the generation of efficient and high-quality task planning schemes.
Patent Information
- Application Number
- CN202510196385.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-17
AI Technical Summary
The existing technology is difficult to effectively solve the problem of scheduling of Earth observation satellite missions, especially in the case of increasing the number of tasks and increasing the demand for timely observation, it is difficult to ensure observation efficiency and meet mission requirements.
Using a method based on deep reinforcement learning, we analyze the observation characteristics of synthetic aperture radar earth observation satellites, establish corresponding mathematical models, model the problem solving process into Markov decision-making process, and use strategic models such as pointer networks and attention models to learn effective decision-making strategies.
It has achieved rapid generation of high-quality task planning schemes, improved the efficiency and quality of satellite mission scheduling, met the demand for timely observation, and has application prospects in future satellite scheduling optimization.
Smart Images

Figure CN120163365A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a SAR imaging satellite mission planning method based on deep reinforcement learning, belonging to the field of satellite control technology. Background Art
[0002] Synthetic aperture radar Earth observation satellites (SEOS) are designed to capture images of the Earth's surface from orbit using onboard microwave sensors. Due to the imaging characteristics of synthetic aperture radar (SAR), SEOS can achieve all-weather, all-time, and high-resolution observations of the Earth's surface, unaffected by weather conditions and lighting conditions. This ability makes SEOS of extremely important value and significance in fields such as environmental monitoring, urban planning, agricultural management, and disaster response. The SAR Earth observation satellite scheduling problem (SEOSSP) involves determining the optimal sequence of observation tasks according to task requirements and satellite characteristics, including the start and end times of each task. With the increase in the number of tasks and the demand for timely observations, developing efficient and high-quality methods is crucial for ensuring observation efficiency and meeting task requirements. The present invention first analyzes the observation characteristics of SEOSs and establishes a mathematical model of SEOSSP. Based on this model, we model the problem-solving process as a Markov decision process and apply two basic policy models (BPMs) to learn effective decision-making strategies. Summary of the Invention
[0003] The object of the present invention is to propose a SAR imaging satellite mission planning method based on deep reinforcement learning for the above technical problems.
[0004] To achieve the above object, the technical solution adopted by the present invention to solve the above technical problems is:
[0005] The ground station obtains the mission planning problem and performs planning preprocessing;
[0006] In the mission planning problem, the operating parameters of the satellite are preprocessed into time windows and observation modes before solving the mission planning.
[0007] During the observation process of the Synthetic Aperture Radar Earth Observation Satellite SEOS, the time period during which the radar can detect the target is called the Visible Time Window VTW of the mission. According to the direction of the radar antenna towards the target, SEOS can operate in different observation modes: strip mode, scan mode, and spotlight mode;
[0008] Each of the above modes corresponds to a different observation method, and each mission has a corresponding observation duration. There is a conversion time required when switching from one mission of a different mode to another. When solving the Earth Observation Satellite Mission Scheduling Problem SEOSSP, the conversion time between different modes needs to be considered;
[0009] Each mission needs to meet the requirements of the observation mode and duration. The mission planning needs to determine the specific observation sequence to satisfy the constraints of the observation mode conversion time and duration; each SEOS mission has an observation benefit attribute, which reflects the importance of the mission among the missions to be observed. The more important the mission, the higher its benefit value; when the mission is completely observed, the corresponding observation benefit can be obtained;
[0010] SEOS mission planning is to generate a high-quality observation sequence that meets the problem constraints by establishing a mathematical model, and the quality of the mission planning is measured by calculating the total observation benefit of the observation sequence.
[0011] Furthermore, before establishing the mathematical model, the following assumptions are made for the problem to simplify the constraints:
[0012] 1) Each mission is observed at most once;
[0013] 2) The observation tasks are limited to point targets, that is, one observation can completely cover the target to be observed;
[0014] 3) Only consider the mission planning in the observation phase, and do not involve the planning problems of other activities including satellite on-orbit charging and data transmission;
[0015] 4) Ignore the satellite memory and battery constraints; assume that the power supply can be guaranteed in time during the satellite operation, and the memory is released by returning the observation information, so as to ensure that the memory and power meet the working requirements;
[0016] 5) Do not consider the physical properties of different observation modes. The observation methods of specific tasks have been designed in the task preprocessing stage, and only the specific observation time window needs to be determined in the planning stage.
[0017] Furthermore, establishing the mathematical model specifically includes:
[0018] The goal of SEOSSP is to maximize the total observation benefit of the observation sequence, which can be expressed as:
[0019] max P=pro i·x i (1)
[0020] The problem has the following constraints:
[0021] (1) The actual observation time of a task should be within the visible time window of the task:
[0022]
[0023] (2) The actual observation duration of a task meets the requirements of the task's continuous observation time:
[0024]
[0025] (3) Each task can be observed at most once:
[0026]
[0027] (4) Calculation of the transition time between two tasks:
[0028]
[0029] (5) Calculation of the SEOS transition time:
[0030] SEOS can observe tasks continuously in the same observation mode. In this case, the transition time between tasks can be ignored; if the observation modes of the current observed task and the next two observed tasks are different, SEOS needs to switch modes; therefore, for SEOSs, the transition time mainly considers the mode transition time between tasks with different observation modes;
[0031]
[0032] Define the observation mode of SEOS for tasks as follows:
[0033]
[0034] Among them, T is the set of candidate tasks, task i is the i-th task to be observed, pro i is the benefit obtained from completing the i-th task, d i is the observation duration of task i vtw i is the visible time window of task i i is the start time of the visible time window of task i i is the end time of the visible time window of task i i is the actual start observation time of task i i For task i The actual end observation time of Trans ij is the transition time between task i and task j, and mode i is the observation mode of task i, and x i is a binary variable indicating whether task i is observed, and P is the total observation benefit of the observation sequence.
[0035] The present invention discloses a SAR imaging satellite mission planning method based on deep reinforcement learning. By deeply analyzing SEOSSP, a corresponding mathematical model is established, and a solution is proposed using deep reinforcement learning DRL. This method first analyzes the observation characteristics of SEOS and establishes a mathematical model of SEOSSP. Based on this model, the problem-solving process is modeled as a Markov decision process, and two basic policy models BPMs are applied to learn effective decision-making strategies, which can quickly generate high-quality solutions, highlighting its application prospects in future SEOS scheduling optimization. Brief Description of the Drawings
[0036] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0037] Figure 1 is the SAR satellite observation mode of the present invention;
[0038] Figure 2 is the training process of the present invention.
[0039] Figure 3 is the BPM solution process of the present invention. Detailed Embodiments
[0040] The following describes exemplary embodiments of the present application in conjunction with the drawings, including various details of the embodiments of the present application to assist understanding, which should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for clarity and conciseness, the description below omits the description of well-known functions and structures.
[0041] Synthetic aperture radar Earth observation satellites (SEOS) are designed to capture images of the Earth's surface from orbit using onboard microwave sensors. Due to the imaging characteristics of synthetic aperture radar (SAR), SEOS can achieve all-weather, all-time, and high-resolution observations of the Earth's surface, unaffected by weather conditions and lighting conditions. This ability makes SEOS of extremely important value and significance in fields such as environmental monitoring, urban planning, agricultural management, and disaster response. The SAR Earth observation satellite scheduling problem (SEOSSP) involves determining the optimal sequence of observation tasks according to task requirements and satellite characteristics, including the start and end times of each task. With the increase in the number of tasks and the demand for timely observations, developing efficient and high-quality methods is crucial for ensuring observation efficiency and meeting task requirements. This invention first analyzes the observation characteristics of SEOSs and establishes a mathematical model of SEOSSP. Based on this model, we model the problem-solving process as a Markov decision process and apply two baseline policy models (BPMs) to learn effective decision-making strategies.
[0042] Problem description
[0043] SAR satellite imaging characteristics
[0044] The ground station obtains the task planning problem and performs preprocessing for planning;
[0045] In the task planning problem, the operating parameters of the satellite are preprocessed into time windows and observation modes before solving the task planning.
[0046] During the SEOS observation process, the time period when the radar can detect the target is called the visible time window (VTW) of the task. According to the direction of the radar antenna towards the target, SEOS can work in different modes: stripmap mode (as shown in Figure 1 (a)), scan mode (as shown in Figure 1 (b)), and spotlight mode (as shown in Figure 1(as shown in (c)). Each mode has a different observation method, and each task has a specific observation duration. Although SEOS can continuously execute observation tasks in the same observation mode, a conversion time is required when switching from one task that requires a different observation mode to another. Therefore, when solving SEOSSP, the importance of considering the conversion time between different observation modes is self-evident.
[0047] Problem assumptions and variables
[0048] Before establishing the mathematical model, the following assumptions are made for the problem to simplify the constraints:
[0049] 1) Each task is observed at most once.
[0050] 2) The observation tasks are limited to point targets, that is, one observation can completely cover the target to be observed.
[0051] 3) Only the task planning in the observation phase is considered, and the planning problems of other activities such as satellite on-orbit charging and data transmission are not involved.
[0052] 4) The satellite memory and battery constraints are ignored. This study assumes that the power supply can be guaranteed in a timely manner during the operation of the satellite, and the memory is released by returning the observation information, so as to ensure that the memory and power meet the working requirements.
[0053] The problem variables are shown in Table 1.
[0054] Table 1 Variables
[0055]
[0056]
[0057] Mathematical model
[0058] The goal of SEOSSP is to maximize the total observation benefit of the observation sequence, which can be expressed as:
[0059] max P = pro i ·x i (1)
[0060] The problem has the following constraints:
[0061] (1) The actual observation time of the task should be within the visible time window of the task:
[0062]
[0063] (2) The actual observation duration of the task meets the requirements of the task's continuous observation time:
[0064]
[0065] (3) Each task is observed at most once:
[0066]
[0067] (4) Calculation of the transition time between two tasks:
[0068]
[0069] (5) Calculation of the SEOS transition time:
[0070] SEOS can continuously observe tasks in the same observation mode. In this case, the transition time between tasks can be ignored. If the observation modes of the current observed task and the next two observed tasks are different, SEOS needs to switch modes. Therefore, for SEOSs, the transition time mainly considers the mode transition time between tasks with different observation modes.
[0071]
[0072] The present invention defines the observation mode of SEOS for tasks as follows:
[0073]
[0074] Problem Solving
[0075] SAR Satellite Mission Planning Based on Markov Decision Process
[0076] The problem uses a constructive solution process to generate a planning scheme. Initially, the construction process starts from an empty set and selects tasks according to a specific strategy (the baseline strategy model used in this article will be introduced later). At the decision point, the model selects a task according to the current state and inserts it into the observation sequence. Subsequently, the state is updated according to the selected task and a constraint check is performed. Repeat the above operations until no new tasks can be added and the solution construction is completed.
[0077] During the solution construction process, task selection is completely determined by the previous decision. Therefore, the problem solving can be expressed as a Markov decision process consisting of four basic components: state set, action set, probability transition function, and reward function. The specific situation is as follows:
[0078] · The state set consists of the satellite and task states at the decision point.
[0079] · The action set represents the action of task selection.
[0080] · The probability transition function represents the probability distribution obtained by applying an action to a state.
[0081] · The reward function represents the total priority of the observation sequence.
[0082] Baseline Policy Model
[0083] The Markov decision model is used to model the problem. During this process, the decision maker sequentially selects tasks to join the observation sequence according to a certain policy. To better solve this problem, two basic policy models (baseline policy model, BPM) are used to fit the policy for solving SEOSSP, and the model is trained through a reinforcement learning algorithm. The BPM solution process is specifically as Figure 3 shown.
[0084] During the solution process, at each decision moment t, the problem scenario features are used as the BPM input, and the BPM outputs the action probability (i.e., the task probability distribution). Based on this, a task is selected to join the current solution sequence. Then it is updated to decision moment t + 1, and the decision continues until the termination condition is reached, that is, no task can be added to the observation sequence. During this process, a Mask mechanism is introduced, which performs task constraint checks according to the current moment state and ensures that tasks violating the problem constraints cannot be selected (probability is 0).
[0085] The specific model introduction is as follows:
[0086] 1) Pointer network (PN)
[0087] PN is a neural network architecture that is good at sequence-to-sequence tasks. A remarkable feature of PN is that it can dynamically generate pointers in the output sequence, and these pointers refer to specific positions in the input sequence. This ability allows the network to directly address elements in the input sequence, enhancing its ability to manage dependencies between segment information during the generation process. Therefore, PN can improve the quality of the generated output and the generalization performance of the model. Chen et al. first proposed an improved pointer network (IPN) to solve the EOSSP of agile EOS, and we modified it to be able to solve SEOSSP. By extracting the problem features of EOSSP as the model input. According to the feature state, the problem features can be divided into static features and dynamic features. Among them, the static features refer to the features that do not change during the decision-making process, including the start time of the visible time window of the task the end time of the visible time window, the observation duration d i and the task benefit value pro i . The dynamic features refer to the states that need to be updated after each decision. Among them, the task end time oe at the previous decision moment t-1 represents the current decision time, and task i the state at decision point t Used to indicate whether the task has been selected in this state. Through the input of problem features, IPN can output the probability distribution of tasks during the decision-making process, guiding task selection and thus achieving problem-solving.
[0088] 2) Attention Model (AM)
[0089] AM has proven to be significantly effective in solving classical problems. This model allows dynamic allocation of attention weights to different parts of sequential data. By calculating the similarity between queries and keys and normalizing the results, the attention mechanism enables the model to more effectively capture key information within the input sequence, thereby improving performance in problem-solving tasks. In the present invention, we use the features of SEOSSP as the input of the model, utilize the attention mechanism to solve this problem, and use the static features of the problem as the feature input of the model (the specific meaning is the same as IPN) so that it can output the probability distribution of tasks during the decision-making process, and thus achieve problem-solving.
[0090] Reinforcement learning training method
[0091] We use the REINFORCE algorithm with a rolling baseline to obtain the optimal parameter θ * , and the pseudocode of the training method is shown in Algorithm 1.
[0092]
[0093] Use the REINFORCE algorithm with a rolling baseline to train IPN and AM respectively, so that the two models can learn the optimal solution parameters through training. During the specific training process, the model first initializes its parameter θ and sets the optimal parameter θ * = θ (as in Line3). Secondly, the model uses the parameter θ to generate a probability vector of actions at each decision step, selects actions according to the probability, and finally obtains the solution and calculates its reward. The baseline of the algorithm is the reward obtained by the model using the historical best parameter θ * . Finally, use the Adam optimizer to calculate and update the policy gradient of the model parameters. When the updated parameter θ performs better than the historical optimal parameter θ * on the validation dataset, update θ * . By continuously updating the parameters, the solution quality of the model is also improved, and finally the optimal model parameter θ * for solving is learned. The obtained optimal solution parameter is the parameter of the baseline policy model, and the task of the mathematical model is to perform constraint checking to avoid selecting tasks that do not meet the constraints.
[0094] Method verification
[0095] Experimental design
[0096] The scenario generation of SEOSSP follows the observation characteristics of SEOS, and the task duration and benefits are randomly obtained within the ranges of [15, 30] and [1, 10]. Due to the limitation of device memory, the model is trained using three task sizes (25, 50, and 100), and each size contains 256,000 scenarios. To evaluate the model performance, validation and test scenario sets with task sizes of 25, 50, 100, 150, and 200 are created, and each scenario set contains 1000 scenarios of the corresponding size. We use Actor-Critic and REINFORCEMENT with a rolling baseline to train AM and IPN respectively. To optimize the balance between training time and efficiency, we trained the model with a batch size of 512 in 20 epochs. The initial learning rate of the Adam optimizer is 0.0001, and the decay rate for each epoch is 0.99.
[0097] Analysis of the training process
[0098] To analyze the training method and its effects, we record the objective function values of each epoch in scenarios of different scales and plot the results in Figure 2 .
[0099] As Figure 2 shown, AM obtains better objective function values than IPN in both training methods and achieves convergence in scenarios of three sizes. This indicates that compared with IPN, AM has better policy representation and higher learning efficiency for this problem.
[0100] Analysis of the solution performance
[0101] To evaluate the effectiveness of the proposed method, we apply the trained model to 1000 different test scenarios to measure and record the average objective function value (Obj) and the calculation time (T) for each scenario.
[0102] Table 2 Experimental results
[0103]
[0104] According to the experimental results in Table 2, the AM solution performance obtained in different training methods is better than that of IPN, indicating that the model structure of AM is more suitable for representing the solution strategy of the problem. Specifically, whether using REINFORCEMENT training with a rolling baseline or Actor-Critic training, the solution quality of AM is similar, with the difference remaining within 0.15%. This shows that the performance of AM is less sensitive to the training method. In contrast, there are significant differences in the performance of IPN trained by the two methods, highlighting that the training method has a greater impact on the effectiveness of IPN. Generally speaking, AM trained with the REINFORCEMENT algorithm with a rolling baseline shows superior performance in solving problems, demonstrating high effectiveness and efficiency, thus proving that it can meet the timeliness and solution quality in real-world applications.
[0105] The present invention proposes a method based on deep reinforcement learning to solve the SEOSSP problem. First, the observation characteristics of SEOSSP are analyzed, and the mathematical model of SEOSSP is introduced. On this basis, two baseline policy models are improved, including the pointer network IPN and the attention model AM, to represent the solution strategy. Through experimental analysis, the performance of these enhanced policy models and related training methods is evaluated. The results show that AM trained by the REINFORCEMENT algorithm with a rolling baseline is significantly better than other comparison methods, proving its effectiveness in solving SEOSSP. This progress emphasizes the practical value and potential applicability of the proposed method in real-world scenarios.
[0106] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A SAR imaging satellite mission planning method based on deep reinforcement learning, characterized in that: The ground station obtains the mission planning problem and performs planning preprocessing; In the mission planning problem, the satellite’s operating parameters are preprocessed and converted into time windows and observation patterns before solving the mission planning; During the observation process of the synthetic aperture radar earth observation satellite SEOS, the time period during which the radar can detect the target is called the visible time window VTW of the mission. Depending on the direction of the radar antenna toward the target, SEOS works in different observation modes: strip mode, scanning mode and beam mode; Each of the above observation modes corresponds to a different observation method, and each task has a corresponding observation duration. When switching from a task in a different mode to another task, conversion time is required. When solving the Earth Observation Satellite Task Scheduling Problem SEOSSP, the conversion time between different modes needs to be considered; Each task needs to meet the requirements of observation mode and duration. Task planning needs to determine the specific observation sequence to meet the observation mode conversion time and duration constraints. Each SEOS task has an observation benefit attribute, which reflects the importance of the task in the tasks to be observed. The more important the task, the higher its benefit value. When the task is fully observed, the corresponding observation benefit can be obtained. SEOS mission planning generates high-quality observation sequences that meet problem constraints by establishing a mathematical model, and measures the quality of mission planning by calculating the total observation benefit of the observation sequence.
2. The planning method according to claim 1, characterized in that: Before building the mathematical model, the following assumptions are made to simplify the constraints: 1) Each task is observed at most once; 2) The observation task is limited to point targets, and the target to be observed can be fully covered by one observation; 3) Only the mission planning during the observation phase is considered, and the planning of other activities including satellite on-orbit charging and data transmission is not involved; 4) Ignore satellite memory and battery constraints; assume that the power supply can be guaranteed in time during the satellite operation, and release the memory by returning observation information, so as to ensure that the memory and power meet the working requirements; 5) Regardless of the physical properties of different observation modes, the observation method of a specific task has been designed in the task preprocessing stage, and only the specific observation time window needs to be determined in the planning stage.
3. The planning method according to claim 1, characterized in that: The establishment of mathematical model specifically includes: The goal of SEOSSP is to maximize the total observation benefit of the observation sequence, which can be expressed as: max P=pro i ·x i (1) The problem has the following constraints: (1) The actual observation time of the task should be within the visible time window of the task: (2) The actual observation time of the mission meets the mission continuous observation time requirements: (3) Each task is observed at most once: (4) Calculation of the two-task conversion time: (5) SEOS conversion time calculation: SEOS can observe tasks continuously with the same observation mode, and the conversion time between tasks can be ignored. If the observation modes of the current observation task and the next two observation tasks are different, SEOS needs to switch modes. For SEOS, the conversion time mainly considers the mode conversion time between tasks with different observation modes. The observation mode of SEOS for tasks is defined as follows: Among them, T is the candidate task set, task i is the i-th task to be observed, pro i is the benefit obtained by completing the i-th task, d i For task i The observation duration, vtw i For task i The visible time window, For task i The visible time window starts at For task i The visible time window ends at For task i The actual start time of observation is For task i The actual end observation time, Trans ij is the transition time between task i and task j, mode i is the observation mode of task i, x i is a binary variable, indicating whether task i is observed, and P is the total observed benefit of the observation sequence.
Citation Information
Cited By
Heterogeneous constellation intelligent task decision-making method for space transaction target observation
CN120974902A
Two-dimensional large field of view SAR satellite autonomous mission planning method and system
CN122694010A