Imaging satellite multi-circle task planning method based on deep reinforcement learning

Through the multi-turn neural strategy network and single-turn neural strategy network based on deep reinforcement learning, the complex correspondence between the task and the visible time window in the multi-turn imaging satellite mission planning is solved, and efficient task planning and observation efficiency are achieved.

CN120124459APending Publication Date: 2025-06-10NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510196002.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The multi-circle imaging satellite mission planning problem has a complex correspondence between the task and the visible time window, more complex constraints and greater scene complexity, which makes it difficult for existing algorithms to effectively plan observation tasks and improve satellite observation efficiency.

Method used

Using a deep reinforcement learning method, by constructing a multi-turn neural strategy network (MNPN) and a single-turn neural strategy network (SNPN), using Markov decision-making process and loop splitting strategies, we directly learn the solution strategies of multi-turn task planning to ensure that the task planning results meet the unique constraints.

Benefits of technology

Through the deep reinforcement learning method, effective solutions to multi-circle imaging satellite mission planning are achieved, satellite observation efficiency is improved, and the unique constraints of mission planning are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124459A_ABST
    Figure CN120124459A_ABST
Patent Text Reader

Abstract

The invention discloses an imaging satellite multi-circle task planning method based on deep reinforcement learning, and analyzes a multi-circle task planning problem. And then designing two multi-circle sub-task planning methods with different solving processes, namely a multi-circle sub-task planning method based on a Markov decision process and a multi-circle sub-task planning method based on circle splitting. For a method based on a Markov decision process, Markov decision process modeling and feature engineering of a problem are carried out; when a multi-turn task planning method based on turn splitting is designed, a multi-turn strategy model training method based on transfer learning is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multi-orbit mission planning method for imaging satellites based on deep reinforcement learning, belonging to the field of satellite control technology. Background Technique

[0002] The mission planning problem of agile imaging satellites (Agile Earth Observation Satellite Scheduling Problem, AEOSSP) involves the planning and scheduling of observation tasks for agile imaging satellites. In AEOSSP, each task corresponds to a ground target to be observed with a visible time window (Visible Time Window, VTW) and a specified duration. After completing the observation task, the satellite will obtain the corresponding benefits of the task. Different from traditional imaging satellites (Earth Observation Satellite, EOS), agile imaging satellites have three degrees of freedom: roll, pitch, and yaw. EOS can only observe directly above the target; while AEOS can observe before or after reaching the target. Therefore, the start time of the task observation of AEOS is variable within the VTW. The solution of AEOSSP not only includes the task observation sequence, but also the specific start and end observation times of each task. The main goal of AEOSSP is to maximize the total observation benefit by adjusting the task observation order and specific start time.

[0003] In practical applications, the commonly used planning period is usually in days. This means that if the satellite flight orbit and the ground target position are appropriate, the satellite may have multiple visible time windows for the same target within the planning period. Compared with the single-orbit mission planning problem, there are the following three main differences in the multi-orbit imaging satellite mission planning problem:

[0004] (1) The correspondence between tasks and visible time windows is different

[0005] In the single-orbit mission planning problem, the relationship between each task and the visible time window is one-to-one. However, in the multi-orbit planning process, each visible time window corresponds to a task, and a task may correspond to one or more visible time windows.

[0006] (2) The constraints of the multi-orbit mission planning problem are more complex

[0007] In the process of multi-orbit imaging satellite mission planning, the "uniqueness" constraint needs to be satisfied, that is, to ensure that each task can be observed at most once. During the decision-making process, once it is selected to observe a task in a certain visible time window, it must be ensured that other visible time windows corresponding to the task of this visible time window will not be selected in subsequent decisions.

[0008] (3) The scenario complexity of the multi-orbit mission planning problem is greater.

[0009] Compared with the single-orbit mission planning problem, the multi-orbit planning cycle is longer, so the number of time windows is more. Therefore, in the case of the same number of tasks to be planned, the scale of the multi-orbit mission scenario features increases exponentially compared with the single-orbit mission scenario, which poses great challenges to the computational efficiency of the algorithm and the storage of computing devices.

[0010] Therefore, for the multi-orbit imaging satellite mission planning problem, it is necessary to design a more reasonable solution process and strategy to achieve effective planning of observation tasks, so as to maximize the satellite observation efficiency. Summary of the Invention

[0011] The purpose of the present invention is to solve the above technical problems and propose a multi-orbit mission planning method for imaging satellites based on deep reinforcement learning.

[0012] To achieve the above purpose, the technical solution adopted by the present invention to solve the above technical problems is:

[0013] A multi-orbit mission planning method for imaging satellites based on deep reinforcement learning, characterized in that the mission planning method includes: a multi-orbit mission planning method based on the Markov decision process and a multi-orbit mission planning method based on orbit splitting;

[0014] Among them, the multi-orbit mission planning method based on the Markov decision process uses a multi-orbit neural policy network MNPN to directly learn the multi-orbit mission planning strategy, while the task planning problem based on orbit splitting uses a single-orbit neural policy network SNPN to learn the single-orbit mission planning strategy in the multi-orbit scenario;

[0015] The multi-orbit mission planning method based on the Markov decision process expands the observation and planning cycle, and directly learns the solution strategy of the multi-orbit mission planning by using a deep neural policy network. In the decision-making process, the dependence on the task concept is reduced. Starting from the initial moment, the multi-orbit neural policy network MNPN selects a suitable visible time window according to the current scene time, and determines the imaging start time of the task according to the imaging duration requirement of the corresponding task; once the specific start observation time is determined, update the observation end time of the task to the current scene time, and perform a visible time window constraint check; MNPN selects the next visible time window again according to the current state, and repeats this process until there is no visible time window that can be selected or the planning cycle ends, and the scene planning process ends, and a multi-orbit imaging scheme is output. The constraint check after selection by the policy network ensures that the task planning result meets the uniqueness constraint;

[0016] The multi-orbit task planning method based on Markov decision process regards all visible time windows of the tasks to be observed in the scene as independent decision windows, takes the observation timeline as the decision timeline to select windows in sequence, and determines the start and end times of specific tasks within the selected windows until there are no observable tasks or the observation period ends, thus solving the multi-orbit imaging satellite task planning problem; gives the 4-tuple {S, A, T, R} representation of the Markov decision process for the multi-orbit task planning problem of imaging satellites, specifically including the state set S, the action set A, the state transition function T, and the reward function R; among them, the state set S = {s 0 , s 1 , …, s T} includes all possible states at each decision step. In the multi-orbit task planning problem, the state s t at each decision step t = {time t , VTW t} consists of the time information time t and the visible time window state VTW t ; the time information time t represents the imaging time corresponding to the decision step t, and the visible time window state contains the states of all visible time windows in the scene. Among them, |VTW| represents the number of visible time windows in this scene, represents whether the task corresponding to the window vtw i has been observed before this moment. If it has not been observed, if it has been observed,

[0017] The multi-orbit task planning method based on orbit splitting calls the single-orbit neural policy network SNPN for single-orbit task planning in sequence during the splitting and solving process, and passes the unfinished tasks to the next orbit, repeating this process until all orbit task planning is completed. During the training process of the neural policy network SNPN, the method of transfer learning is used. First, use the SNPN with the characteristics of the multi-orbit task planning scene to learn the optimal strategy for single-orbit task planning, and transfer the trained policy network to the multi-orbit task planning training scene for further training. This method ensures the one-to-one correspondence between tasks and visible time windows during the decision-making process by splitting the multi-orbit planning problem into single-orbit task planning problems; when the imaging window of a task is determined, the subsequent single-orbit task planning will no longer consider the window corresponding to this task, thus ensuring that the final multi-orbit task planning result meets the uniqueness constraint of the problem; the training method based on transfer learning can also enable SNPN to take into account the ability of single-orbit decision-making and single-orbit decision-making in the multi-orbit scene during the process of single-orbit task planning;

[0018] The multi-round task planning method based on round splitting uses a single-round task planning neural policy network SNPN to represent the strategy for constructing single-round observation sequences. For each multi-round task scenario, first, the task set to be observed is updated. It can be seen that the number of visible time windows for each task in the scenario is different. The current task set to be observed is sequentially input into the single-round planning model. The model outputs the single-round observation plan for this round and updates the task set to be observed according to the observation results until all rounds of planning are completed or there are no tasks to be observed. During the model training process, when the model has completed the task planning for all rounds in the scenario, the model will update the parameters of the model according to the overall observation benefits of multiple rounds.

[0019] Further, the decision-making of the multi-round task planning method based on the Markov decision process is to select visible time windows. Therefore, the feature engineering of the problem also focuses on all visible time windows in the observation scenario. The features of the problem are divided into static features and dynamic features. The static features are known after the scenario is established and remain unchanged during the decision-making process, while the dynamic features are updated in real-time as the decision-making progresses; the numbers of the visible time windows are defined. It is known that there are M visible time windows in the problem scenario, and its window set VTW scenario ={vtw 1 , vtw 2 , …, vtw m , …, vtw M}}. The tasks to be observed corresponding to each visible time window vtw m are Analyze the self-attributes, corresponding task attributes, relationships between the same task windows, and window attributes to be planned of the visible time windows in the problem scenario one by one, and set appropriate features as the input for model decision-making.

[0020] Further, the multi-round task planning method based on the Markov decision process uses a multi-round neural policy network MNPN to directly learn the construction strategy of multiple rounds. MNPN uses five normalized features as the static feature input of the network. The Mask identification length of the method is the number of visible time windows. During each decision-making, the Mask mechanism will mask all visible time windows of the observed tasks and the visible time windows that violate the current constraints to ensure the feasibility of the solution. After the decision-making is completed, the Mask is updated.

[0021] Furthermore, in the multi-round task planning method based on round splitting, a training framework for the policy network based on transfer learning is designed. During the network training process, first, the SNPN with multi-round feature input is trained using the single-round training dataset; after the training is completed, model transfer is performed, and the SNPN trained on a single track is used as the initial model for multi-track scenario training; after the training of the single-round task planning sub-problem model is completed, the optimal parameter θ * s of the model will be used as the model initialization parameters for multi-round task scenario training; the multi-round task scenario training method is improved based on the REINFORCE algorithm with a rolling baseline to enable efficient training of the model; during the training process, each batch of training scenarios will be split into multiple single-round scenarios and solved sequentially. After each round of solution, Mask1 and Mask * 1 will be updated according to the task completion situation, so as to ensure that the planned tasks will not be selected in subsequent rounds; when all tasks in this batch are solved, the algorithm calculates the overall observed reward and updates the parameters.

[0022] Furthermore, the multi-round task planning method based on round splitting uses the single-round neural policy network SNPN to learn the solution strategy of the single-round task planning sub-problem after splitting the multi-round task planning problem. To ensure that the policy network can be effectively used to solve the multi-round problem, the feature input of SNPN is consistent with that of MNPN. The length of the Mask identifier is the number of scenario tasks. After each round of planning, SNPN will update the identifier Mask1 to represent the observed tasks. During the single-round planning process, Mask2 is used to represent the tasks that violate the problem constraints in the current state; SNPN will update Mask2 in each decision-making process, and the tasks finally masked by the Mask mechanism are the union of Mask1 and Mask2, Mask = Mask1 ∪ Mask2; during the training process, SNPN first conducts policy learning in the single-round training scenario to ensure the solution performance of the single-round task planning sub-problem; then, the trained SNPN is trained in the multi-round task planning scenario using the method of transfer learning to learn the strategy of single-round task planning in the multi-round scenario.

[0023] The present invention proposes a solution method for the multi-loop task planning problem. First, two multi-loop task planning methods with different solution processes are constructed: one based on the Markov decision process and the other based on loop splitting. The method based on the Markov decision process includes detailed modeling and feature engineering, and a corresponding MNPN policy model is designed to adapt to the complexity of multi-loop task planning; while the method based on loop splitting introduces a strategy model training method based on transfer learning to train the corresponding SNPN model to accelerate and optimize the learning process of the multi-loop task planning model. In the experimental verification stage, through comparative analysis, the performance of the two methods in the multi-loop task planning problem is demonstrated. By comparing MNPN and SNPN, the effectiveness of the proposed solution process and the multi-loop model training method based on transfer learning is analyzed. Through a more reasonable solution process and strategy, the effective planning of the observation task is realized, so as to maximize the satellite observation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 is the multi-loop task planning process of the present invention based on the Markov decision process;

[0026] Figure 2 is the multi-loop task planning process of the present invention based on loop splitting;

[0027] Figure 3 is the Markov decision process of the multi-loop task planning problem of the present invention;

[0028] Figure 4 is the multi-loop task solving process of the present invention based on loop splitting;

[0029] Figure 5 is the training process of the policy network based on transfer learning of the present invention;

[0030] Figure 6 is the analysis of the training processes of MNPN and SNPN of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The exemplary embodiments of the present application will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present application are included to assist in understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0032] The Agile Earth Observation Satellite Scheduling Problem (AEOSSP) involves the planning and scheduling of observation tasks for agile imaging satellites. In AEOSSP, each task corresponds to a ground target to be observed with a Visible Time Window (VTW) and a specified duration. After completing the observation task, the satellite will obtain the corresponding benefit of the task. Different from traditional Earth Observation Satellites (EOS), agile imaging satellites have three degrees of freedom: roll, pitch, and yaw. EOS can only observe directly above the target; while AEOS can observe before or after reaching above the target. Therefore, the start time of the task observation of AEOS is variable within the VTW. The solution of AEOSSP not only includes the task observation sequence but also the specific start and end observation times of each task. The main objective of AEOSSP is to maximize the total observation benefit by adjusting the task observation order and the specific start time.

[0033] In practical applications, the commonly used planning period is usually in days. This means that if the satellite flight orbit and the ground target position are appropriate, the satellite may have multiple visible time windows for the same target within the planning period. Compared with the single-orbit mission planning problem, there are the following three main differences in the multi-orbit imaging satellite mission planning problem:

[0034] (1) The correspondence between tasks and visible time windows is different

[0035] In the single-orbit mission planning problem, the relationship between each task and the visible time window is one-to-one. However, in the multi-orbit planning process, each visible time window corresponds to a task, and one task may correspond to one or more visible time windows.

[0036] (2) The constraints of the multi-orbit mission planning problem are more complex

[0037] In the process of multi - loop imaging satellite mission planning, the "uniqueness" constraint needs to be satisfied, that is, to ensure that each mission can be observed at most once. During the decision - making process, once a mission is selected for observation in a certain visible time window, it is necessary to ensure that other visible time windows corresponding to this mission will not be selected in subsequent decisions.

[0038] (3) The scenario complexity of the multi - loop mission planning problem is greater.

[0039] Compared with the single - loop mission planning problem, the multi - loop planning cycle is longer and the number of visible time windows is larger. Therefore, in the case of the same number of missions to be planned, the scale of the multi - loop mission scenario characteristics increases exponentially compared with the single - loop mission scenario, which poses great challenges to the computational efficiency of the algorithm and the storage of computing devices.

[0040] Therefore, for the multi - loop imaging satellite mission planning problem, it is necessary to design a more reasonable solution process and strategy to achieve effective planning of observation missions, so as to maximize the satellite observation efficiency.

[0041] Problem description

[0042] Problem parameters

[0043] Table 1 Model variables

[0044]

[0045]

[0046] Problem model

[0047] The optimization goal of the agile imaging satellite mission planning problem is to maximize the total observation benefit of the observation sequence, which can be expressed as:

[0048]

[0049] The problem has the following constraints:

[0050] (1) The actual observation time of the mission should be within the visible time window of the mission:

[0051]

[0052] (2) The actual observation duration of the mission meets the requirements of the mission's continuous observation time:

[0053]

[0054] (3) In each scheduling period, the mission can be observed at most once:

[0055]

[0056] (4) Conversion time constraint between two tasks:

[0057]

[0058] (5) Calculation of conversion time for agile imaging satellites:

[0059] The attitude conversion time of the agile satellite has a time-dependent characteristic. According to the satellite attitude, its conversion time can be expressed by a piecewise function, and the specific calculation formula is as follows:

[0060]

[0061] Among them, Δθ ij represents the attitude transformation angle between two tasks task i and task j , and the calculation formula is as follows:

[0062]

[0063] According to the attitude change angle, a piecewise function can be used to calculate the conversion time of the agile satellite. For two given tasks task i and task j :

[0064]

[0065] Problem solving

[0066] Solution framework and solution process design for multi-orbit mission planning method

[0067] Based on the single-orbit mission planning method, considering different mission planning requirements, the multi-orbit mission planning method proposes two solution processes for multi-orbit mission planning: the multi-orbit mission planning method based on the Markov decision process and the multi-orbit mission planning method based on orbit splitting. The solution frameworks of both are constructed by a neural policy network. The difference is that the method based on the Markov decision process uses a multi-circle neural policy network (MNPN) to directly learn the multi-orbit mission planning strategy, while the mission planning problem based on orbit splitting uses a single-circle neural policy network (SNPN) to learn the single-orbit mission planning strategy in the multi-orbit scenario.

[0068] (1) Multi-orbit mission planning process based on the Markov decision process

[0069] The solution process of the multi-orbit mission planning method based on the Markov decision process is as Figure 1As shown in the figure. This method expands the observation and planning cycle and directly learns the solution strategy for multi-orbit task planning using a deep neural policy network. During the decision-making process, this method reduces the dependence on task concepts. From the initial moment, the multi-orbit neural policy network (MNPN) selects an appropriate visible time window according to the current scene time and determines the imaging start time of the task according to the imaging duration requirements of the corresponding task. Once the specific start observation time is determined, the observation end time of this task is updated to the current scene time, and a visible time window constraint check is performed. MNPN selects the next visible time window again according to the current state and repeats this process until there is no visible time window to select or the planning cycle ends. The scene planning process then ends, and a multi-orbit imaging plan is output. The method ensures that the task planning result meets the uniqueness constraint through the constraint check after the selection by the policy network.

[0070] (2) Multi-orbit task planning method based on orbit splitting

[0071] The solution process of the multi-orbit task planning method based on orbit splitting is as Figure 2 shown in the figure. This method splits the multi-orbit task planning problem into multiple single-orbit task planning sub-problems.

[0072] In the splitting solution process, the single-orbit neural policy network (SNPN) is called in sequence for single-orbit task planning, and the unfinished tasks are passed to the next orbit, repeating this process until all orbit task planning is completed. During the training process of the neural policy network SNPN, the present invention uses the method of transfer learning. First, SNPN with the characteristics of multi-orbit task planning scenarios is used to learn the optimal strategy for single-orbit task planning, and the trained policy network is transferred to the multi-orbit task planning training scenario for further training. This method ensures the one-to-one correspondence between tasks and visible time windows during the decision-making process by splitting the multi-orbit planning problem into single-orbit task planning problems. When the imaging window of a task is determined, subsequent single-orbit task planning will no longer consider the corresponding window of this task, thereby ensuring that the final multi-orbit task planning result meets the uniqueness constraint of the problem. At the same time, the training method based on transfer learning can also enable SNPN to take into account the ability of single-orbit decision-making and single-orbit decision-making in multi-orbit scenarios during the process of single-orbit task planning.

[0073] Multi-orbit task planning method based on Markov decision process

[0074] Markov decision process modeling of the multi-orbit task planning problem of imaging satellites

[0075] The Markov decision process for the multi-orbit mission planning problem of imaging satellites regards all visible time windows of the tasks to be observed in the scene as independent decision windows. Taking the observation timeline as the decision timeline, window selection is carried out in sequence, and the start and end times of specific tasks are determined within the selected windows until there are no observable tasks or the observation period ends, thus completing the solution to the multi-orbit imaging satellite mission planning problem.

[0076] Figure 3 The Markov decision process of multi-orbit mission planning for imaging satellites under the guidance of two different solution strategies is introduced. The number of each visible time window represents "task number - visible time window number". In the initial state, the imaging time corresponding to decision step t0 is the start time of the scene, and at this time all visible time windows can be selected. In this state, both strategies select window "4-1". After selecting window "4-1", the state is updated to decision step t1, and the corresponding imaging time at this time is the imaging end time of task 4 in window "4-1". In this state, window "3-1" cannot be selected because it does not meet the imaging time constraint, and the other windows "4-2" and "4-3" corresponding to task 4 cannot be selected because task 4 has already been observed. At this time, window "1-1" and window "5-1" are relatively close, but task 1 has three visible time windows while task 5 has only one. According to different strategies, Strategy 1 selects window "1-1" (as Figure 3 (b)), and Strategy 2 selects window "5-1" (as Figure 3 (c)). Due to the decision differences at this decision step, the subsequent decisions of the two strategies are also different. After a series of "state update - decision", Strategy 1 finally obtains the observation sequence of tasks as "4→1→2→3", while Strategy 2 obtains the observation sequence "4→5→2→3→1". Thus, it can be seen that in the process of solving the multi-orbit mission planning problem of imaging satellites, it is very necessary to consider the relationship between multiple windows of a single task to improve the overall planning performance of the strategy.

[0077] Next, the present invention gives the 4-tuple {S, A, T, R} representation of the Markov decision process for the multi-orbit mission planning problem of imaging satellites. The specific state set S, action set A, state transition function T, and reward function R are defined as follows:

[0078] (1) State set S

[0079] The state set S = {s 0 , s 1 , …, s T} includes all possible states at each decision step. In the multi-orbit mission planning problem, the state s t at each decision step t = {time t , VTW t} It includes time information time t and the visible time window state VTW t to form. The time information time t represents the imaging time corresponding to the decision step t, and the visible time window state includes the states of all visible time windows in the scene. Among them, |VTW| represents the number of visible time windows in this scene, indicating whether the task corresponding to the window vtw i has been observed before this moment. If it has not been observed, if it has been observed,

[0080] (2) Action set A

[0081] In the multi-orbit task planning problem, for each state s t the action a that can be taken t corresponds to the visible time window that can be selected. To ensure that the satellite performs only one task observation at a time, each action a t corresponds to a visible time window. At the decision step t, the action set A t includes all visible time windows except those that violate the imaging time constraint or the windows corresponding to the observed tasks. In addition, if there are N st selectable visible time windows at the decision step t, then the length of the action set A t is N st +1, and the action indicates the end of the decision-making process.

[0082] (3) State transition function T

[0083] If the task corresponding to the visible time window vtw selected at the decision step t st is task st , then the method to update the next state s t+1 is as follows:

[0084]

[0085] Equation (9) indicates that the end time of the task task corresponding to the window selected at the decision step t st is the time of the state s t+1 ; Equation (10) indicates that if the task task is observed at the decision step t st , then all visible time windows corresponding to this task cannot be selected at the decision step t + 1.

[0086] (4) Reward function R

[0087] The reward function of the multi-orbit task planning problem is the same as that of the single-orbit task planning problem and can also be expressed as Among them, r t represents the reward obtained from the task corresponding to the visible time window selected for decision step t.

[0088] Feature engineering

[0089] The decision-making of the multi-round task planning method based on the Markov decision process is to select the visible time window. Therefore, the feature engineering of the problem also focuses on all the visible time windows in the observation scenario. The features of the problem are divided into static features and dynamic features. The static features are known after the scenario is established and remain unchanged during the decision-making process, while the dynamic features are updated in real time as the decision-making progresses. For the convenience of description, we define the numbers of the visible time windows. It is known that there are M visible time windows in the problem scenario, and its window set VTW scenario ={vtw 1 , vtw 2 , …, vtw m , …, vtw M}, and the task to be observed corresponding to each visible time window vtw m is Therefore, the self-attributes of the visible time windows in the problem scenario, the corresponding task attributes, the relationship attributes between the same task windows, and the attributes to be planned for the windows are analyzed one by one, and appropriate features are set as the input for model decision-making, and their rationality is explained.

[0090] (1) Self-attributes of visible time windows

[0091] The self-attributes of the visible time windows represent the position information of the windows within the planning period. The method selects the start time and the end time of the visible time window to characterize its self-attributes.

[0092] (2) Corresponding task attributes of visible time windows

[0093] The ultimate goal of selecting the visible time window is to complete the observation of the task. Then, certain features are needed to characterize the characteristics of the task corresponding to the window. The present invention selects the imaging duration and the task benefit of the task corresponding to the visible time window as features. Among them, the imaging duration represents the time required for the task within the planning period, and its normalized result is the ratio to the length of the planning period; while the task benefit reflects the importance of the task in the entire task to be observed, and its ratio to the maximum task benefit is used as the normalized result.

[0094] (3) Relationship attributes between visible time windows of the same task

[0095] ​The biggest feature of the multi-orbit task planning scenario is that a single task may have multiple visible time windows. If a task preempts the observation opportunity of a task with only one visible window, it will lead to a waste of observation resources. Therefore, extracting features to represent the relationship between different visible windows of the same task is very important for ensuring the accuracy and efficiency of model decision-making. When selecting the attribute features of the visible time window relationship, the present invention proposes two attributes for analysis: task number i m and the remaining number of visible time windows Among them, task number i m represents the number of the task to be observed corresponding to this visible time window in the set, and uses the ratio of it to the maximum task number as the normalization result to identify the same task. However, considering that the tasks to be observed are in an equal relationship, and the task numbers cause differences in different visible time windows due to the task order, which is not conducive to model decision-making. And the remaining number of visible time windows represents the number of remaining visible time windows of this task after the corresponding moment of this visible time window, and performs normalization processing using the ratio of it to the maximum number of visible time windows in the scenario. This avoids introducing task-related information and can effectively link the visible time windows of the same task, which is more in line with the original intention of selecting visible time windows in the multi-orbit Markov decision-making process.

[0096] (4) Attributes of the to-be-planned state of the visible time window

[0097] Attribute features of the visible time window's plannability Indicates whether this window can be selected at decision step t. If the task corresponding to this window has been selected before decision step t, then Otherwise,

[0098] Design of the multi-orbit neural policy network

[0099] The multi-orbit task planning method based on the Markov decision-making process uses a multi-orbit neural policy network (MNPN) to directly learn the construction strategy of multiple orbits. MNPN uses five normalized features as the static feature input of the network, expressed as The Mask identification length of the method is the number of visible time windows. At each decision, the Mask mechanism will mask all visible time windows of the observed tasks and the visible time windows that violate the current constraints to ensure the feasibility of the solution. After the decision is completed, the Mask is updated.

[0100] Multi-orbit task planning method based on orbit splitting and solution

[0101] Solution process of the multi-orbit task planning method based on orbit splitting and solution

[0102] The solution process of the multi-orbit task planning method based on orbit splitting is as followsFigure 4 As shown in the figure, the method uses a single-loop sub-task planning neural policy network (SNPN) to construct a policy representation for the single-loop observation sequence. For each multi-loop task scenario, the method first updates the set of tasks to be observed. It can be seen that the number of visible time windows for each task in the scenario is different. Based on this, the current set of tasks to be observed is input into the single-loop planning model in sequence. The model outputs the single-loop observation plan for this loop, and updates the set of tasks to be observed according to the observation results. The above steps are repeated in sequence until all loop planning is completed or there are no tasks to be observed. During the model training process, when the model has completed the task planning for all loops in the scenario, the model will update the parameters of the model according to the overall observation benefit of the multi-loop.

[0103] Single-loop neural policy network design

[0104] The task planning method based on loop splitting uses a single-loop neural policy network (SNPN) to learn the solution strategy for the single-loop task planning sub-problem after splitting the multi-loop task planning problem. To ensure that the policy network can be effectively used to solve the multi-loop problem, the feature input of SNPN is consistent with that of MNPN. The Mask identification length of the method is the number of scenario tasks. After each loop planning is completed, SNPN will update the identification Mask1 to represent the observed tasks, and Mask2 is used to represent the tasks that violate the problem constraints in the current state during the single-loop planning process. SNPN will update Mask2 in each decision-making process, and the tasks that the final Mask mechanism will mask are the union of Mask1 and Mask2, Mask = Mask1 ∪ Mask2. During the training process, SNPN first conducts policy learning in the single-loop training scenario to ensure the solution performance of the single-loop task planning sub-problem; then uses the method of transfer learning to train the trained SNPN in the multi-loop task planning scenario to learn the strategy of single-loop task planning in the multi-loop scenario. The specific SNPN training method will be introduced in detail later.

[0105] Multi-loop model training method based on transfer learning

[0106] In the multi-loop task planning method based on the loop splitting solution process, the present invention designs a policy network training framework based on transfer learning, as Figure 5 shown.

[0107] During the network training process, the method first uses the single-loop training data set to train the SNPN with multi-loop feature input. After the training is completed, model migration is performed, and the SNPN trained on the single track is used as the initial model for multi-track scenario training. After the training of the single-loop task planning sub-problem model is completed, the optimal parameter number θ of the model *s will be used as the model initialization parameters for multi-round task scenario training. The pseudo-code of the multi-round task scenario training algorithm is shown in Algorithm 1. The method is improved based on the REINFORCE algorithm with a rolling baseline to enable efficient training of the model. During the training process, each batch of training scenarios will be split into multiple single-round scenarios and solved sequentially (Steps 7-10). After each round of solution, Mask1 and Mask will be updated according to the task completion status (Step 10), which can ensure that the planned tasks will not be selected in subsequent rounds. When all tasks in this batch are solved, the algorithm calculates the overall observed reward and updates the parameters (Steps 12-15). * 1 (Step 10), which can ensure that the planned tasks will not be selected in subsequent rounds. When all tasks in this batch are solved, the algorithm calculates the overall observed reward and updates the parameters (Steps 12-15).

[0108]

[0109] Experimental Analysis

[0110] To verify the effectiveness of the two methods in solving this problem, the high efficiency of the method in solving quality and the rationality of the solution process are respectively illustrated through comparative experiments and result analysis. The experiment generated a multi-round scenario set with task sizes of 50, 100, and 150. Among them, the size of the training scenario set is 256,000, and the size of the test scenario is 1,000. The experiment improved the construction heuristic rules and reconstruction mechanism to solve the multi-round task planning problem, and then verified the effectiveness of the proposed method. All model training and experiments were written in Python 3.8. Deep reinforcement learning was based on PyTorch 1.9.0, the compilation tool was PyCharm, and the experimental environment was ubuntu18.04 + cuda11.1. The GPU model of all experimental and model training platforms was TITAN RTX, and the CPU was Intel i9-11900K CPU.

[0111] Analysis of the Training Process

[0112] During the training process, this experiment trained the proposed MNPN and SNPN using training scenarios with task sizes of 50, 100, and 150. Due to device memory limitations, 256,000 training data with a scenario size were used for training MNPN. The scenario batch sizes for tasks with sizes of 50, 100, and 150 were 512, 128, and 128 respectively, and each size was trained for 20 epochs. SNPN was trained using 256,000 scenarios of different sizes in a single-loop scenario. The batch size for scenarios under the three task sizes was 512, and each size was trained for 20 epochs. In the multi-loop scenario, a total of 64,000 training scenarios were used for training with a batch size of 128, and each size was trained for 5 epochs. The training benefit iteration diagrams of the two models are as shown in Figure 6 shown.

[0113] Analyzing the benefit iteration diagrams, all training processes have gone through a process of benefit improvement and tending to be stable, proving the stability and effectiveness of the training method. Specifically, analyze the training situation in different scenarios. The multi-scenario training of MNPN can quickly achieve benefit improvement at the initial stage of training for the three task sizes, and tend to be stable and have a slight increase after 5 epochs. The single-loop scenario training of SNPN can also achieve convergence, and the benefit is significantly improved in the first 5 epochs, then maintains a small increase, and remains stable after 15 epochs. During the multi-loop scenario training of SNPN, SNPN can achieve convergence after 2 epochs when training in the scenario with a task size of 50, while the benefit still has a significant increase in the 5th epoch when training in the scenarios with task sizes of 100 and 150.

[0114] Solving Performance Analysis

[0115] The present invention respectively designs MNPN for the Markov decision process and SNPN for loop splitting according to different solving processes to solve the multi-loop problem scenario. The calculation results of directly solving 1,000 test scenarios with task sizes of 50, 100, and 150 by the constructed heuristic rules and MNPN and SNPN are counted to analyze the solving performance of the policy network. The calculation results of the experiment are shown in Table 2.

[0116] Table 2 Experimental Results of Solving Performance of Policy Network

[0117]

[0118] It can be seen from the experimental results that both MNPN and SNPN can achieve relatively high average returns under different task scales. The average return of SNPN is higher than that of MNPN in the scenario with a task scale of 50, while in the scenarios with 100 and 150 tasks, the return of MNPN is higher. In contrast, the average returns of other methods are generally low and their performance is relatively unstable. Analyzing from the aspect of computing time, both MNPN and SNPN have short computing times in the scenarios and are more stable as the scale increases, indicating that they can balance high solution performance and relatively high computing efficiency.

[0119] Analysis of the Training Method of the Policy Network Based on Transfer Learning

[0120] In this experiment, MNPN trained with multi-round scenarios, SNPN (SNPN-1) trained with single-round scenarios, and SNPN (SNPN-2) trained with multi-round scenarios were used to solve the test scenarios with task scales of 50, 100, and 150, so as to verify the effectiveness of the proposed training method of the policy network based on transfer learning and provide a basis for further research.

[0121] Table 3 Experimental Results of the Test of the Models Trained by Different Processes

[0122]

[0123]

[0124] The experimental results are shown in Table 3. In addition to the average return, the experiment also calculated the return growth ratios of SNPN-2 and MNPN compared with SNPN-1. It can be observed from the table content that in all test scenarios, the solution returns of SNPN-2 are significantly better than those of SNPN-1, indicating that the proposed training method of the policy network based on transfer learning can effectively improve the solution performance of the single-round policy model for solving multi-round task planning scenarios. Although MNPN has better solution ability than SNPN-2 in most scenarios, the training cost of SNPN-2 is relatively small. Through the training method of transfer learning, it only needs 5 rounds, with a short iteration period and stronger operability.

[0125] The present invention deeply studies the solution methods for the multi-loop mission planning problem. First, two multi-loop mission planning methods with different solution processes are constructed: one based on Markov decision process and the other based on loop splitting. The method based on Markov decision process includes detailed modeling and feature engineering, and a corresponding MNPN policy model is designed to adapt to the complexity of multi-loop mission planning; while the method based on loop splitting introduces a strategy model training method based on transfer learning to train the corresponding SNPN model to accelerate and optimize the learning process of the multi-loop mission planning model. In the experimental verification stage, through comparative analysis, the performance of the two methods in the multi-loop mission planning problem is demonstrated. In addition, in the experimental stage, by comparing MNPN and SNPN, the effectiveness of the proposed solution process and the multi-loop model training method based on transfer learning is analyzed.

[0126] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A multi-turn mission planning method for imaging satellite based on deep reinforcement learning, characterized in that: The task planning method includes: a multi-turn task planning method based on Markov decision process and a multi-turn task planning method based on turn splitting; The multi-turn task planning method based on Markov decision process uses a multi-turn neural policy network MNPN to directly learn the multi-turn task planning strategy, while the task planning problem based on turn splitting uses a single-turn neural policy network SNPN to learn the single-turn task planning strategy in the multi-turn scenario; The multi-turn task planning method based on the Markov decision process directly learns the solution strategy of multi-turn task planning by extending the observation and planning cycle and using the deep neural policy network. In the decision-making process, the dependence on the task concept is reduced. Starting from the initial moment, the multi-turn neural policy network MNPN selects a suitable visible time window according to the current scene time, and determines the imaging start time of the task according to the imaging duration requirement of the corresponding task; once the specific start observation time is determined, the observation end time of the task is updated to the current scene time, and a visible time window constraint check is performed; MNPN selects the next visible time window again according to the current state, and repeats the process until there is no visible time window to be selected or the planning cycle ends, the scene planning process is declared over, and a multi-turn imaging plan is output, and the constraint check after the policy network selection ensures that the task planning result satisfies the uniqueness constraint; The multi-turn mission planning method based on Markov decision process regards all visible time windows of the tasks to be observed in the scene as independent decision windows, selects windows in sequence with the observation timeline as the decision timeline, and determines the start and end time of the specific task in the selected window until there are no observable tasks or the observation cycle ends, thereby completing the solution of the multi-turn imaging satellite mission planning problem; a 4-tuple {S, A, T, R} expression of the Markov decision process of the imaging satellite multi-turn mission planning problem is given, which specifically includes a state set S, an action set A, a state transfer function T and a reward function R; wherein the state set S = {s0, s1, …, S T } includes all possible states of each decision step. In the multi-turn task planning problem, the state s of each decision step t is t ={time t ,VTW t }Including time information time t and visible time window status VTW t Composition; time information time t Indicates the imaging time corresponding to the decision step t, the visible time window state Contains the status of all visible time windows in the scene, where |VTW| represents the number of visible time windows in the scene. Indicates that the window vtw before this time i Whether the corresponding task has been observed, if not, If it has been observed, The multi-turn task planning method based on lap splitting sequentially calls the single-turn neural policy network SNPN to perform single-turn task planning in the splitting and solving process, and passes the unfinished tasks to the next lap, and repeats the process until all lap task planning is completed. In the training process of the neural policy network SNPN, a transfer learning method is used. First, the SNPN with multi-turn task planning scene features is used to learn the optimal strategy for single-turn task planning, and the trained strategy network is migrated to the multi-turn task planning training scene for further training. This method splits the multi-turn planning problem into a single-turn task planning problem to ensure a one-to-one correspondence between tasks and visible time windows in the decision-making process; when the task determines the imaging window, the subsequent single-turn task planning will no longer consider the window corresponding to the task, so as to ensure that the final multi-turn task planning result meets the uniqueness constraint of the problem; the training method based on transfer learning can also enable SNPN to take into account the ability of single-turn decision-making and single-turn decision-making in multi-turn scenarios during the process of single-turn task planning; The multi-lap task planning method based on lap splitting uses a single-lap task planning neural strategy network SNPN to perform single-lap observation sequence construction strategy representation. For each multi-lap task scenario, the task set to be observed is first updated. It can be seen that the number of visible time windows for each task in the scenario is different. The current task set to be observed is input into the single-lap planning model in turn. The model outputs the single-lap observation plan for the lap and updates the task set to be observed according to the observation results until all lap planning is completed or there are no tasks to be observed. During the model training process, when the model completes task planning for all laps in the scenario, the model will update the model parameters according to the overall observation benefits of multiple laps.

2. The method according to claim 1, characterized in that The decision of the multi-turn task planning method based on the Markov decision process is to select the visible time window. Therefore, the feature engineering of the problem is also carried out around all the visible time windows in the observation scene. The characteristics of the problem are divided into static characteristics and dynamic characteristics. The static characteristics are known after the scene is established and remain unchanged during the decision-making process, while the dynamic characteristics will be updated in real time with the decision. The number of the visible time window is defined. It is known that there are M visible time windows in the problem scene, and its window set VTW scenario ={vtw1,vtw2,…,vtw m ,…,vtw M }, each visible time window vtw m The corresponding task to be observed is The properties of the visible time window in the problem scenario, the corresponding task properties, the relationship properties between windows of the same task, and the properties of the window to be planned are analyzed one by one, and appropriate features are set as model decision input.

3. The method according to claim 1, characterized in that The multi-turn task planning method based on Markov decision process uses a multi-turn neural policy network MNPN to directly learn the multi-turn construction strategy. MNPN uses five normalized features as static feature inputs of the network. The Mask identification length of the method is the number of visible time windows. At each decision, the Mask mechanism will cover up all visible time windows of the observed tasks and the visible time windows that violate the current constraints to ensure the feasibility of the solution. After the decision is made, the Mask is updated.

4. The method according to claim 1, characterized in that: In the multi-turn task planning method based on lap splitting, a policy network training framework based on transfer learning is designed. In the network training process, the single-lap training data set is first used to train the SNPN with multi-lap feature input; after the training is completed, the model is transferred, and the SNPN trained on the single track is used as the initial model for multi-track scenario training; after the single-lap task planning sub-problem model training is completed, the optimal parameter number θ of the model is * s will be used as the model initialization parameter for multi-turn mission scenario training; the multi-turn mission scenario training method is improved on the basis of the REINFORCE algorithm with rolling baseline, so that it can train the model efficiently; During the training process, each batch of training scenarios will be split into multiple single-cycle scenarios and solved in sequence. After each cycle is solved, Mask1 and Mask2 will be updated according to the task completion status. * 1. This ensures that the planned tasks will no longer be selected in subsequent rounds. When all tasks in the batch are solved, the algorithm calculates the overall observation benefit and updates the parameters.

5. The method according to claim 1, characterized in that The multi-turn task planning method based on round splitting uses a single-turn neural policy network SNPN to learn the solution strategy of the single-turn task planning sub-problem after the multi-turn task planning problem is split. In order to ensure that the policy network can be effectively used for solving multi-turn problems, the feature input of SNPN is consistent with that of MNPN, and the length of the Mask identifier is the number of scenario tasks. After each round of planning, SNPN will update the identifier Mask1 to indicate the observed tasks, and Mask2 is used in the single-turn planning process to indicate the tasks that violate the problem constraints in the current state; SNPN will update Mask2 in each decision-making process, and the task that will be blocked by the final Mask mechanism is the union of Mask1 and Mask2, Mask = Mask1∪Mask2; during the training process, SNPN first performs strategy learning in a single-lap training scenario to ensure the performance of solving the single-lap task planning sub-problem; then the trained SNPN is trained in a multi-lap task planning scenario using the transfer learning method to learn the strategy of single-lap task planning in a multi-lap scenario.