Heterogeneous space platform collaborative earth observation autonomous task planning method and device
By employing a two-stage interactive solution process based on deep reinforcement learning, the problem of low planning efficiency for collaborative Earth observation missions across heterogeneous space platforms was solved, achieving efficient task allocation and planning, and improving observation benefits and resource utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF CHINESE ACAD OF SCI
- Filing Date
- 2025-12-12
- Publication Date
- 2026-05-01
AI Technical Summary
Existing methods for planning collaborative Earth observation missions using heterogeneous space platforms are inefficient, have long decision-making times, and require multiple communications between space and ground, leading to decision-making delays. These methods are insufficient to meet the needs of highly dynamic and complex observation missions requiring rapid response.
A two-stage interactive solution process based on deep reinforcement learning is adopted, including adaptive allocation operators and destruction operators in the task allocation stage, and attention mechanism decision-making of encoder and decoder in the task planning stage, to establish a multi-agent system for task allocation and planning.
It improved the utilization rate of time resources of heterogeneous space platforms, maximized observation benefits, and reduced the runtime of mission planning.
Smart Images

Figure CN121961035A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of space-based Earth observation technology, and is a method and apparatus for autonomous mission planning of heterogeneous space-based Earth observation based on deep reinforcement learning. Background Technology
[0002] Space-based Earth observation has a wide range of applications, playing a vital role in disaster relief, agriculture, forestry, animal husbandry, fisheries, and environmental monitoring. To adapt to the demands of highly dynamic and complex observation missions, it is necessary to coordinate heterogeneous space platforms such as satellites, airships, and unmanned aerial vehicles (UAVs) to leverage their complementary advantages. Space-based Earth observation involves various uncertainties, such as the emergence of new missions, space platform malfunctions, and obstruction of observation tasks, necessitating autonomous decision-making to meet new requirements such as rapid response in space-based Earth observation. Therefore, it is necessary to develop autonomous mission planning methods for collaborative Earth observation using heterogeneous space platforms.
[0003] Current research on collaborative Earth observation mission planning for heterogeneous space platforms is limited. Mission planning efficiency is constrained by the metaheuristic algorithms and frameworks employed. For example, metaheuristic algorithms have long decision-making times and require multiple ground-to-space communications, leading to decision delays. This invention addresses the autonomous mission planning problem for collaborative Earth observation using heterogeneous space platforms such as Earth observation satellites, aerostats, and UAVs. It considers point target observation tasks and maximizes the total observation mission benefit under constraints such as the visible time window of the satellite observation task, satellite attitude transition time, aerostat and UAV cruise time, and maximum endurance of the aerostat and UAV. A deep reinforcement learning-based autonomous mission planning method is proposed, comprising a two-stage solution framework of task allocation and task planning. In the task allocation stage, where deep reinforcement learning methods struggle to determine rewards, a metaheuristic approach is adopted to effectively allocate tasks through various heuristic rules. The task planning stage employs a deep reinforcement learning algorithm, reducing the runtime of the mission planning process. A multi-agent system is established to describe the autonomous mission planning problem for collaborative Earth observation on heterogeneous space platforms. This system includes one allocation agent (AA) and several scheduling agents (SA). The AA is responsible for allocating observation tasks to the scheduling agents, and the SAs are responsible for autonomous mission planning. This invention proposes a deep reinforcement learning-based method for autonomous mission planning in collaborative Earth observation on heterogeneous space platforms, providing support for the field of space-based Earth observation technology. Summary of the Invention
[0004] The purpose of this invention is to provide a method and apparatus for autonomous mission planning of collaborative Earth observation on heterogeneous space platforms based on deep reinforcement learning. A two-stage interactive solution process is designed. In the mission allocation stage, an adaptive allocation process based on multiple allocation operators is proposed, designing five allocation operators and five violation operators. The allocation and violation operators coordinate to solve the mission allocation sub-problem. In the mission planning stage, an attention mechanism based on encoders and decoders is developed to determine the scheduling sequence of tasks and solve the observation mission sequence sub-problem.
[0005] The present invention provides a method for autonomous mission planning of collaborative Earth observation on heterogeneous space platforms based on deep reinforcement learning, comprising the following steps: S1. Establish an autonomous planning model for collaborative Earth observation using heterogeneous space platforms; S2. Generate initial solutions based on deep reinforcement learning; S3. Generating neighborhood solutions based on deep reinforcement learning; S4: Update the current solution, the optimal solution, and the policy weights; S5. Repeat S3 and S4. If the best solution has not been updated after the preset number of iterations, the algorithm terminates and outputs the current best solution.
[0006] Optionally, S1 further includes: S101: defining the autonomous planning problem for collaborative Earth observation by heterogeneous space platforms on a directed graph G = (N, A); N is the set of network nodes, A is the set of directed arcs, and the decision period is T; the set N = {D∪C} consists of virtual nodes. D = {0, σ} and observation task nodes C = {1, 2, …, n Composed of}, where {0} is the virtual starting point, { σ} represents the virtual endpoint obtained by copying {0}; the set of directed arcs is . ,in N 0 = {0}∪ C , N σ = C ∪{ σ}; For observation tasks i ∈ C Observational benefits are The observation duration is The observation time requirement is The cruising distance between the airship and the drone between the two observation missions is... and The cruising speeds of the airship and the drone are respectively and The Earth observation satellite set is Sa ={ s | s = 1, 2, …, n s Earth observation satellites i For observation tasks j The number of visible time windows is The set of visible time windows is , of which k The time window is Earth observation satellites s In observation mission i and j The attitude transition time between them is In other words, the arc { in the directed graph i , j The length of}, and the conversion time between virtual points and observation task nodes is 0; S102: Establish a multi-agent system to describe the autonomous task planning problem of collaborative Earth observation on heterogeneous space platforms. This multi-agent system includes a task allocation agent AA and several task planning agents SA. AA is used to allocate observation tasks to the task planning agents, and SA is used to perform autonomous task planning. For the task allocation problem, establish a task allocation model with the goal of maximizing the total observation benefits. For the autonomous task planning problem, establish a task planning model based on Markov decision process. The goal of the task allocation model is to maximize the total revenue of all space platforms: Max (1) M A collection of space platforms N For the set of observation tasks, x ij For binary decision variables, when the observation task j Assigned to space platform i hour x ij = 1, otherwise 0; z ij As a binary variable, when the space platform i The observation mission was completed. j hour, z ij = 1, otherwise 0; p j For observation mission j The benefits; The task allocation process satisfies a uniqueness constraint, meaning that a task can only be assigned to one space platform: ST , (2) The feasibility of the space platform for the mission was considered during the mission allocation process. The space platform met the mission's payload type requirements, resolution requirements, and the visible time window of the observation mission. Furthermore, preprocessing was performed during the mission and space platform matching phase to ensure that the space platform could execute the assigned mission. After the AA assigns observation tasks to various space platforms, the SA performs autonomous mission planning, establishing a mission planning model based on a Markov decision process. The system starts from the initial state. S Starting from 0, SA is based on S 0 Perform an action A 0, received by the environment A After 0, calculate the action. A 0 reward R 1, and reach the next stage state. S 1; SA is based on S i Continuous output action A i This continues until a termination state is reached; each interaction between the SA and the environment is called a time step. (1) States: The set of states forms the state space: (3) S i It is the first t The state is defined for each time step, and each state contains a set of attributes that describe the state information at a specific time step. These attributes are divided into two parts: static and dynamic information. For satellites, static information includes the benefits and duration of the observation mission, while dynamic information includes the total remaining observation opportunities, the earliest time that the observation can be completed, the latest time that the observation can begin, and whether the mission has been allocated to a space platform and the execution time. For airships and UAVs, static information includes the benefits and duration of the observation mission, while dynamic information includes the total time consumed by the space platform from the current point to the completion of the observation mission and the remaining available time of the space platform. (2) Action: Action is the decision made by SA based on the current state, that is, to select an observation task; (3) Rewards: The rewards that SA receives from the environment are not always equal to the observation gains of the selected observation task, but rather the increment of the total observation gains after taking action; at each time step, the environment receives the decision action of SA and automatically feeds back the corresponding rewards and updated status. (4) Value function: The value function defines the average reward of long-term actions. A deep network is built as an estimate of the state value function.
[0007] Optionally, S2 includes: generating an initial solution includes: initializing the weights of the allocation strategy and the destruction strategy, initializing network parameters; performing initial task allocation using the task allocation strategy, and initial task planning based on a deep reinforcement learning method; S201: The allocation strategies used for task allocation are divided into allocation strategies between heterogeneous space platforms and between homogeneous space platforms, as well as upper-level strategies. The strategies between heterogeneous space platforms include allocation strategies between satellites and other space platforms, and allocation strategies between airships and UAVs. The strategies between homogeneous space platforms are allocation strategies between satellites. The upper-level strategies are the conditions that must be met when calling other allocation strategies, including that tasks are allocated to feasible space platforms, and the total number of observation tasks allocated to a certain type or a certain space platform cannot exceed the upper limit. ① Allocation between satellites and the other two types of space platforms: First allocation strategy 1: random allocation; without violating the upper-level strategy, AA will allocate observation tasks to satellites or other space platforms with equal probability; Second allocation strategy 2: Taboo strategy; observation tasks are allocated according to a random strategy, but not to space platforms that received the task in the previous round but did not execute it, unless the upper-level strategy is violated. ② Allocation between airships and drones: Third allocation strategy 3: random allocation; without violating the upper-level strategy, AA will allocate observation tasks to airships or drones with equal probability; Fourth allocation strategy 4: allocation based on proximity; without violating the upper-level strategy, AA will allocate the observation task to the nearest airship or UAV; ③ Satellite allocation: Fifth allocation strategy 5: random allocation; AA means that satellites are selected to perform observation tasks with equal probability; Sixth allocation strategy 6: Observation opportunity allocation; the longer the visibility time, the more likely it is to complete the observation task. Therefore, in this strategy, AA will allocate the observation task to the satellite with the longest total feasible observation time for the task. Seventh Allocation Strategy 7: Task Conflict Allocation; The longer the overlap between the observation time windows of observation tasks, the greater the possible observation conflict between observation tasks on that satellite; Therefore, in this strategy, AA will allocate observation tasks to the satellite with the lowest degree of conflict. Eighth allocation strategy 8: Historical experience allocation; In multiple iterations, for each observation task assigned to each satellite, the mean of the objective function values of the most recent iterations is calculated. In this strategy, AA will assign the observation task to the satellite that performed the task in the previous iterations and has the largest mean objective value. S202: AA selects the first allocation strategy 1, the third allocation strategy 3, and the fifth allocation strategy 5 to allocate tasks. SA performs task planning based on deep reinforcement learning methods to obtain an initial solution; deep reinforcement learning methods include: ① Establish a policy network model The policy network model consists of an encoder model and a decoder model. The encoder uses a one-dimensional convolutional layer as the embedding layer to map the data sequence composed of static and dynamic information of the system state into an encoded matrix. The decoder uses a gated recurrent unit (GRU) to store the information of the decoded sequence. At each decoding time step, based on the encoder output, the hidden state of the GRU, and the mask vector, the probability distribution of task selection is output based on the attention mechanism, and the decision for the current time step is sampled. ② Establish a value network model The value network uses a one-dimensional convolutional layer as the embedding layer to map the data sequence composed of static and dynamic information of the system state into an encoded matrix. After gathering the two encoded matrices, the output of the Critic network is obtained through two fully connected layers, which serves as an estimate of the current state value function. ③ Proximal policy optimization training policy network model Step 1: Initialize the policy network, value function network, and other algorithm parameters; Step 2: Under the current strategy, determine the observation task and collect trajectory data, including status, actions, rewards, and whether the trajectory is complete; Step 3: Use the value function network to calculate the advantage function for each state, which is the estimated difference between the future cumulative reward and the state value; Step 4: Calculate the importance sampling ratio of the old and new strategies; Step 5: Calculate the loss function and update the policy network; the loss function consists of two parts: advantage-weighted policy loss, which measures the improvement of the new policy relative to the old policy and maximizes the advantage function of experience; and a pruning term, which limits the magnitude of a single update and prevents excessive policy changes. Step 6: Update the policy network parameters using the stochastic gradient descent optimization algorithm to minimize the calculated loss function; Step 7: Repeat steps 2 through 6 until the predetermined number of training iterations or performance metrics are reached; S203: Generate a feasible initial solution S 0, and calculate its objective function value. f 0; will S 0 is set as the current solution. S C and best solution S B and the current objective function value fC ← f 0 and the optimal objective value f B ← f 0.
[0008] Optionally, S3 includes: destroying the current solution, that is, AA selects a destruction strategy based on the selection probability of the destruction strategy, removes some tasks from the current solution, and then performs reallocation and replanning to generate a neighborhood solution; S301: Sabotage strategies include: ① General strategy: First sabotage strategy 1: random sabotage; on each space platform, AA selects N observation tasks to remove with equal probability; Second Disruption Strategy 2: Observational Gain Destruction; In each space platform, AA selects and removes the N observation tasks with the lowest observational gains. ② Regarding satellites: Third Disruption Strategy 3: Observation Opportunity Disruption; Observation tasks with more observation opportunities are more flexible than those with fewer observation opportunities, and observation tasks with lower flexibility will be retained; AA selects the N observation tasks with the longest total feasible observation time for removal. Fourth Disruption Strategy 4: Observation Task Conflict Removal; On a space platform, the longer the overlap between the observation time windows of other observation tasks, the more likely an observation time conflict will occur; AA will remove the N observation tasks with the greatest degree of conflict; ③ For airships and drones: Fifth Disruption Strategy 5: Remove the N tasks in the task sequence that take the longest time from the end of the previous observation task to the end of this task; S302: Then, AA determines the allocation strategy based on the selection probability of the allocation strategy and redistributes the unplanned and removed observation tasks. After redistribution, SA replans to obtain a new current solution.
[0009] 5. The method for autonomous mission planning of collaborative Earth observation using heterogeneous space platforms according to claim 4, characterized in that: step S4 includes: S401: If the newly generated solution S * C (corresponding objective function value) f * C )and S B If the newly generated solution is better, then set... S B ← S * C ,S C ← S * C 、f B ← f * C and f C ← f * C If the newly generated solution is only better than the current solution, then S C ← S * C and f C ← f * C If none of the above apply, the Metropolis criterion will be used to determine whether to accept the newly generated solution. S402: Based on the results of S401, update the selection probabilities of the allocation and destruction strategies; set two parameters for each allocation and destruction strategy: weight and score; the initial values of the score and weight are 0; after each iteration, if the new solution becomes... S B If the chosen allocation strategy and destruction strategy score increase by 30, then the score increases by 30; if the new solution does not become... S B If the new solution is worse than the current solution but is accepted, the scores of the chosen allocation and destruction strategies increase by 10; if the new solution is worse than the current solution but is accepted, the scores of the chosen allocation and destruction strategies increase by 1; otherwise, the scores remain unchanged. ϕ In each iteration, the weight of a strategy is the ratio of its score at the current weight value to the sum of the scores of all similar strategies; the strategy selection probability is calculated based on the weight ratio among strategies; each ϕ Each iteration updates the probability of policy selection, the probability value being calculated from the original probability and this... ϕ The cumulative score for each iteration is determined, and after the update, the scores of the strategies are initialized.
[0010] Optionally, S5 includes: repeating S3 and S4. If the best solution is not updated after 10 iterations, the algorithm terminates and outputs the best solution, which is the solution to the autonomous mission planning problem for collaborative Earth observation by heterogeneous space platforms.
[0011] Secondly, this invention provides a device for autonomous mission planning of collaborative Earth observation on heterogeneous space platforms based on deep reinforcement learning, the device comprising: The first processing unit is used to establish an autonomous planning model for collaborative Earth observation of heterogeneous space platforms and generate initial solutions based on deep reinforcement learning. The second processing unit is used to generate neighborhood solutions based on deep reinforcement learning, and update the current solution, the optimal solution and the policy weights. It then generates neighborhood solutions again based on deep reinforcement learning, and updates the current solution, the optimal solution and the policy weights. If the best solution is not updated after a preset number of iterations, the algorithm terminates and outputs the current best solution.
[0012] Optionally, the first processing unit is further configured to define the autonomous planning problem for collaborative Earth observation by heterogeneous space platforms on a directed graph G = (N, A); N is the set of network nodes, A is the set of directed arcs, and the decision period is T; the set N = {D∪C} consists of virtual nodes. D = {0, σ} and observation task nodes C = {1, 2, …, n Composed of}, where {0} is the virtual starting point, { σ} represents the virtual endpoint obtained by copying {0}; the set of directed arcs is . ,in N 0 = {0}∪ C , N σ = C ∪{ σ}; For observation tasks i ∈ C Observational benefits are The observation duration is The observation time requirement is The cruising distance between the airship and the drone between the two observation missions is... and The cruising speeds of the airship and the drone are respectively and The Earth observation satellite set is Sa ={ s | s = 1, 2, …, n s Earth observation satellites i For observation tasks j The number of visible time windows is The set of visible time windows is , of which k The time window is Earth observation satellites s In observation mission i and j The attitude transition time between them is In other words, the arc { in the directed graph i , jThe length of the virtual point is determined, and the conversion time between the virtual point and the observation task node is 0. A multi-agent system is established to describe the autonomous task planning problem of collaborative Earth observation on heterogeneous space platforms. This multi-agent system includes a task allocation agent AA and several task planning agents SA. AA is responsible for allocating observation tasks to the task planning agents, and SA is responsible for autonomous task planning. For the task allocation problem, a task allocation model with the goal of maximizing the total observation revenue is established. For the autonomous task planning problem, a task planning model based on the Markov decision process is established. The goal of the task allocation model is to maximize the total revenue of all space platforms: Max (1) M A collection of space platforms N For the set of observation tasks, x ij For binary decision variables, when the observation task j Assigned to space platform i hour x ij = 1, otherwise 0; z ij As a binary variable, when the space platform i The observation mission was completed. j hour, z ij = 1, otherwise 0; p j For observation mission j The benefits; The task allocation process must satisfy a unique constraint, meaning that a task can only be assigned to one space platform: ST , (2) To ensure that the space platform can execute the assigned tasks, the feasibility of the space platform for the tasks is also considered during the task allocation process. For example, the space platform must meet the payload type requirements, resolution requirements, and the visible time window of the space platform for the observation tasks. Furthermore, preprocessing is carried out during the matching stage between the observation tasks and the space platform. After the AA assigns observation tasks to various space platforms, the SA performs autonomous mission planning, establishing a mission planning model based on a Markov decision process. The system starts from the initial state. S Starting from 0, SA is based on S 0 Perform an action A 0, received by the environment A After 0, calculate the action. A 0 reward R 1, and reach the next stage state. S 1; SA is based onS i Continuous output action A i This continues until a termination state is reached; each interaction between the SA and the environment is called a time step. (1) State: The set of states is the state space, which can be represented as follows: (3) S i It is the first t The state is defined for each time step, and each state contains a set of attributes that describe the state information at a specific time step. These attributes are divided into two parts: static and dynamic information. For satellites, static information includes the benefits and duration of the observation mission, while dynamic information includes the total remaining observation opportunities, the earliest time that the observation can be completed, the latest time that the observation can begin, and whether the mission has been allocated to a space platform and the execution time. For airships and UAVs, static information includes the benefits and duration of the observation mission, while dynamic information includes the total time consumed by the space platform from the current point to the completion of the observation mission and the remaining available time of the space platform. (2) Action: Action is the decision made by SA based on the current state, namely "selecting an observation task"; (3) Rewards: The rewards that SA receives from the environment are not always equal to the observation gains of the selected observation task, but rather the increment of the total observation gains after taking action; at each time step, the environment receives the decision action of SA and automatically feeds back the corresponding rewards and updated status. (4) Value function: The value function defines the average reward of long-term actions. A deep network is built as an estimate of the state value function.
[0013] Optionally, the second processing unit is further configured to, if a newly generated solution S * C (corresponding objective function value) f * C )and S B If the newly generated solution is better, then set... S B ← S * C , S C ← S * C 、f B ← f * C andf C ← f * C If the newly generated solution is only better than the current solution, then S C ← S * C and f C ← f * C If none of the above apply, the Metropolis criterion is used to determine whether to accept the newly generated solution; the root update includes the selection probabilities of the allocation and destruction strategies; two parameters, weight and score, are set for each allocation and destruction strategy; the score and weight are initialized to 0; after each iteration, if the new solution becomes... S B If the chosen allocation strategy and destruction strategy score increase by 30, then the score increases by 30; if the new solution does not become... S B If the new solution is worse than the current solution but is accepted, the scores of the chosen allocation and destruction strategies increase by 10; if the new solution is worse than the current solution but is accepted, the scores of the chosen allocation and destruction strategies increase by 1; otherwise, the scores remain unchanged. ϕ In each iteration, the weight of a strategy is the ratio of its score at the current weight value to the sum of the scores of all similar strategies; the strategy selection probability is calculated based on the weight ratio among strategies; each ϕ Each iteration updates the probability of policy selection, the probability value being calculated from the original probability and this... ϕ The cumulative score for each iteration is determined, and after the update, the scores of the strategies are initialized.
[0014] Thirdly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements any of the methods described above.
[0015] The beneficial effects of this invention are as follows: This invention addresses applications in the field of space-based Earth observation technology, providing a method and apparatus for autonomous mission planning of collaborative Earth observation on heterogeneous space platforms based on deep reinforcement learning. The method includes the following steps: establishing an autonomous mission planning model for collaborative Earth observation on heterogeneous space platforms; generating an initial solution based on the proposed deep reinforcement learning method and a random allocation strategy; generating neighborhood solutions based on the proposed deep reinforcement learning method, multiple allocation and disruption strategies; updating the current solution and the optimal solution, and adaptively updating the selection weights of the allocation and disruption strategies; continuing to generate neighborhood solutions based on the new strategy and weights; and outputting the best solution as the solution to the autonomous mission planning problem for collaborative Earth observation on heterogeneous space platforms when the algorithm reaches the termination condition. This invention has the positive effects of improving the utilization rate of time resources of heterogeneous space platforms and maximizing observation benefits. The above description is merely an overview of the technical solution of this invention. To better understand the technical means of this invention and to implement it according to the contents of the specification, and to make the above and other objects, features, and advantages of this invention more apparent, specific embodiments of this invention are described below. Attached Figure Description
[0016] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 This is a flowchart illustrating an autonomous mission planning method for collaborative Earth observation on heterogeneous space platforms based on deep reinforcement learning, provided in an embodiment of the present invention. Figure 2 This is a flowchart illustrating the implementation of an autonomous mission planning method for collaborative Earth observation on heterogeneous space platforms based on deep reinforcement learning, according to the present invention. Figure 3 This is a network structure diagram of the policy network and value network in this invention; Figure 4 This is a schematic diagram of the structure of an autonomous mission planning device for collaborative Earth observation based on heterogeneous space platforms using deep reinforcement learning, provided in an embodiment of the present invention. Detailed Implementation
[0017] like Figure 1 , Figure 2 As shown, the autonomous mission planning method for Earth observation on heterogeneous space platforms based on deep reinforcement learning provided by this invention mainly includes the following steps: S1: Establish an autonomous planning model for collaborative Earth observation using heterogeneous space platforms.
[0018] S101: The problem of autonomous planning for collaborative Earth observation by heterogeneous space platforms is defined on a directed graph G = (N,A). N is the set of network nodes, A is the set of directed arcs, and the decision period is T. The set N = {D∪C} consists of virtual nodes. D = {0, σ} and observation task nodes C = {1, 2, …, n Composed of}, where {0} is the virtual starting point, { σ} represents the virtual endpoint obtained by copying {0}. The set of directed arcs is... ,in N 0 = {0}∪ C , N σ = C ∪{ σ For observation tasks i ∈ C Observational benefits are The observation duration is The observation time requirement is The cruising distance between the airship and the drone between the two observation missions is... and The cruising speeds of the airship and the drone are respectively and The Earth observation satellite set is as follows: Sa = { s | s = 1, 2, …, n s Earth observation satellites i For observation tasks j The number of visible time windows is The set of visible time windows is , of which k The time window is Earth observation satellites s In observation mission i and j The attitude transition time between them is In other words, the arc { in the directed graph i , j The length of}, and the conversion time between virtual points and observation task nodes is 0.
[0019] S102: Establish a multi-agent system to describe the autonomous mission planning problem for collaborative Earth observation on heterogeneous space platforms. This multi-agent system includes one task allocation agent (AA) and several task planning agents (SA). AA is responsible for allocating observation tasks to the task planning agents, and SA are responsible for autonomous task planning. For the task allocation problem, establish a task allocation model with the objective of maximizing total observation revenue; for the autonomous task planning problem, establish a task planning model based on Markov decision processes.
[0020] The goal of the task allocation model is to maximize the total revenue of all space platforms: Max (1) M A collection of space platforms N For the set of observation tasks, x ij For binary decision variables, when the observation task j Assigned to space platform i hour x ij = 1, otherwise 0. z ij As a binary variable, when the space platform i The observation mission was completed. j hour, z ij = 1, otherwise 0. p j For observation mission j The benefits.
[0021] The task allocation process must satisfy a unique constraint, meaning that a task can only be assigned to one space platform: ST , (2) To ensure that the space platform can execute the assigned tasks, the feasibility of the space platform for the tasks is also considered during the task allocation process. For example, the space platform must meet the payload type requirements, resolution requirements, and the visible time window of the observation tasks. Furthermore, preprocessing is carried out during the matching stage between the observation tasks and the space platform.
[0022] After the AA assigns observation tasks to various space platforms, the SA performs autonomous mission planning, establishing a mission planning model based on a Markov decision process. The system starts from the initial state. S Starting from 0, SA is based on S 0 Perform an action A 0, received by the environment A After 0, calculate the action. A 0 reward R 1, and reach the next stage state.S 1. SA is based on S i Continuous output action A i This continues until a termination state is reached. Each interaction between the SA and the environment is called a time step.
[0023] (1) State: The set of states is the state space, which can be represented as follows: (3) S i It is the first t The state is defined for each time step, and each state contains a set of attributes describing the state information at a specific time step. These attributes are divided into two parts: static and dynamic information. For satellites, static information includes the observation mission's benefits and duration, while dynamic information includes the total remaining observation opportunities, the earliest time that observations can be completed, the latest time that observations can begin, and whether the mission has been allocated to a space platform and the execution time. For airships and UAVs, static information includes the observation mission's benefits and duration, while dynamic information includes the total time consumed by the space platform from the current point to the completion of the observation mission and the remaining available time of the space platform.
[0024] (2) Action: Action is the decision made by SA based on the current state, namely, "selecting an observation task".
[0025] (3) Rewards: The rewards that SA receives from the environment are not always equal to the observation gains of the selected observation task, but rather the increment of the total observation gains after taking action. At each time step, the environment receives the decision action from SA and automatically provides the corresponding rewards and updated status.
[0026] (4) Value function: The value function defines the average reward of long-term actions. A deep network is built as an estimate of the state value function.
[0027] S2: Initial solution generation based on deep reinforcement learning.
[0028] Generating the initial solution includes: initializing the weights of the assignment and violation strategies, initializing network parameters; performing initial task assignment using the task assignment strategy, and initial task planning based on deep reinforcement learning methods.
[0029] S201: The allocation strategies used for task assignment are divided into allocation strategies between heterogeneous space platforms, allocation strategies between homogeneous space platforms, and upper-level strategies. Strategies between heterogeneous space platforms include allocation strategies between satellites and other space platforms, and allocation strategies between aerostats and unmanned aerial vehicles (UAVs). Strategies between homogeneous space platforms are allocation strategies between satellites. Upper-level strategies define the conditions that must be met when invoking other allocation strategies, including that tasks are assigned to feasible space platforms, and that the total number of observation tasks assigned to a certain type or platform cannot exceed the upper limit.
[0030] ① Allocation between satellites and the other two types of space platforms: Allocation Strategy 1: Random Allocation. Without violating the upper-level strategy, AA will allocate observation tasks to satellites or other space platforms with equal probability.
[0031] Allocation Strategy 2: Taboo Strategy. Observation tasks are allocated according to a random strategy, but not to space platforms that received the task in the previous round but did not execute it, unless the upper-level strategy is violated.
[0032] ② Allocation between airships and drones: Allocation Strategy 3: Random Allocation. Without violating the upper-level strategy, AA allocates observation tasks to either aerostats or drones with equal probability.
[0033] Allocation Strategy 4: Assignment based on proximity. Without violating the upper-level strategy, AA will assign observation tasks to the nearest airship or drone.
[0034] ③ Satellite allocation: Allocation Strategy 5: Random Allocation. AA selects satellites to perform observation tasks with equal probability.
[0035] Allocation Strategy 6: Observation Opportunity Allocation. The longer the observation time, the more likely it is to complete the observation task. Therefore, in this strategy, AA will allocate the observation task to the satellite with the longest total feasible observation time for the task.
[0036] Allocation Strategy 7: Task Conflict Allocation. The longer the overlap between the observation windows of observation tasks, the greater the potential observation conflict between observation tasks on that satellite. Therefore, in this strategy, AA allocates observation tasks to the satellite with the lowest degree of conflict.
[0037] Allocation Strategy 8: Historical Experience Allocation. In multiple iterations, for each observation task assigned to each satellite, the mean of the objective function values from the most recent iterations is calculated. In this strategy, the AA assigns the observation task to the satellite that performed the task in the previous iterations and has the largest mean objective function value.
[0038] S202: AA selects assignment strategies 1, 3, and 5 to allocate tasks, while SA uses deep reinforcement learning methods for task planning to obtain an initial solution. Deep reinforcement learning methods include: ① Establish a policy network model The policy network model consists of an encoder model and a decoder model. The encoder uses one-dimensional convolutional layers as embedding layers to map the data sequence, composed of static and dynamic information of the system state, into an encoded matrix. The decoder uses gated recurrent units (GRUs) to store information of the decoded sequence. At each decoding time step, based on the encoder output, the hidden state of the GRU, and the mask vector, the probability distribution of task selection is output using an attention mechanism, and the decision for the current time step is sampled.
[0039] ② Establish a value network model The value network uses a one-dimensional convolutional layer as the embedding layer, which maps the data sequence composed of static and dynamic information of the system state into an encoded matrix. After gathering the two encoded matrices, the output of the Critic network is obtained through two fully connected layers, which serves as an estimate of the current state value function.
[0040] ③ Proximal policy optimization training policy network model Step 1: Initialize the policy network, value function network, and other algorithm parameters.
[0041] Step 2: Under the current strategy, determine the observation task and collect trajectory data, including information such as status, action, reward, and whether the trajectory is complete.
[0042] Step 3: Use the value function network to calculate the advantage function for each state, which is the estimated difference between the future cumulative return and the state value.
[0043] Step 4: Calculate the importance sampling ratio of the old and new strategies.
[0044] Step 5: Calculate the loss function and update the policy network. The loss function consists of two parts: advantage-weighted policy loss, which measures the improvement of the new policy relative to the old policy and maximizes the advantage function of experience; and a pruning term, which limits the magnitude of a single update and prevents excessive policy changes.
[0045] Step 6: Update the policy network parameters using the stochastic gradient descent optimization algorithm to minimize the calculated loss function.
[0046] Step 7: Repeat steps 2 through 6 until the predetermined number of training iterations or performance metrics are reached.
[0047] In this embodiment of the invention, the network structures of the policy network and the value network are specifically as follows: Figure 3 As shown.
[0048] S203: Generate a feasible initial solution S 0, and calculate its objective function value. f 0. Will S 0 is set as the current solution. S C and best solution S B and the current objective function value f C ← f 0 and the optimal objective value f B ← f 0.
[0049] S3: Neighborhood solution generation based on deep reinforcement learning.
[0050] Disrupt the current solution, that is, AA selects a disruption strategy based on the probability of the disruption strategy selection, removes some tasks from the current solution, and then performs redistribution and replanning to generate neighborhood solutions.
[0051] S301: Sabotage strategies include: ① General strategy: Sabotage Strategy 1: Random Sabotage. On each space platform, AA selects N observation missions to remove with equal probability.
[0052] Disruption Strategy 2: Observational Gain Destruction. On each space platform, the AA selects and removes the N observation tasks with the lowest observational gains.
[0053] ② Regarding satellites: Disruption Strategy 3: Observation Opportunity Destruction. Observation tasks with more observation opportunities are more flexible than those with fewer opportunities, and less flexible observation tasks will be retained. AA selects the N observation tasks with the longest total feasible observation time for removal.
[0054] Disruption Strategy 4: Observation Task Conflict Removal. On a space platform, the longer the overlap between the observation time windows of other observation tasks, the more likely an observation time conflict will occur. AA removes the N observation tasks with the highest degree of conflict.
[0055] ③ For airships and drones: Disruption Strategy 5: Remove the N tasks from the task sequence that take the longest time from the end of the previous observation task to the end of this task.
[0056] S302: Then, AA determines the allocation strategy based on the selection probability of the allocation strategy and redistributes the unplanned and removed observation tasks. After redistribution, SA replans to obtain a new current solution.
[0057] S4: Update the current solution, the optimal solution, and the policy weights.
[0058] S401: If the newly generated solution S * C (corresponding objective function value) f * C )and S B If the newly generated solution is better, then set... S B ← S * C , S C ← S * C 、f B ← f * C and f C ← f * C If the newly generated solution is only better than the current solution, then S C ← S * C and f C ← f * C If none of the above apply, the Metropolis criterion will be used to determine whether to accept the newly generated solution.
[0059] S402: Based on the results of S401, update the selection probabilities of the allocation and destruction strategies. Set two parameters for each allocation and destruction strategy: weight and score. The score and weight are initially set to 0. After each iteration, if the new solution becomes... S B If the chosen allocation strategy and destruction strategy score increase by 30, then the score increases by 30; if the new solution does not become... S B If the new solution is worse than the current solution, the scores of the chosen allocation and destruction strategies increase by 10; if the new solution is worse than the current solution but is accepted, the scores of the chosen allocation and destruction strategies increase by 1; otherwise, the scores remain unchanged. ϕIn each iteration, the weight of a strategy is the ratio of its score at the current weight value to the sum of the scores of all similar strategies. The strategy selection probability is calculated based on the weight ratios among the strategies. ϕ Each iteration updates the probability of policy selection, the probability value being calculated from the original probability and this... ϕ The cumulative score for each iteration is determined, and after the update, the scores of the strategies are initialized.
[0060] S5: Repeat S3 and S4, for 10 iterations. S B If no update is found, the algorithm terminates and outputs the result. S B .
[0061] S6: The autonomous task planning method for heterogeneous space platform Earth observation based on deep reinforcement learning in this invention is compared with the independent planning results of three types of space platforms to verify the effectiveness of the method in this invention. In independent planning, the random allocation strategy in this invention is used to allocate tasks, and task planning is performed based on deep reinforcement learning. The collaborative process between multiple platforms is no longer performed. The number of observation tasks for each type of platform is counted, and the independent planning results are output.
[0062] This invention addresses the problem of autonomous mission planning for collaborative Earth observation using heterogeneous space platforms. Multiple sets of computational examples of varying scales are constructed and compared with independent planning for the three types of space platforms to verify the effectiveness of the proposed method. The observation missions are distributed across the region 120°E–123°E, 30°N–33°N. The scheduling period is 24 hours (October 1, 2023, 00:00:00 to October 1, 2023, 24:00:00). Mission locations are randomly generated within the region. The scales of the three types of space platforms (satellite / airship / UAV) in the test cases are (5 / 20 / 30), (10 / 40 / 60), (10 / 100 / 100), (15 / 100 / 100), and (20 / 100 / 100). All observation tasks are standard point target tasks, with task sizes of 500, 1000, 1500, 1750, and 2000 targets, offering the same observation benefits. Observation durations are [10s, 30s], and there are specific time requirements. A random value between the start time and the end time. The earliest start time is added to 4 hours, but not exceeding the end time. The algorithm is coded using Python 3.9 and runs on a computer with a Windows 11 64-bit operating system, an Intel Core i9-13900K@3.00 GHz processor, 32 MB of RAM, and an Nvidia RTX A6000 48GB CPU. Earth observation satellite orbital parameters are shown in Table 1, and parameters related to the aerostat and UAV are shown in Table 2.
[0063] Table 1. Orbital parameters of Earth observation satellites Semi-major axis (m) Track inclination angle (°) Right ascension of the ascending node (°) Eccentricity Perigee argument (°) Angle of approach (°) 7569345 100.74 83.2388 0.0025382 10.9656 90.17 7467869 63.40 338.3441 0.0278024 4.928 129.66 7580228 100.63 39.861 0.0005816 260.4264 207.29 7467799 63.40 116.9629 0.018683 4.7392 318.22 7011361 98.30 84.095 0.001895 218.1331 250.30 7620677 100.28 190.5985 0.0021603 30.9156 234.61 7580045 100.05 283.1983 0.0009035 207.7205 245.68 7075827 98.32 218.5334 0.0000559 290.4737 117.31 7011677 97.86 206.4439 0.0017091 53 323.38 7467884 63.39 48.286 0.0302859 10.0047 24.61 7578175 100.09 290.073 0.000815 343.9494 86.55 7467835 63.41 281.7681 0.0085931 355.9612 231.60 6843529 97.67 240.9719 0.0002185 89.2653 83.12 6996702 97.91 296.9113 0.0001947 87.1681 319.38 7467687 63.38 35.4537 0.0452227 14.8066 346.571 7000790 97.96 287.6775 0.0001726 87.1112 4.8288 7013670 98.42 76.1724 0.0037755 168.5903 191.6171 6873477 97.61 65.5052 0.0001482 94.5257 265.6149 7467768 63.39 121.457 0.033844 7.8224 352.7896 7073152 98.16 163.625 0.0000692 52.1286 279.94 Table 2 Parameters of Airships and UAVs parameter airship drones Base station location (121.3°E, 32.3°N) (121.5°E, 32.1°N) Maximum battery life 4h 2h movement speed 70km / h 150km / h Coverage radius 5.8km - Table 3 Experimental Results Table 3 presents the experimental results of the proposed method and independent planning for three types of space platforms. It includes the total number of observation tasks for each platform, the number of observation tasks for each type of platform, the improvement rate of the proposed method compared to independent planning, and the running time of the proposed algorithm. The results in the table represent the most conservative results from five runs, i.e., the results with the smallest average total number of observation tasks across the five runs. It can be observed that the proposed method outperforms the independent planning method in all examples, with the highest improvement rate in the number of completed tasks reaching 82.39%, and the algorithm's running time is less than 1.5 minutes, verifying the effectiveness of the proposed method.
[0064] Meanwhile, this invention also provides a device for autonomous mission planning of collaborative Earth observation on heterogeneous space platforms based on deep reinforcement learning, see [link to related documentation]. Figure 4 The device includes: The first processing unit is used to establish an autonomous planning model for collaborative Earth observation of heterogeneous space platforms and generate initial solutions based on deep reinforcement learning. The second processing unit is used to generate neighborhood solutions based on deep reinforcement learning, and update the current solution, the optimal solution and the policy weights. It then generates neighborhood solutions again based on deep reinforcement learning, and updates the current solution, the optimal solution and the policy weights. If the best solution is not updated after a preset number of iterations, the algorithm terminates and outputs the current best solution.
[0065] Furthermore, in this embodiment of the invention, the first processing unit is also used to define the autonomous planning problem of collaborative Earth observation by heterogeneous space platforms on a directed graph G = (N, A); N is a set of network nodes, A is a set of directed arcs, and the decision period is T; the set N = {D∪C} consists of virtual nodes. D = {0, σ} and observation task nodes C = {1, 2, …, n Composed of}, where {0} is the virtual starting point, { σ} represents the virtual endpoint obtained by copying {0}; the set of directed arcs is . ,in N 0 = {0}∪ C , N σ = C ∪{ σ}; For observation tasks i ∈ C Observational benefits are The observation duration is The observation time requirement is The cruising distance between the airship and the drone between the two observation missions is... and The cruising speeds of the airship and the drone are respectively and The Earth observation satellite set is Sa ={ s | s = 1, 2, …, n s Earth observation satellites i For observation tasks j The number of visible time windows is The set of visible time windows is , of which k The time window is Earth observation satellites s In observation mission i and j The attitude transition time between them is In other words, the arc { in the directed graph i , j The length of the virtual point is determined, and the conversion time between the virtual point and the observation task node is 0. A multi-agent system is established to describe the autonomous task planning problem of collaborative Earth observation on heterogeneous space platforms. This multi-agent system includes a task allocation agent AA and several task planning agents SA. AA is responsible for allocating observation tasks to the task planning agents, and SA is responsible for autonomous task planning. For the task allocation problem, a task allocation model with the goal of maximizing the total observation revenue is established. For the autonomous task planning problem, a task planning model based on the Markov decision process is established. The goal of the task allocation model is to maximize the total revenue of all space platforms: Max (1) M A collection of space platforms N For the set of observation tasks, x ij For binary decision variables, when the observation task j Assigned to space platform i hour x ij = 1, otherwise 0; z ij As a binary variable, when the space platform i The observation mission was completed. jhour, z ij = 1, otherwise 0; p j For observation mission j The benefits; The task allocation process must satisfy a unique constraint, meaning that a task can only be assigned to one space platform: ST , (2) To ensure that the space platform can execute the assigned tasks, the feasibility of the space platform for the tasks is also considered during the task allocation process. For example, the space platform must meet the payload type requirements, resolution requirements, and the visible time window of the space platform for the observation tasks. Furthermore, preprocessing is carried out during the matching stage between the observation tasks and the space platform. After the AA assigns observation tasks to various space platforms, the SA performs autonomous mission planning, establishing a mission planning model based on a Markov decision process. The system starts from the initial state. S Starting from 0, SA is based on S 0 Perform an action A 0, received by the environment A After 0, calculate the action. A 0 reward R 1, and reach the next stage state. S 1; SA is based on S i Continuous output action A i This continues until a termination state is reached; each interaction between the SA and the environment is called a time step. (1) State: The set of states is the state space, which can be represented as follows: (3) S i It is the first t The state is defined for each time step, and each state contains a set of attributes that describe the state information at a specific time step. These attributes are divided into two parts: static and dynamic information. For satellites, static information includes the benefits and duration of the observation mission, while dynamic information includes the total remaining observation opportunities, the earliest time that the observation can be completed, the latest time that the observation can begin, and whether the mission has been allocated to a space platform and the execution time. For airships and UAVs, static information includes the benefits and duration of the observation mission, while dynamic information includes the total time consumed by the space platform from the current point to the completion of the observation mission and the remaining available time of the space platform. (2) Action: Action is the decision made by SA based on the current state, namely "selecting an observation task"; (3) Rewards: The rewards that SA receives from the environment are not always equal to the observation gains of the selected observation task, but rather the increment of the total observation gains after taking action; at each time step, the environment receives the decision action of SA and automatically feeds back the corresponding rewards and updated status. (4) Value function: The value function defines the average reward of long-term actions. A deep network is built as an estimate of the state value function.
[0066] Furthermore, the second processing unit is also used to, if the newly generated solution S * C (corresponding objective function value) f * C )and S B If the newly generated solution is better, then set... S B ← S * C , S C ← S * C 、f B ← f * C and f C ← f * C If the newly generated solution is only better than the current solution, then S C ← S * C and f C ← f * C If none of the above apply, the Metropolis criterion is used to determine whether to accept the newly generated solution; the root update includes the selection probabilities of the allocation and destruction strategies; two parameters, weight and score, are set for each allocation and destruction strategy; the score and weight are initialized to 0; after each iteration, if the new solution becomes... S B If the chosen allocation strategy and destruction strategy score increase by 30, then the score increases by 30; if the new solution does not become... S B If the new solution is worse than the current solution but is accepted, the scores of the chosen allocation and destruction strategies increase by 10; if the new solution is worse than the current solution but is accepted, the scores of the chosen allocation and destruction strategies increase by 1; otherwise, the scores remain unchanged. ϕ In each iteration, the weight of a strategy is the ratio of its score at the current weight value to the sum of the scores of all similar strategies; the strategy selection probability is calculated based on the weight ratio among strategies; each ϕ Each iteration updates the probability of policy selection, the probability value being calculated from the original probability and this... ϕ The cumulative score for each iteration is determined, and after the update, the scores of the strategies are initialized.
[0067] In addition, embodiments of the present invention also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the methods described above.
[0068] The relevant content of the device embodiment and storage medium embodiment of the present invention can be understood by referring to the method embodiment of the present invention, and will not be discussed in detail here.
[0069] Although preferred embodiments of the invention have been disclosed for illustrative purposes, those skilled in the art will recognize that various modifications, additions, and substitutions are possible, and therefore the scope of the invention should not be limited to the embodiments described above.
Claims
1. A method for autonomous mission planning of collaborative Earth observation on heterogeneous space platforms based on deep reinforcement learning, characterized in that, include: S1. Establish an autonomous planning model for collaborative Earth observation using heterogeneous space platforms; S2. Generate initial solutions based on deep reinforcement learning; S3. Generating neighborhood solutions based on deep reinforcement learning; S4. Update the current solution, the optimal solution, and the policy weights; S5. Repeat S3 and S4. If the best solution has not been updated after the preset number of iterations, the algorithm terminates and outputs the current best solution.
2. The method for autonomous mission planning of collaborative Earth observation using heterogeneous space platforms according to claim 1, characterized in that: S1 further includes: S101: The problem of autonomous planning for collaborative Earth observation by heterogeneous space platforms is defined on a directed graph G = (N, A); N is the set of network nodes, A is the set of directed arcs, and the decision period is T; the set N = {D∪C} consists of virtual nodes. D = {0, σ } and observation task nodes C = {1, 2, …, n Composed of}, where {0} is the virtual starting point, { σ } represents the virtual endpoint obtained by copying {0}; the set of directed arcs is . ,in N 0 = {0}∪ C , N σ = C ∪{ σ }; For observation tasks i ∈ C Observational benefits are The observation duration is The observation time requirement is The cruising distance between the airship and the drone between the two observation missions is... and The cruising speeds of the airship and the drone are respectively and The Earth observation satellite set is Sa = { s | s = 1, 2, …, n s Earth observation satellites i For observation tasks j The number of visible time windows is The set of visible time windows is , of which k The time window is Earth observation satellites s In observation mission i and j The attitude transition time between them is In other words, the arc { in the directed graph i , j The length of}, and the conversion time between virtual points and observation task nodes is 0; S102: Establish a multi-agent system to describe the autonomous task planning problem of collaborative Earth observation on heterogeneous space platforms. This multi-agent system includes a task allocation agent AA and several task planning agents SA. AA is used to allocate observation tasks to the task planning agents, and SA is used to perform autonomous task planning. For the task allocation problem, establish a task allocation model with the goal of maximizing the total observation benefits. For the autonomous task planning problem, establish a task planning model based on Markov decision process. The goal of the task allocation model is to maximize the total revenue of all space platforms: Max (1) M A collection of space platforms N For the set of observation tasks, x ij For binary decision variables, when the observation task j Assigned to space platform i hour x ij = 1, otherwise 0; z ij As a binary variable, when the space platform i The observation mission was completed. j hour, z ij = 1, otherwise 0; p j For observation mission j The benefits; The task allocation process satisfies a uniqueness constraint, meaning that a task can only be assigned to one space platform: S.T. , (2) The feasibility of the space platform for the mission was considered during the mission allocation process. The space platform met the mission's payload type requirements, resolution requirements, and the visible time window of the observation mission. Furthermore, preprocessing was performed during the mission and space platform matching phase to ensure that the space platform could execute the assigned mission. After the AA assigns observation tasks to various space platforms, the SA performs autonomous mission planning, establishing a mission planning model based on a Markov decision process. The system starts from the initial state. S Starting from 0, SA is based on S 0 Perform an action A 0, received by the environment A After 0, calculate the action. A 0 reward R 1, and reach the next stage state. S 1; SA is based on S i Continuous output action A i This continues until a termination state is reached; each interaction between the SA and the environment is called a time step. (1) States: The set of states forms the state space: (3) S i It is the first t The state is defined for each time step, and each state contains a set of attributes that describe the state information at a specific time step. These attributes are divided into two parts: static and dynamic information. For satellites, static information includes the benefits and duration of the observation mission, while dynamic information includes the total remaining observation opportunities, the earliest time that the observation can be completed, the latest time that the observation can begin, and whether the mission has been allocated to a space platform and the execution time. For airships and UAVs, static information includes the benefits and duration of the observation mission, while dynamic information includes the total time consumed by the space platform from the current point to the completion of the observation mission and the remaining available time of the space platform. (2) Action: Action is the decision made by SA based on the current state, that is, to select an observation task; (3) Rewards: The rewards that SA receives from the environment are not always equal to the observation gains of the selected observation task, but rather the increment of the total observation gains after taking action; at each time step, the environment receives the decision action of SA and automatically feeds back the corresponding rewards and updated status. (4) Value function: The value function defines the average reward of long-term actions. A deep network is built as an estimate of the state value function.
3. The method for autonomous mission planning of heterogeneous space platform collaborative Earth observation based on deep reinforcement learning according to claim 2, characterized in that: S2 includes: Generating an initial solution includes: initializing the weights of the assignment and violation strategies, initializing network parameters; performing initial task assignment using the task assignment strategy, and initial task planning based on deep reinforcement learning methods; S201: The allocation strategies used for task allocation are divided into allocation strategies between heterogeneous space platforms and between homogeneous space platforms, as well as upper-level strategies. The strategies between heterogeneous space platforms include allocation strategies between satellites and other space platforms, and allocation strategies between airships and UAVs. The strategies between homogeneous space platforms are allocation strategies between satellites. The upper-level strategies are the conditions that must be met when calling other allocation strategies, including that tasks are allocated to feasible space platforms, and the total number of observation tasks allocated to a certain type or a certain space platform cannot exceed the upper limit. ① Allocation between satellites and the other two types of space platforms: First allocation strategy 1: random allocation; without violating the upper-level strategy, AA will allocate observation tasks to satellites or other space platforms with equal probability; Second allocation strategy 2: Taboo strategy; observation tasks are allocated according to a random strategy, but not to space platforms that received the task in the previous round but did not execute it, unless the upper-level strategy is violated. ② Allocation between airships and drones: Third allocation strategy 3: random allocation; without violating the upper-level strategy, AA will allocate observation tasks to airships or drones with equal probability; Fourth allocation strategy 4: allocation based on proximity; without violating the upper-level strategy, AA will allocate the observation task to the nearest airship or UAV; ③ Satellite allocation: Fifth allocation strategy 5: random allocation; AA means that satellites are selected to perform observation tasks with equal probability; Sixth allocation strategy 6: Observation opportunity allocation; the longer the visibility time, the more likely it is to complete the observation task. Therefore, in this strategy, AA will allocate the observation task to the satellite with the longest total feasible observation time for the task. Seventh Allocation Strategy 7: Task Conflict Allocation; The longer the overlap between the observation time windows of observation tasks, the greater the possible observation conflict between observation tasks on that satellite; Therefore, in this strategy, AA will allocate observation tasks to the satellite with the lowest degree of conflict. Eighth allocation strategy 8: Historical experience allocation; In multiple iterations, for each observation task assigned to each satellite, the mean of the objective function values of the most recent iterations is calculated. In this strategy, AA will assign the observation task to the satellite that performed the task in the previous iterations and has the largest mean objective value. S202: AA selects the first allocation strategy 1, the third allocation strategy 3, and the fifth allocation strategy 5 to allocate tasks. SA performs task planning based on deep reinforcement learning methods to obtain an initial solution; deep reinforcement learning methods include: ① Establish a policy network model The policy network model consists of an encoder model and a decoder model. The encoder uses a one-dimensional convolutional layer as the embedding layer to map the data sequence composed of static and dynamic information of the system state into an encoded matrix. The decoder uses a gated recurrent unit (GRU) to store the information of the decoded sequence. At each decoding time step, based on the encoder output, the hidden state of the GRU, and the mask vector, the probability distribution of task selection is output based on the attention mechanism, and the decision for the current time step is sampled. ② Establish a value network model The value network uses a one-dimensional convolutional layer as the embedding layer to map the data sequence composed of static and dynamic information of the system state into an encoded matrix. After gathering the two encoded matrices, the output of the Critic network is obtained through two fully connected layers, which serves as an estimate of the current state value function. ③ Proximal policy optimization training policy network model Step 1: Initialize the policy network, value function network, and other algorithm parameters; Step 2: Under the current strategy, determine the observation task and collect trajectory data, including status, actions, rewards, and whether the trajectory is complete; Step 3: Use the value function network to calculate the advantage function for each state, which is the estimated difference between the future cumulative reward and the state value; Step 4: Calculate the importance sampling ratio of the old and new strategies; Step 5: Calculate the loss function and update the policy network; the loss function consists of two parts: advantage-weighted policy loss, which measures the improvement of the new policy relative to the old policy and maximizes the advantage function of experience; and a pruning term, which limits the magnitude of a single update and prevents excessive policy changes. Step 6: Update the policy network parameters using the stochastic gradient descent optimization algorithm to minimize the calculated loss function; Step 7: Repeat steps 2 through 6 until the predetermined number of training iterations or performance metrics are reached; S203: Generate a feasible initial solution S 0, and calculate its objective function value. f 0; will S 0 is set as the current solution. S C and best solution S B and the current objective function value f C ← f 0 and the optimal objective value f B ← f 0.
4. The method for autonomous mission planning of heterogeneous space platform collaborative Earth observation based on deep reinforcement learning according to claim 3, characterized in that: S3 includes: Disrupt the current solution, that is, AA selects a disruption strategy based on the probability of the disruption strategy selection, removes some tasks from the current solution, and then performs reallocation and replanning to generate neighborhood solutions; S301: Sabotage strategies include: ① General strategy: First sabotage strategy 1: random sabotage; on each space platform, AA selects N observation tasks to remove with equal probability; Second Disruption Strategy 2: Observational Gain Destruction; In each space platform, AA selects and removes the N observation tasks with the lowest observational gains. ② Regarding satellites: Third Disruption Strategy 3: Observation Opportunity Disruption; Observation tasks with more observation opportunities are more flexible than those with fewer observation opportunities, and observation tasks with lower flexibility will be retained; AA selects the N observation tasks with the longest total feasible observation time for removal. Fourth Disruption Strategy 4: Observation Task Conflict Removal; On a space platform, the longer the overlap between the observation time windows of other observation tasks, the more likely an observation time conflict will occur; AA will remove the N observation tasks with the greatest degree of conflict; ③ For airships and drones: Fifth Disruption Strategy 5: Remove the N tasks in the task sequence that take the longest time from the end of the previous observation task to the end of this task; S302: Then, AA determines the allocation strategy based on the selection probability of the allocation strategy and redistributes the unplanned and removed observation tasks. After redistribution, SA replans to obtain a new current solution.
5. The method for autonomous mission planning of collaborative Earth observation using heterogeneous space platforms according to claim 4, characterized in that: S4 includes: S401: If the newly generated solution S * C (corresponding objective function value) f * C )and S B If the newly generated solution is better, then set... S B ← S * C , S C ← S * C 、f B ← f * C and f C ← f * C If the newly generated solution is only better than the current solution, then S C ← S * C and f C ← f * C If none of the above apply, the Metropolis criterion will be used to determine whether to accept the newly generated solution. S402: Based on the results of S401, update the selection probabilities of the allocation and destruction strategies; set two parameters for each allocation and destruction strategy: weight and score; the initial values of the score and weight are 0; after each iteration, if the new solution becomes... S B If the chosen allocation strategy and destruction strategy score increase by 30, then the score increases by 30; if the new solution does not become... S B If the new solution is worse than the current solution but is accepted, the scores of the chosen allocation and destruction strategies increase by 10; if the new solution is worse than the current solution but is accepted, the scores of the chosen allocation and destruction strategies increase by 1; otherwise, the scores remain unchanged. ϕ In each iteration, the weight of a strategy is the ratio of its score at the current weight value to the sum of the scores of all similar strategies; the strategy selection probability is calculated based on the weight ratio among strategies; each ϕ Each iteration updates the probability of policy selection, the probability value being calculated from the original probability and this... ϕ The cumulative score for each iteration is determined, and after the update, the scores of the strategies are initialized.
6. The method for autonomous mission planning of collaborative Earth observation using heterogeneous space platforms according to claim 5, characterized in that: S5 includes: Repeat S3 and S4. If the best solution is not updated after 10 iterations, the algorithm terminates and outputs the best solution, which is the solution to the autonomous mission planning problem for collaborative Earth observation by heterogeneous space platforms.
7. A heterogeneous space platform collaborative Earth observation autonomous mission planning device based on deep reinforcement learning, characterized in that, The device includes: The first processing unit is used to establish an autonomous planning model for collaborative Earth observation of heterogeneous space platforms and generate initial solutions based on deep reinforcement learning. The second processing unit is used to generate neighborhood solutions based on deep reinforcement learning, and update the current solution, the optimal solution and the policy weights. It then generates neighborhood solutions again based on deep reinforcement learning, and updates the current solution, the optimal solution and the policy weights. If the best solution is not updated after a preset number of iterations, the algorithm terminates and outputs the current best solution.
8. The heterogeneous space platform collaborative Earth observation autonomous mission planning device according to claim 7, characterized in that: The first processing unit is further configured to define the autonomous planning problem for collaborative Earth observation by heterogeneous space platforms on a directed graph G = (N, A); N is the set of network nodes, A is the set of directed arcs, and the decision period is T; the set N = {D∪C} consists of virtual nodes. D = {0, σ } and observation task nodes C = {1, 2, …, n Composed of}, where {0} is the virtual starting point, { σ } represents the virtual endpoint obtained by copying {0}; the set of directed arcs is . ,in N 0 = {0}∪ C , N σ = C ∪{ σ }; For observation tasks i ∈ C Observational benefits are The observation duration is The observation time requirement is The cruising distance between the airship and the drone between the two observation missions is... and The cruising speeds of the airship and the drone are respectively and The Earth observation satellite set is Sa = { s | s = 1, 2, …, n s Earth observation satellites i For observation tasks j The number of visible time windows is The set of visible time windows is , of which k The time window is Earth observation satellites s In observation mission i and j The attitude transition time between them is In other words, the arc { in the directed graph i , j The length of} and the conversion time between virtual points and observation task nodes is 0; establish a multi-agent system to describe the autonomous task planning problem of collaborative Earth observation of heterogeneous space platforms. The multi-agent system includes a task allocation agent AA and several task planning agents SA. AA is responsible for allocating observation tasks to task planning agents, and SA is responsible for autonomous task planning. For the task allocation problem, a task allocation model with the goal of maximizing total observed benefits is established; for the autonomous task planning problem, a task planning model based on Markov decision process is established. The goal of the task allocation model is to maximize the total revenue of all space platforms: Max (1) M A collection of space platforms N For the set of observation tasks, x ij For binary decision variables, when the observation task j Assigned to space platform i hour x ij = 1, otherwise 0; z ij As a binary variable, when the space platform i The observation mission was completed. j hour, z ij = 1, otherwise 0; p j For observation mission j The benefits; The task allocation process must satisfy a unique constraint, meaning that a task can only be assigned to one space platform: S.T. , (2) To ensure that the space platform can execute the assigned tasks, the feasibility of the space platform for the tasks is considered during the task allocation process. The space platform meets the payload type requirements, resolution requirements, and the visible time window of the space platform for the observation tasks. Furthermore, preprocessing is performed during the matching stage between the observation tasks and the space platform. After the AA assigns observation tasks to various space platforms, the SA performs autonomous mission planning, establishing a mission planning model based on a Markov decision process. The system starts from the initial state. S Starting from 0, SA is based on S 0 Perform an action A 0, received by the environment A After 0, calculate the action. A 0 reward R 1, and reach the next stage state. S 1; SA is based on S i Continuous output action A i This continues until a termination state is reached; each interaction between the SA and the environment is called a time step. (1) States: The set of states forms the state space: (3) S i It is the first t The state is defined for each time step, and each state contains a set of attributes that describe the state information at a specific time step. These attributes are divided into two parts: static and dynamic information. For satellites, static information includes the benefits and duration of the observation mission, while dynamic information includes the total remaining observation opportunities, the earliest time that the observation can be completed, the latest time that the observation can begin, and whether the mission has been allocated to a space platform and the execution time. For airships and UAVs, static information includes the benefits and duration of the observation mission, while dynamic information includes the total time consumed by the space platform from the current point to the completion of the observation mission and the remaining available time of the space platform. (2) Action: Action is the decision made by SA based on the current state, namely "selecting an observation task"; (3) Rewards: The rewards that SA receives from the environment are not always equal to the observation gains of the selected observation task, but rather the increment of the total observation gains after taking action; at each time step, the environment receives the decision action of SA and automatically feeds back the corresponding rewards and updated status. (4) Value function: The value function defines the average reward of long-term actions. A deep network is built as an estimate of the state value function.
9. The heterogeneous space platform collaborative Earth observation autonomous mission planning device according to claim 7, characterized in that: The second processing unit is further configured to, if a newly generated solution S * C (corresponding objective function value) f * C )and S B If the newly generated solution is better, then set... S B ← S * C , S C ← S * C 、f B ← f * C and f C ← f * C If the newly generated solution is only better than the current solution, then S C ← S * C and f C ← f * C If none of the above apply, the Metropolis criterion is used to determine whether to accept the newly generated solution; the root update includes the selection probabilities of the allocation and destruction strategies; two parameters, weight and score, are set for each allocation and destruction strategy; the score and weight are initially set to 0. After each iteration, if the new solution becomes S B If the chosen allocation strategy and destruction strategy score increase by 30, then the score increases by 30; if the new solution does not become... S B If the new solution is worse than the current solution, the scores of the selected allocation strategy and destruction strategy increase by 10; if the new solution is worse than the current solution but is accepted, the scores of the selected allocation strategy and destruction strategy increase by 1. Otherwise, the score remains unchanged; after ϕ In each iteration, the weight of a strategy is the ratio of its score at the current weight value to the sum of the scores of all similar strategies; the strategy selection probability is calculated based on the weight ratio among strategies. Every ϕ Each iteration updates the probability of policy selection, the probability value being calculated from the original probability and this... ϕ The cumulative score for each iteration is determined, and after the update, the scores of the strategies are initialized.
10. A computer-readable storage medium storing a computer program that, when executed by a processor, implements the method of any one of claims 1-6.