Satellite resource adaptive allocation method and system based on deep intensity learning

Through a deep intensity learning method, a task planning model and Markov decision-making process are constructed, combined with Transformer and Actor-Critic algorithms to optimize satellite resource allocation, solve the efficient coordination problem of multiple satellites under limited resources, and realize independent decision-making for dynamic task scheduling.

CN120430573APending Publication Date: 2025-08-05CHINA UNIV OF GEOSCIENCES (WUHAN)
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510565345.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

It is difficult for the existing technology to achieve efficient and coordinated implementation of multiple satellites under limited satellite resource conditions, and traditional resource allocation methods are difficult to cope with dynamically changing observation requirements and environmental conditions, resulting in lagging system response.

Method used

Using a deep intensity learning method, a task planning problem model is constructed, ground goals are decomposed into multiple optional observation activities, Markov decision-making process model is established, and satellite resource allocation strategies are optimized through Transformer's policy network combined with Actor-Critic algorithm, supplemented by heuristic algorithm to determine the task execution order and time period.

Benefits of technology

It realizes efficient coordination of multiple satellites under the conditions of limited satellite resources, can dynamically adjust resource allocation strategies, meet the needs of multiple observation tasks, and improves the independent decision-making ability of satellite mission planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430573A_ABST
    Figure CN120430573A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of aerospace engineering application, and particularly discloses a satellite resource adaptive allocation method and system based on deep intensity learning. The method comprises the steps that a task planning problem model is constructed, ground targets are decomposed into a plurality of optional observation activities based on the model, each ground target corresponds to an independent sub-task in a visible time window of the ground target, and the sub-tasks are defined as meta-tasks; based on the meta-task set, constructing a Markov decision process model; on the basis of a Markov decision process model, a strategy network model based on Transform is adopted, an Actor-Critic algorithm is combined to iteratively optimize a satellite resource allocation strategy for network model parameters, and then a heuristic algorithm is adopted to determine the specific execution sequence and execution time period of each satellite resource provider task. According to the invention, under the condition of limited satellite resources, efficient cooperation of multiple satellites can be realized so as to complete various observation tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of aerospace engineering applications, and more specifically, relates to a method and system for adaptive allocation of satellite resources based on deep intensity learning. Background Art

[0002] With the rapid development of aerospace technology, imaging satellites have become an indispensable observation tool in remote sensing. Relying on high-precision imaging equipment, imaging satellites are capable of high-resolution observations of the Earth's surface and are widely used in fields such as agricultural monitoring, forestry protection, urban planning, and marine environmental monitoring. However, imaging satellite mission planning faces numerous challenges. The core issue is how to achieve efficient coordination among multiple satellites to complete various observation missions within limited satellite resources. Imaging satellites typically operate in low-Earth orbit and fly at high speeds, resulting in a limited window of visibility when they fly over ground targets. To ensure image quality, satellites often require attitude adjustments, which are subject to multiple constraints such as storage capacity, power supply, and maneuverability. This means that satellites cannot meet the observation needs of all users. How to optimize resource allocation and maximize observation needs within these constraints has become a core issue in imaging satellite mission planning research.

[0003] Traditional resource allocation methods mainly rely on predefined rules or static optimization strategies. This fixed decision-making model is difficult to cope with dynamically changing observation requirements and environmental conditions. Its offline optimization characteristics will also lead to delayed system response and lack the ability to adjust strategies according to real-time status during task execution, which limits the system's performance in dynamic environments.

[0004] Therefore, how to achieve efficient coordination of multiple satellites to complete various observation tasks under limited satellite resources is a difficult problem in current research. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the purpose of this application is to provide a satellite resource adaptive allocation method and system based on deep intensity learning, which can achieve efficient collaboration of multiple satellites to complete various observation tasks under limited satellite resource conditions.

[0006] To achieve the above objectives, in a first aspect, the present application provides a method for adaptively allocating satellite resources based on deep intensity learning, comprising the following steps: S10, constructing a task planning problem model. Based on this model, the ground target is decomposed into multiple optional observation activities. Each ground target corresponds to an independent subtask within its visible time window. These subtasks are defined as meta-tasks. S20, constructs a Markov decision process model based on the meta-task set; S30, based on the Markov decision process model, adopting a Transformer-based policy network model and combining the Actor-Critic algorithm to iteratively optimize the satellite resource allocation strategy for the network model parameters, and then using a heuristic algorithm to determine the specific execution order and execution time period of each satellite resource provider task. The beneficial effects of the present application are as follows: the satellite resource adaptive allocation method based on deep intensity learning provided by the present application converts ground targets into meta-tasks with complete attributes through task meta-taskization, and establishes an MDP model based on this to make the state space and action space clear and controllable. Then, a Transformer-based policy network is combined with the Actor-Critic algorithm, and after training, it is supplemented by heuristic fine-tuning to form a "macro DRL decision-making + micro rule execution" scheduling mechanism, thereby realizing efficient collaboration of multiple satellites to complete various observation tasks under the conditions of limited satellite resources.

[0007] As a further preference, in step S10, the optimization goal of the task planning problem model is to maximize the benefits during the task planning process.

[0008] As a further preference, in step S10, the constraints of the task planning problem model include task uniqueness constraint, task visibility constraint, observation window and resource availability constraint, minimum conversion time constraint and task conflict constraint and resource capacity constraint.

[0009] As a further preferred embodiment, the task uniqueness constraint is that each task can only be executed once at most during the planning and scheduling process, and can only be assigned within a visible time window; The task visibility constraint is that the specific start time and end time of the task must fall within the time window of the scenario to which it belongs; The observation window and resource availability constraints are that for each task in the planning and scheduling process, if resources and corresponding visible time windows are allocated to the task, the observation time period of the task must fall completely within the allocated visible time window; The minimum conversion duration constraint and the task conflict constraint are that when the same satellite resource performs two consecutive observation tasks, the start time between the two tasks should meet the minimum conversion duration constraint, and the execution time segments do not overlap; The resource capability constraint is that the storage capacity and energy consumption required by the task sequence allocated to each resource must not exceed the maximum data storage capacity and total power storage upper limit of the resource.

[0010] As further preferred, the Markov decision process model includes a state representation, an action space and a reward function; The state representation consists of the static state characteristics of the meta-task and the dynamic state characteristics of the resource. The action space is the selection of meta-tasks. Selecting a meta-task means assigning the task to a specified satellite resource. The reward function comprehensively considers the task benefits and conflicts between tasks, providing an evaluation criterion for strategy optimization.

[0011] As further preferred, in step S30, the policy network model includes an encoder and a decoder; The encoder is used to extract high-dimensional representations from the static and dynamic features of tasks and resources; the decoder is used to gradually output satellite resource allocation decisions based on the contextual features generated by the encoder and combined with historical mission information.

[0012] As a further preference, in step S30, the Actor is used to output the action probability distribution of each meta-task according to the environment state, and to achieve optimal allocation of resources by maximizing the expected reward; the Critic is used to evaluate the quality of these strategies and provide feedback to optimize the Actor's decision.

[0013] In a second aspect, the present application provides a satellite resource adaptive allocation system based on deep intensity learning, comprising: The mission planning problem construction module is used to construct a mission planning problem model. Based on this model, the ground target is decomposed into multiple optional observation activities. Each ground target corresponds to an independent subtask within its visible time window. These subtasks are defined as meta-tasks. MDP model building module, used to build a Markov decision process model based on a set of meta-tasks; The satellite resource allocation module is used to iteratively optimize the satellite resource allocation strategy based on the Markov decision process model, a Transformer-based policy network model, and the actor-critic algorithm. The heuristic algorithm is then used to determine the specific execution order and execution time period of each satellite resource provider task. On the third aspect, the present application also provides an application of adaptive allocation of satellite resources based on deep intensity learning as described above, which is applied in the fields of agricultural monitoring, forestry protection, urban planning or marine environment monitoring.

[0014] It can be understood that the beneficial effects of the second and third aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flowchart of a method for adaptively allocating satellite resources based on deep intensity learning provided by an embodiment of the present application; Figure 2This is a schematic diagram of the strategic network model structure provided by a specific embodiment of the present application; Figure 3 This is a framework diagram of the satellite resource adaptive allocation method based on deep reinforcement learning provided in a specific embodiment of the present application. DETAILED DESCRIPTION

[0016] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0017] It should be understood that, in the description of this application, the term "several" means at least one, such as one, two, etc., unless otherwise clearly and specifically defined; the term "plurality" means two or more, unless otherwise clearly and specifically defined; the terms "first" and "second" etc. are used to distinguish different objects, rather than to describe a specific order of objects; the term "and / or" includes any and all combinations of one or more related listed items.

[0018] Additionally, references throughout this specification to "one embodiment," "one embodiment," "an example," or similar language indicate that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. Thus, appearances of the phrase "in one embodiment," "in one embodiment," and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.

[0019] like Figure 1 As shown, the present application provides a satellite resource adaptive allocation method based on deep intensity learning, which can be applied to fields such as agricultural monitoring, forestry protection, urban planning, and marine environment monitoring, including steps S10 to S30, which are detailed as follows: Step S10 constructs a multi-constraint optimization model for the mission planning problem, i.e., the mission planning problem model. This model formally describes factors such as satellite resources, mission sets, visibility time windows, and constraints. Based on this model, ground targets are decomposed into multiple optional observation activities. Each ground target corresponds to an independent subtask within its visibility time window. These subtasks are defined as metatasks to facilitate subsequent satellite resource allocation and mission scheduling.

[0020] In this embodiment, a subtask is the smallest observation unit that divides a ground target within a visible time window. Each subtask contains only one observation activity and its necessary time window information. These subtasks are defined as metatasks.

[0021] Step S20: constructing a Markov Decision Process (MDP) model based on the meta-task set.

[0022] It should be noted that the meta-task set is the sum of all meta-tasks obtained after decomposing all ground targets in the system, which constitutes the action space in MDP modeling.

[0023] In this embodiment, the Markov decision process model includes a state representation, an action space, and a reward function.

[0024] The state representation consists of the static state characteristics of the meta-task and the dynamic state characteristics of the resource. The action space is the selection of meta-tasks. Selecting a meta-task is to assign the task to a specified satellite resource. The reward function comprehensively considers the task benefits and conflicts between tasks, providing an evaluation criterion for subsequent strategy optimization.

[0025] In step S30, a Transformer-based policy network model is used, based on a Markov decision process model, and the Actor-Critic algorithm is used to iteratively optimize the network model parameters. Through continuous training feedback, the model continuously optimizes the resource allocation strategy. A heuristic algorithm is then used to determine the specific execution order and timeframe for each satellite's resource allocation tasks. In this embodiment, deep reinforcement learning is used to allocate resources at a high level, while a heuristic algorithm is used to schedule specific tasks at a low level, thereby ensuring the reasonable execution of tasks after resource allocation.

[0026] The satellite resource adaptive allocation method based on deep intensity learning provided in this embodiment transforms ground targets into meta-tasks with complete attributes through task meta-tasking. Based on this, an MDP model is established to make the state space and action space clear and controllable. Then, a Transformer-based policy network is combined with the Actor-Critic algorithm. After training, it is supplemented by heuristic fine-tuning to form a "macro DRL decision-making + micro rule execution" scheduling mechanism. This allows for efficient collaboration of multiple satellites to complete various observation tasks under limited satellite resources.

[0027] Based on the same inventive concept, the present application also provides a satellite resource adaptive allocation system based on deep intensity learning, including a mission planning problem construction module, an MDP model construction module and a satellite resource allocation module.

[0028] Among them, the task planning problem construction module is used to construct a task planning problem model. Based on this model, the ground target is decomposed into multiple optional observation activities. Each ground target corresponds to an independent subtask within its visible time window, and these subtasks are defined as meta-tasks.

[0029] The MDP model building module is used to build a Markov decision process model based on a set of meta-tasks.

[0030] The satellite resource allocation module is used to iteratively optimize the satellite resource allocation strategy based on the Markov decision process model, adopting a Transformer-based policy network model and combining the Actor-Critic algorithm to optimize the network model parameters. It then uses a heuristic algorithm to determine the specific execution order and execution time period of each satellite resource provider task. It should be noted that the functions of the modules provided in this embodiment can be found in the introduction of the aforementioned method embodiment and will not be repeated here.

[0031] The following describes the satellite resource adaptive allocation method based on deep intensity learning provided by this application in conjunction with specific embodiments.

[0032] Step 1: Constructing the mission planning problem model.

[0033] To formally describe the imaging satellite mission planning problem, this example introduces some necessary definitions and symbols, clarifies the core concepts and variables in mission planning, considers multiple factors such as time constraints, resource constraints, and attitude adjustments, and provides a detailed description of the imaging satellite collaborative mission planning problem model.

[0034] (1) Define the set as follows: 1) Scene Imaging Task Collection , Represents the number of tasks in the scene collection.

[0035] 2) Scene Imaging Resource Collection , Represents the number of available resources in the scene.

[0036] 3) Mission Occupied resources The set of allowed visible time windows { ,…, } , = [ , ] , k ={1,2,…, } .in, Indicates a task Occupied resources The number of visible time windows allowed at a time. and Respectively represent the tasks Occupied resources The first The start and end time of the visible time window.

[0037] 4) Resources The effective execution range set of in, It is a resource The total number of valid time intervals, and Represents resources No. The start and end time of a time period.

[0038] 5) Definition For all resources that can be arranged The set of tasks to be executed, To meet the observation mission The requested resource collection, .

[0039] (2) Define the parameters as follows: 1) : Scene start time.

[0040] 2) : The end time of the scene.

[0041] 3) :Indicates a task The specific start time when the task is selected for scheduling. .

[0042] 4) :Indicates a task The specific end time when the task is selected for scheduling. .

[0043] 5) :Task The weight of . , .

[0044] 6) :Task The duration required to complete. , >0 . 7) :Task Complete the required storage capacity consumption. , >0 .

[0045] 8) :The shortest conversion time of the task. and tasks Being scheduled and occupying resources at the same time ,Task After the execution is completed, continue to execute the task The conversion time is , .

[0046] 9) : Indicates resources The total data storage capacity limit.

[0047] (3) Define the optimization objective as follows: Remember tasks In Resources On the The execution status on the visible time window is a Boolean variable, Indicates a task Assigned to a visible time window Due to the differences in the importance of tasks, different tasks are assigned different weights. The weight of a task is regarded as the benefit brought by successful execution. In this embodiment, the optimization goal is set to maximize the benefit during the task planning process:

[0048] (4) Define the constraints as follows: 1) Task uniqueness constraint: Each task It can only be executed once during the planning and scheduling process, and can only be allocated within a visible time window.

[0049]

[0050] 2) Task visibility constraints: Task The specific start time of the scheduled schedule and end time Must fall within the scene time window If ,but:

[0051]

[0052] 3) Observation window and resource availability constraints: For each task in the planning and scheduling process, if the task Resources allocated and the corresponding visible time window, the observation period of the task must fall completely within the allocated visible time window. , ,if ,but:

[0053] 4) Minimum transition time constraint and task conflict constraint: When the same satellite resource performs two consecutive observation tasks, it needs to complete the transition from one observation attitude to another within a certain period of time. Therefore, the start time between the two tasks should meet the minimum transition time constraint, and the execution time segments should not overlap. That is, for any and any two consecutive observation task sequences assigned to this resource , ,like and If both are established, then:

[0054]

[0055] 5) Resource capacity constraint: The storage capacity and energy consumption required by the task sequence assigned to each resource must not exceed the maximum data storage capacity and total power storage limit of the resource:

[0056] Step 2: Deep reinforcement learning decision-making for resource allocation problems Based on the concept of meta-task, meta-task It not only represents the results of task decomposition but also records the time window information bound to resources, so that each meta-task has attributes directly related to resource allocation. On this basis, the agent acts as a decision maker, using the MDP model as an environment model and learning the optimal strategy through interaction with the environment. The specific MDP definition is as follows: (1) Status indication In the resource allocation process, assuming the total time step is , the current state It consists of two components: Static State, ) and Dynamic State, ). The static state stores each meta-task The basic information of , and remains unchanged during the task allocation process. The formal expression is:

[0057] Among them, a single meta-task Static characteristics of Depend on composition. , Respectively represent the meta-task in the resource The start and end time of the time window selected above. If the task In Resources exist Time window set , then the time window with the smallest timing conflict is selected first:

[0058]

[0059] in, It is a time window The timing conflict index is calculated as follows: Recorded time window The corresponding minimum timing violation indicator.

[0060]

[0061]

[0062] The dynamic state records the real-time status of resources and is updated gradually as the allocation process progresses. It can be formally expressed as:

[0063] in, Represents the current time step resource The current remaining capacity of the resource is used to evaluate resource load. The sum of the timing conflicts of all assigned meta-tasks on the current resource is accumulated to reflect the conflict pressure of the resource in real time. Indicates the resources currently allocated A set of meta-tasks.

[0064]

[0065] (2) Action Space In this model, the action Defined as "select meta-task" or "terminate". Actions include the following two types: 1) Select meta-task : Schedule tasks to resources , waiting to be scheduled for execution; 2) If in time step There is no meta-task to select. Select "Terminate" to end the current resource allocation process.

[0066] In this process, in order to reduce unnecessary actions, the model introduces a mask mechanism before action execution to avoid invalid meta-task selection.

[0067] (3) State transition rules When the agent is in state Next action When the agent moves to the next state State transition is mainly affected by the following two factors: 1) Task allocation update: When a meta-task is assigned to a specific resource, the storage capacity and conflict index of the resource will be updated according to the allocation requirements.

[0068] 2) Meta-task queue update: Assigned meta-tasks are removed from the task queue to avoid affecting the selection order and status evaluation of unassigned tasks.

[0069] (4) Reward Function The reward function is used to evaluate the effectiveness of the actions performed by the agent at each time step. It is defined as follows: 1) If in time step A meta-task is selected, and the reward value is its corresponding weight .

[0070] 2) If no meta-task is selected or there is no available task at the current time step, the reward .

[0071] The total reward R of the entire decision-making process is expressed as:

[0072] In this formula, Reflects the agent's The selected meta-task reward, and The total reward is reduced to encourage the agent to balance resource conflicts when selecting tasks.

[0073] Step 3: Transformer-based policy network model After completing the definition of the MDP model, in order to effectively solve the complex dependencies and efficient decision-making requirements in the dynamic resource allocation problem, this paper designs a policy network model based on the Transformer architecture. The model consists of two main modules: the encoder and the decoder. The two complement each other. The encoder is used to extract high-dimensional representations from the static and dynamic features of tasks and resources, and the decoder gradually outputs allocation decisions based on the contextual features generated by the encoder. The model structure is as follows Figure 2 shown.

[0074] (1) Encoder The input of the encoder is the state vector, which comes from the state information defined in the MDP model and is composed of the static feature sequence and dynamic feature sequences Since the number of meta-tasks for each task is inconsistent, in order to ensure the consistency of the encoder input dimension, the static feature sequence is specified as follows: If the task In Resources There is no available time window on , that is, there is no corresponding meta-task, then .

[0075] Static feature sequence and dynamic feature sequences Each is mapped to a high-dimensional embedding space through an independent linear projection layer to generate an embedding vector and , the specific mapping process is as follows:

[0076]

[0077] In linear transformation, represents the weight matrix, The embedding dimension is set to 128 for the bias term. Then, the static embedding and dynamic embedding are concatenated to obtain the feature vector To further capture the potential dependencies between tasks and resources, the concatenated feature vectors are input into the multi-head attention mechanism for feature fusion. In the attention calculation, the hidden features of the input is mapped into query, key, and value vectors:

[0078] in, The weight matrix of query, key and value is a trainable model parameter. and key The dot product calculation result is:

[0079] The dot product result needs to be scaled and applied Function, which converts the relevance score into a probability distribution, indicating the degree of attention each task pays to different resources. Scaling factor It is used to prevent the gradient from being too large and ensure the stability of model training. Combined to generate attention weights:

[0080] Attention weight The outputs of are concatenated and linearly transformed to generate the final output of the multi-head attention layer:

[0081] is the weight matrix of the linear transformation after concatenation. The output of the attention layer is further processed by the feedforward network (FF), and the hidden features of the next layer are generated through skip connections and normalization operations (Batch Normalization, BN):

[0082] The feedforward network consists of two layers of linear transformation and The activation function is calculated as follows:

[0083] After multiple layers of stacked attention calculations and feature feedforward processing, the encoder ultimately outputs a high-dimensional contextual feature representation for each task and resource. This feature effectively captures the potential dependencies between tasks and the dynamic changes in resource status, providing rich contextual information for the decoder's decision-making.

[0084] (2) Decoder In the decoding stage, based on the context features generated by the encoder and combined with historical task information, a decision sequence is gradually generated. , the selected meta-task Represents the current partial solution, which is used to update the status of task assignment. , the decoder first performs the historical meta-task sequence assigned to the resource Perform the maximum pooling operation (Max Pooling) to extract the aggregated features:

[0085] in, Indicates the total number of resources. Aggregate characteristics of all resources Splicing into global context features , then processed by linear transformation and feedforward network to generate time steps Embedding vector of :

[0086] is the weight matrix and bias term of the feedforward network, which can be learned through training. In order to combine the context features generated by the encoder with the current feature representation of the decoder, the decoder embeds the encoder output and the current time step Perform concatenation and generate comprehensive feature representation through linear transformation:

[0087] Represents the trainable parameters of the linear transformation layer. The dependency between historical allocation decisions and current state is effectively balanced. Function pair comprehensive features Perform normalization to generate the probability distribution of all candidate meta-tasks:

[0088] The decoder adopts a sampling method based on probability distribution. Select the action for the current time step (i.e., the selected meta-task).

[0089] In this embodiment, adaptive resource allocation for satellite missions utilizes a Transformer-based policy network model, incorporating the actor-critic algorithm from deep reinforcement learning. The actor network generates a series of possible meta-task selection strategies based on the environment state and optimizes resource allocation by maximizing expected rewards. The critic network evaluates the quality of these strategies and provides feedback to optimize the actor's decision-making. Furthermore, a heuristic rule-based single-satellite mission planning solution is introduced as a supplementary approach to determine the specific execution order and timeframe for tasks on each resource.

[0090] Through the above steps, the satellite resource adaptive allocation method based on deep reinforcement learning proposed in this embodiment has the following framework: Figure 3 As shown in the figure, by combining deep reinforcement learning with traditional scheduling methods, through effective resource allocation and task scheduling optimization, it can achieve efficient, intelligent and adaptive allocation of satellite resources under multiple task requests, limited resource conditions and dynamically changing environments, and enhance the autonomous decision-making ability of satellite mission planning.

[0091] The first purpose of this embodiment is to propose a mission planning problem model, which aims to accurately describe the key state variables and decision factors in the satellite resource scheduling process, and provide high-quality input for subsequent reinforcement learning training and decision-making.

[0092] The second goal of this example is to build a task scheduling model based on a Markov decision process. This model explicitly defines the state representation of meta-tasks, the action space, and the reward function that integrates task benefits and conflicts, providing a quantitative basis for performance evaluation and iterative optimization of the policy network.

[0093] The third objective of this embodiment is to design a Transformer-based policy network model, use the Actor-Critic algorithm to perform online training and iterative updates on network parameters, and use reward feedback as a driver to achieve online updates and dynamic adjustments to the policy. It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A satellite resource adaptive allocation method based on deep intensity learning, characterized in that: The steps include: S10, constructing a task planning problem model. Based on this model, the ground target is decomposed into multiple optional observation activities. Each ground target corresponds to an independent subtask within its visible time window. These subtasks are defined as meta-tasks. S20, constructs a Markov decision process model based on the meta-task set; S30, based on the Markov decision process model, adopting a Transformer-based policy network model and combining the Actor-Critic algorithm to iteratively optimize the satellite resource allocation strategy for the network model parameters, and then using a heuristic algorithm to determine the specific execution order and execution time period of each satellite resource provider task.

2. The satellite resource adaptive allocation method based on deep intensity learning according to claim 1, characterized in that In step S10, the optimization goal of the task planning problem model is to maximize the benefits during the task planning process.

3. The satellite resource adaptive allocation method based on deep intensity learning according to claim 1, characterized in that In step S10, the constraints of the task planning problem model include task uniqueness constraint, task visibility constraint, observation window and resource availability constraint, minimum conversion time constraint and task conflict constraint and resource capacity constraint.

4. The satellite resource adaptive allocation method based on deep intensity learning according to claim 3, characterized in that The task uniqueness constraint is that each task can only be executed once at most during the planning and scheduling process, and can only be assigned within a visible time window; The task visibility constraint is that the specific start time and end time of the task must fall within the time window of the scenario to which it belongs; The observation window and resource availability constraints are that for each task in the planning and scheduling process, if resources and corresponding visible time windows are allocated to the task, the observation time period of the task must fall completely within the allocated visible time window; The minimum conversion duration constraint and the task conflict constraint are that when the same satellite resource performs two consecutive observation tasks, the start time between the two tasks should meet the minimum conversion duration constraint, and the execution time segments do not overlap; The resource capability constraint is that the storage capacity and energy consumption required by the task sequence allocated to each resource must not exceed the maximum data storage capacity and total power storage upper limit of the resource.

5. The satellite resource adaptive allocation method based on deep intensity learning according to claim 1, characterized in that The Markov decision process model includes a state representation, an action space, and a reward function; The state representation consists of the static state characteristics of the meta-task and the dynamic state characteristics of the resource. The action space is the selection of meta-tasks. Selecting a meta-task means assigning the task to a specified satellite resource. The reward function comprehensively considers the task benefits and conflicts between tasks, providing an evaluation criterion for strategy optimization.

6. The satellite resource adaptive allocation method based on deep intensity learning according to claim 1, characterized in that In step S30, the policy network model includes an encoder and a decoder; The encoder is used to extract high-dimensional representations from the static and dynamic features of tasks and resources; the decoder is used to gradually output satellite resource allocation decisions based on the contextual features generated by the encoder and combined with historical mission information.

7. The satellite resource adaptive allocation method based on deep intensity learning according to claim 1, characterized in that: In step S30, the Actor is used to output the action probability distribution of each meta-task based on the environment state and achieve optimal resource allocation by maximizing the expected reward; the Critic is used to evaluate the quality of these strategies and provide feedback to optimize the Actor's decision.

8. A satellite resource adaptive allocation system based on deep intensity learning, characterized in that: include: The mission planning problem construction module is used to construct a mission planning problem model. Based on this model, the ground target is decomposed into multiple optional observation activities. Each ground target corresponds to an independent subtask within its visible time window. These subtasks are defined as meta-tasks. MDP model building module, used to build a Markov decision process model based on a set of meta-tasks; The satellite resource allocation module is used to iteratively optimize the satellite resource allocation strategy based on the Markov decision process model, a Transformer-based policy network model, and the actor-critic algorithm. The heuristic algorithm is then used to determine the specific execution order and execution time period of each satellite resource provider task.

9. An application of satellite resource adaptive allocation based on deep intensity learning according to any one of claims 1 to 7, characterized in that: It is used in the fields of agricultural monitoring, forestry protection, urban planning or marine environment monitoring.

Citation Information

Cited By

  • Planning method for agile observation satellite

    CN120746211A

  • Space-based target visible window calculation and detection performance analysis method considering multiple constraints

    CN121477223A

  • Satellite ground station task resource allocation method and device and electronic equipment

    CN122204154A