A heterogeneous constellation intelligent task decision method for spatial anomaly target observation

By establishing a multi-constraint, multi-objective optimization model and a Markov decision model, and combining it with an intelligent decision-making algorithm based on empirical learning and iterative improvement, the problem that traditional scheduling methods are difficult to meet the observation of space-moving targets was solved. This enabled efficient task allocation and real-time planning for heterogeneous constellations, improving observation quality and system efficiency.

CN120974902BActive Publication Date: 2026-03-31TIANJIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Traditional fixed observation and scheduling methods are difficult to meet the timeliness and accuracy requirements of space-moving targets. How to carry out efficient task allocation and coordination decisions in complex dynamic environments, especially multi-satellite relay observation of high-speed moving targets to improve space early warning capabilities.

Method used

A multi-constraint, multi-objective optimization model was established, a Markov decision model was designed, and an intelligent decision algorithm based on experience learning, mask mechanism, and iterative improvement was adopted to realize task decomposition and real-time planning for heterogeneous constellations. Satellite selection was performed in conjunction with a deep pointer network to optimize double coverage and system switching frequency to improve observation quality.

Benefits of technology

It enables efficient relay tracking and observation of dynamic targets, improves space early warning capabilities, ensures observation quality and system efficiency, and meets the requirements of real-time performance and optimality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974902B_ABST
    Figure CN120974902B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of intelligent algorithm, heterogeneous satellite planning and dynamic target monitoring, and particularly relates to a heterogeneous constellation intelligent task decision method for space dynamic target observation, comprising: step 1: establishing a multi-constraint multi-objective optimization model for dynamic target observation, and designing a Markov dynamic decision model, i.e. determining the elements of the state set of the task decision model; step 2: designing a heterogeneous star cluster intelligent decision algorithm architecture based on experience learning-Mask mechanism-iterative improvement, i.e. realizing offline training of the heterogeneous star cluster task decision algorithm; and step 3: online decision, i.e. using the trained network to perform real-time relay observation on dynamic targets. The present application proposes a general relay observation task scheduling framework which comprehensively considers double coverage, system switching times and observation quality, realizes relay observation on dynamic targets by selecting the optimal multiple stars in the heterogeneous star cluster to maintain target positioning, and improves the space early warning capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of intelligent algorithms, heterogeneous satellite planning and dynamic target monitoring technology, and specifically relates to a heterogeneous constellation intelligent mission decision-making method for observing space-moving targets. Background Technology

[0002] Faced with the complex and ever-changing space environment and increasingly frequent space anomalies, such as rapidly drifting debris, non-cooperative spacecraft maneuvers, and low-Earth orbit anomalies, traditional fixed observation and scheduling methods are no longer sufficient to meet the timeliness and accuracy requirements of mission response. Heterogeneous constellations, composed of multiple types of satellites with different payloads (such as optical, radar, and electromagnetic) and orbital characteristics, possess stronger all-weather and spacetime coverage capabilities and multi-modal observation advantages. Utilizing heterogeneous constellations to achieve efficient perception and continuous tracking of space anomalies has become an important development direction for current space situational awareness. However, different satellites differ in observation capabilities, communication links, energy budgets, and maneuverability. How to conduct efficient mission allocation and coordination decisions in complex and dynamic environments has become a key technical challenge. Furthermore, compared to static mission planning for ground targets, dynamic mission planning for satellite constellations has two very distinct characteristics. First, non-cooperative targets are typically high-speed spacecraft or moving satellites, in a state of high-speed motion, whereas in static mission planning for ground targets, the target's motion trend can be almost ignored. Second, multiple targets to be observed have a high degree of spatiotemporal coupling with our observation satellites. On a temporal scale, our satellites need to continuously relay-track and observe the target. For the same non-cooperative dynamic target, multiple different observation combinations are usually required to complete the relay, and frequent switching of observation combinations will reduce the system's efficiency. On a spatial scale, within a certain scheduling cycle, there are multiple visible observation satellites for the same target, requiring decisions on the combination of observation resources. At the same time, due to the high-speed motion characteristics of the target, the relative motion trends between different satellites and the target are different, which will further lead to differences in observation quality. Again, decisions on the combination of observation resources are needed to improve observation quality.

[0003] Therefore, this invention provides a heterogeneous constellation intelligent mission decision-making method for observing space-moving targets. This invention focuses on the constellation mission decision-making problem for dynamic target observation missions and proposes a general relay observation mission scheduling framework that comprehensively considers double coverage, system switching times, and observation quality. This framework enables the selection of the optimal multiple stars in a heterogeneous constellation to conduct relay observations of dynamic targets in order to maintain target positioning and improve space early warning capabilities. Summary of the Invention

[0004] The purpose of this invention is to provide an intelligent mission decision-making method for heterogeneous constellations for observing space-moving targets. This invention considers the relay observation problem of globally distributed, high-speed moving dynamic targets. First, it establishes constraints such as visibility constraints, dual observation constraints, and relay observation constraints for dynamic targets, and considers comprehensive optimization objectives such as dual coverage of agile imaging satellites, system switching times, and observation quality to establish an optimization model for threat monitoring tasks. Then, addressing the need for continuous relay observation of dynamic targets by heterogeneous satellites, it designs an intelligent decision-making algorithm architecture based on empirical learning, a mask mechanism, iterative improvement, and online decision-making to achieve dual relay tracking observation of dynamic targets. To this end, the technical solution adopted by this invention is as follows: first, based on the space target observation requirements, corresponding constraints and a comprehensive benefit objective function are established, and a Markov decision model for heterogeneous constellation mission planning is designed. Then, an intelligent decision-making algorithm architecture based on empirical learning, a mask mechanism, and iterative improvement is established. Finally, an offline-trained mission decision-making model is used for online planning of satellite observation windows.

[0005] The specific technical solution adopted by this invention is as follows:

[0006] A heterogeneous constellation intelligent mission decision-making method for observing space-moving targets includes the following steps:

[0007] Step 1: Establish a multi-constraint, multi-objective optimization model for dynamic target observation, and design a Markov dynamic decision model, i.e., determine the various elements of the state set of the task decision model:

[0008] First, considering the global distribution characteristics of heterogeneous satellite constellations and dynamic targets, a task decomposition framework is determined. The relay observation process of multiple dynamic targets is divided into multiple consecutive sub-tasks, and solutions are performed in each sub-task. The relative motion relationship between the target and the observation satellite, the unobservable blind zone, double coverage, and communication link constraints are considered. The objective function is to maximize the completion of continuous relay observation of the target. A dynamic satellite constellation task decision model is established to determine the order of observation of the target by each satellite and the handover relationship.

[0009] Secondly, with the comprehensive optimization objectives of double coverage, number of heterogeneous satellite switching and observation quality for the overall operation of dynamic targets, an optimization model for dynamic target observation tasks was established, and on this basis, a Markov decision model for relay observation tasks was established.

[0010] Step 2: Design a heterogeneous constellation intelligent decision-making algorithm architecture based on experience learning, mask mechanism, and iterative improvement, i.e., realize offline training of the heterogeneous constellation task decision-making algorithm:

[0011] Based on the established constellation task allocation model, considering different satellite observation resources and satellite communication capabilities, an intelligent allocation solution strategy based on experience learning, Mask mechanism, iterative improvement, and online decision-making is designed. The Mask mechanism is used to handle different selection orders of heterogeneous satellites such as electronic reconnaissance and optical satellites. Combined with historical experience data provided by the simulation platform and online network adjustment, a dynamic target continuous relay observation offline training strategy that balances real-time performance and optimality is obtained.

[0012] Step 3: Online decision-making, which involves using a trained network for real-time relay observation of dynamic targets:

[0013] After completing the offline learning process through steps 1 and 2, the parameters of the entire task decision network for dynamic target observation are determined. At this point, in each scheduling cycle, the Actor network estimates the future cumulative revenue of the allocation strategy based on the visibility information of different types of satellites for the space target to be observed, determines the type of observation satellite and the observation time window, and determines the next cycle's relay satellite and observation window at the end of each scheduling cycle, thus realizing online real-time task planning until the entire observation of the dynamic target is completed.

[0014] The technical effects achieved by this invention are as follows:

[0015] This invention addresses the constellation mission decision-making problem for space dynamic target observation missions. It proposes a general relay observation mission scheduling framework that comprehensively considers double coverage, system switching times, heterogeneous satellite characteristics, communication links, and observation quality. The dynamic target relay observation mission is decomposed into a series of sub-tasks for solution.

[0016] This invention addresses the problem of solving continuous relay observation tasks for dynamic targets by proposing a flexible and adaptive experience-mask mechanism-iterative improvement heterogeneous constellation intelligent task decision-making algorithm. A dynamic mask mechanism is designed to better handle the practical engineering problem of changes in action space caused by different selection numbers of heterogeneous satellites such as electronic reconnaissance and optical satellites during relay observation. Combined with a depth pointer network, the algorithm performs real-time dynamic selection of heterogeneous satellites, meeting the "plug and play" requirement and ensuring the optimal satellite selection in real time. Attached Figure Description

[0017] Figure 1 This is a diagram of the overall technical solution of the present invention;

[0018] Figure 2 This is a schematic diagram of solar avoidance in this invention;

[0019] Figure 3 This is a schematic diagram of heterogeneous satellite observation in this invention;

[0020] Figure 4This is a diagram illustrating the depth pointer network calculation process in this invention;

[0021] Figure 5 This is a diagram of the Critic network structure in this invention;

[0022] Figure 6 This is a diagram showing the task allocation results from 0 to 250 seconds in this invention;

[0023] Figure 7 This is a diagram showing the task allocation results for 250-500 seconds in this invention;

[0024] Figure 8 This is a diagram showing the task allocation results for 500-750 seconds in this invention;

[0025] Figure 9 This is a diagram showing the task allocation results for 750-1000 seconds in this invention;

[0026] Figure 10 This is the double coverage map of the observed target in this invention. Detailed Implementation

[0027] To make the objectives and advantages of this invention clearer, the invention will be specifically described below with reference to embodiments. It should be understood that the following text is merely used to describe one or more specific embodiments of the invention and does not strictly limit the scope of protection specifically claimed by the invention.

[0028] like Figure 1 As shown, a heterogeneous constellation intelligent mission decision-making method for observing space-moving targets is disclosed. This invention relates to the fields of intelligent algorithms, heterogeneous satellite planning, and dynamic target monitoring, and particularly to a heterogeneous constellation intelligent mission decision-making method for observing space-moving targets. Specifically, it involves implementing real-time mission planning for heterogeneous constellation satellites using an intelligent decision-making algorithm architecture of experience learning-mask mechanism-iterative improvement-online decision-making, including the following steps:

[0029] Step 1: Establish a multi-constraint, multi-objective optimization model for dynamic target observation, and design a Markov dynamic decision model, i.e., determine the various elements of the state set of the task decision model:

[0030] First, considering the global distribution characteristics of heterogeneous satellite constellations and dynamic targets, a task decomposition framework is determined. The relay observation process of multiple dynamic targets is divided into multiple consecutive sub-tasks, and solutions are performed in each sub-task. The relative motion relationship between the target and the observation satellite, the unobservable blind zone, double coverage, and communication link constraints are considered. The objective function is to maximize the completion of continuous relay observation of the target. A dynamic satellite constellation task decision model is established to determine the order of observation of the target by each satellite and the handover relationship.

[0031] Secondly, with the comprehensive optimization objectives of double coverage, number of heterogeneous satellite switching and observation quality for the overall operation of dynamic targets, an optimization model for dynamic target observation tasks was established, and on this basis, a Markov decision model for relay observation tasks was established.

[0032] Preferably, in step 1, during the optimization model establishment process: during the tracking and observation of multiple targets, due to the high-speed movement of the targets, the spatial positions of the targets and the observation satellites are constantly changing, and the visible window is also constantly changing. Therefore, directly providing the observation results for the entire life cycle of multiple dynamic targets is very complex. It is necessary to decompose the observation task and divide the life cycle of the dynamic targets into a series of continuous time periods, i.e., scheduling cycles. Then, within each scheduling cycle, the scheduling results of the observation satellites are given. When the current scheduling cycle ends, the observation satellites in the scheduling results of the next scheduling cycle take over the observation of the dynamic targets.

[0033] Preferably, in step 1, the following six constraints are first established based on the observation characteristics of constellation satellites of dynamic space targets in each scheduling cycle:

[0034] a. Constraints of Deep Space Background Observation: Low-Earth Orbit agile satellites employ deep space background observation for non-cooperative dynamic targets in space. This means that the satellite's line of sight to the target must be higher than the Earth's atmospheric altitude, specifically the observation angle θ formed by the satellite's line of sight and the line connecting the satellite to the Earth's center. tar It must be greater than the adjacent observation angle θ ob ,Right now:

[0035] θ tar ≥ θ ob (1-1)

[0036] The calculation method for the near-edge observation angle is as follows:

[0037]

[0038] Among them, R e H represents the Earth's radius. a H represents the altitude of the atmosphere. sat Indicates the satellite's altitude;

[0039] b. Solar avoidance constraint: When satellites are conducting relay observations, they need to avoid sunlight affecting the detection process. There is a solar avoidance angle constraint. When this constraint is not met, the target cannot be observed during this period.

[0040] α sun >α ob (1-3)

[0041] Where α sun For the line-of-sight solar angle, α obSolar avoidance angle; see definition. Figure 2 As shown;

[0042] c. Maximum detection range constraint: When a non-cooperative dynamic target satisfies the deep space background observation constraint for the observation satellite, the maximum detection range constraint of the optical payload for the non-cooperative dynamic target must also be satisfied. The maximum detection range depends on the detection capability of the payload, the target strength, and the background strength.

[0043] d Heterogeneous satellite observation constraints

[0044] To obtain the position and velocity information of a dynamic target, it is first necessary to determine a suitable electronic satellite based on the observation characteristics of heterogeneous satellites to determine the approximate position information of the dynamic target. Furthermore, at least two optical or SAR satellites are needed simultaneously to perform stereo positioning of the moving target, achieving relay observation within the scheduling cycle. That is, the optical or SAR satellites must meet the dual observation constraint; for example... Figure 3 As shown,

[0045] e. Relay observation constraints

[0046] When switching observation satellite combinations for the same non-cooperative dynamic target, it is necessary to ensure that before the observation window of the previous group of observation satellites releases the observation resources, the next group of observation satellites has completed the preliminary positioning of the target to be observed, captured the target to be observed, and is ready to start relay observation at any time. After the previous group of observation satellites releases the observation resources, the relay observation combination can immediately enter the monitoring state and complete the positioning of the target to be observed.

[0047]

[0048] in, This represents the start time of dynamic objective j in the k-th scheduling cycle. The duration of the double observation window for dynamic target j is represented by Targets, which represents the set of targets.

[0049] f. Communication link constraints

[0050] For heterogeneous satellites with different payloads, such as electronic reconnaissance satellites, optical satellites, or SAR satellites, their communication capabilities, data transmission rates, and link durations vary significantly. During mission scheduling, satellites with stable communication links and the ability to transmit observation data in real-time or near real-time should be prioritized for observation missions. Especially in scenarios involving rapid response observation of non-cooperative dynamic targets, if some optical satellites, despite having the necessary observation capabilities, experience communication link interruptions or severe delays, it will impact data transmission efficiency and subsequent decision-making effectiveness. Therefore, while meeting detection requirements, a communication link constraint mechanism needs to be introduced to comprehensively evaluate the communication performance of various payload satellites, dynamically adjust mission allocation strategies, and ensure that target information can be effectively transmitted within time constraints.

[0051] Secondly, determine the objective function for dynamic target observation:

[0052] The goal of constellation dynamic mission planning for threat monitoring is to achieve double-coverage observation of non-cooperative dynamic targets in space, ensuring that the system has high working efficiency and observation conditions while tracking and locating more dynamic targets.

[0053]

[0054] In the formula, C j (T k () represents the double coverage rate of dynamic target j in the k-th scheduling period. This represents the double coverage rate of all J dynamic targets over a cumulative K scheduling cycles. This represents the cumulative number of switches among N observation satellites. This represents the cumulative observation factor of N observation satellites, which is related to the relative motion trend of the target to be observed and the observation satellites, and reflects the quality of the observation conditions. α1, α2, and α3 represent weighting coefficients.

[0055] Preferably, the establishment of a Markov decision model for observing dynamic targets using heterogeneous satellites includes the following steps:

[0056] First: Establish a state set S; during the relay observation of a dynamic target, the satellite's visibility to the target and its relative motion characteristics both affect the target observation effect; therefore, the distance between the observing satellite and the target, the relative motion between the satellite and the target, the observable time window of the satellite, and the number of switching are used as state s. T Then the state set can be represented as follows:

[0057]

[0058] In the formula, Dis represents the status information of the i-th satellite. i,j ,rela i,j These represent the distance between the target and the satellite, and their relative motion, respectively. i,j This represents the visible time window of satellite i relative to target j, change i This indicates the number of times the i-th satellite has been switched;

[0059] Next, establish: establish action set A; based on the heterogeneous satellite observation constraints, plan the decision center as J dynamic targets to be observed, first select J electronic reconnaissance satellites for preliminary reconnaissance, and further select 2*J optical observation satellites that can communicate with the confirmed electronic reconnaissance satellites according to the satellite communication link, and assign the corresponding visible time window type as action a.

[0060] a = {dian1, dian2, ..., dian} J ,guang1,guang2,...,guang 2*J}, a∈A (1-7)

[0061] Among them dian j This represents the electronic reconnaissance satellite selected to observe the j-th target. 2*j ,guang 2*j-1 This indicates the optical satellite selected for observing the j-th target;

[0062] Then, the immediate benefit is R; in the learning process of collaborative observation tasks for dynamic targets, considering the optimization objectives defined above, if the coverage of the dynamic target is higher while maintaining high system efficiency and observation conditions, the overall observation benefit will be greater, and the reward will be larger. Therefore, the reward design is as follows:

[0063]

[0064] Finally, the discount factor γ is similar to that in static task planning. γ represents the importance of future returns relative to current returns, with 0 < γ < 1. When γ = 0, it is equivalent to considering only current returns and not future returns. When γ = 1, it means that future returns and current returns are considered equally important.

[0065] Step 2: Design a heterogeneous constellation intelligent decision-making algorithm architecture based on experience learning, mask mechanism, and iterative improvement, i.e., realize offline training of the heterogeneous constellation task decision-making algorithm:

[0066] Based on the established constellation task allocation model, considering different satellite observation resources and satellite communication capabilities, an intelligent allocation solution strategy based on experience learning, Mask mechanism, iterative improvement, and online decision-making is designed. The Mask mechanism is used to handle different selection orders of heterogeneous satellites such as electronic reconnaissance and optical satellites. Combined with historical experience data provided by the simulation platform and online network adjustment, a dynamic target continuous relay observation offline training strategy that balances real-time performance and optimality is obtained.

[0067] Preferably, in step 2, considering the different observation functions of electronic reconnaissance satellites and optical satellites in heterogeneous constellations, during the observation of dynamic targets, due to the dynamic nature of the observed targets, electronic reconnaissance satellites, optical satellites, and SAR satellites need to conduct time-sharing relay observations. According to the heterogeneous satellite observation constraints established earlier, one electronic reconnaissance satellite needs to be selected in each scheduling observation cycle, while optical satellites and SAR satellites need to meet the dual observation requirements. Therefore, this will cause the action space of the intelligent task decision network to change depending on the selected satellite. Traditional intelligent decision networks are difficult to solve this problem. Therefore, this patent proposes a heterogeneous constellation intelligent decision algorithm based on experience learning-mask mechanism-iterative improvement.

[0068] First, experience-based learning: data sampling based on expert experience.

[0069] During the data acquisition process, the state s is obtained according to the dynamic target observation task requirements and sent to the task decision network to output the corresponding operation a, that is, for the satellites allocated to the current task, including electronic reconnaissance satellites, optical satellites, or SAR satellites and the corresponding available time windows, the reward r is obtained; after the above operation is executed, the data will be stored in the following form in each iteration round: (s,a,r,s′).

[0070] This process requires numerous rounds to acquire sufficient experience data, all of which is stored in its own replay experience pool. Therefore, this patent employs a method of collecting and training the network simultaneously. At the beginning of training, it collects batch-size decision data from its own experience pool with probability p, and samples from the expert experience pool with probability 1-p. The decision data in the expert experience pool has high reference value, thereby improving the learning speed and training stability of the algorithm. As the number of training iterations increases, p is gradually increased, and the network gradually uses its own data for learning to improve its adaptability.

[0071] Then, the Mask mechanism: adaptively changes the selection of requirements based on the characteristics of heterogeneous satellites;

[0072] In each scheduling cycle of dynamic target observation, it is necessary to select appropriate satellites to perform observation tasks based on the established constraints and objective function. Ultimately, this is a permutation and combination problem. Therefore, a satellite constellation task decision network architecture based on a depth pointer network is designed. Furthermore, considering the engineering requirements that the action space changes constantly due to the different characteristics of heterogeneous satellites, the depth pointer network is further improved. A heterogeneous satellite constellation task decision network architecture based on the Mask mechanism is designed to realize the selection of different satellites and corresponding time windows to ensure real-time relay tracking of dynamic targets.

[0073] The constellation dynamic mission planning model for dynamic target observation missions is based on the Actor-Critic network architecture; these will be introduced separately below.

[0074] Based on the Actor-Critic network architecture design, the following sections will introduce each aspect:

[0075] Actor Network:

[0076] Based on the dynamic nature of target observation and the requirements for selecting heterogeneous satellites, the Actor network is designed as a deep pointer network architecture based on the Mask mechanism. It is responsible for calculating the collaborative task planning results for each scheduling cycle in real time, and its objective is to approximate the optimal task planning strategy π. * It is about maximizing the cumulative task rewards:

[0077]

[0078] Where, θ * The network parameters represent the optimal strategy; the input of the depth pointer network based on the Mask mechanism is the state sequence of each type of satellite's observation of the target in the k-th scheduling period, and the output of the current scheduling period is the observation combination corresponding to each target, including the selection combination of one electronic reconnaissance satellite and two optical or SAR satellites;

[0079] The overall network structure of the pointer network is as follows: Figure 4 As shown, the overall network structure of the pointer network consists of an encoder and a decoder. The encoder is composed of a one-layer one-dimensional convolutional neural network, which is responsible for processing the input sequence information s. T,k Feature extraction is performed to obtain embedded information.

[0080] The decoder is the core component of the pointer network. It calculates the selection probability of each satellite in the input sequence through an attention mechanism. In the online scheduling stage, a greedy strategy is adopted to select the satellite with the highest probability value as the observation resource each time, until the observation satellites corresponding to J dynamic targets in the current scheduling cycle are selected. Considering the uniqueness observation constraint and the heterogeneous satellite observation constraint, that is, each satellite can only observe one target at a time, and the engineering requirement that the action space changes at any time due to the different characteristics of heterogeneous satellites, a dynamic mask mechanism is innovatively designed in each decoding process. That is, when calculating the selection probability of each satellite in the input sequence, the probability of the selected satellite and the non-selected satellite type is set to 0. For example, if an electronic reconnaissance satellite is to be selected, the probability of optical satellite or SAR satellite is set to 0, thereby avoiding the repeated selection of the same satellite and satisfying the selection requirements of heterogeneous satellite characteristics.

[0081] Critic Network:

[0082] Critic network structure Figure 5 As shown, in the offline training phase of the stellar dynamic task planning algorithm, a dual-Critic network structure is used for training. The input is the current state information sequence of the stellar cluster, and the output is the future cumulative expected reward under the current state, which is used to evaluate the quality of the current state. The dual-Critic network architecture is exactly the same, consisting of an input layer, four layers of one-dimensional convolutional neural networks, and an output layer, but the initialized network parameters are different.

[0083] Finally, iterative improvement: parameter optimization based on the policy gradient method:

[0084] A policy gradient-based reinforcement learning algorithm is used to optimize the parameters of the Actor and Critic networks. During training, the policy gradient-based reinforcement learning method iteratively improves the scheduling policy by estimating the gradient of the expected reward based on the satellite state and policy parameters. However, due to factors such as parameter initialization and data noise, the expected reward may be overestimated, leading to gradient estimation errors and thus slowing down the convergence speed of the scheduling policy. Therefore, the policy gradient method proposed in this paper uses two Critic networks, and their minimum value is used to estimate the expected reward for any given state.

[0085] θ, φ1, φ2 are the parameters of the Actor network and the two Critic networks, respectively, J(π θ ) represents the scheduling policy π θ Expected return without discount, D k Represents the set of scheduling decision trajectories; expected return J(π) θ The gradient of ) is calculated as follows:

[0086]

[0087] Where τ represents the scheduling decision trajectory in the decomposed subtask. For scheduling strategy π θ The advantage function represents the advantage of a scheduling strategy relative to the average value given a given state, and is calculated using the following formula:

[0088]

[0089] Where, r k This indicates the k-th scheduling cycle. This represents the output of the i-th Critic network;

[0090] Critic networks are used to approximate the state value V. π (s T,k ), that is, in state s T,k The following is based on the scheduling strategy π θThe average return; the update objective of the dual-Critic network is to minimize the mean square error of the approximate objective:

[0091]

[0092] in, The formula for calculating the cumulative future return is as follows:

[0093]

[0094] Step 3: Online decision-making, which involves using a trained network for real-time relay observation of dynamic targets:

[0095] After completing the offline learning process through steps 1 and 2, the parameters of the entire task decision network for dynamic target observation are determined. At this point, in each scheduling cycle, the Actor network estimates the future cumulative revenue of the allocation strategy based on the visibility information of different types of satellites for the space target to be observed, determines the type of observation satellite and the observation time window, and determines the next cycle's relay satellite and observation window at the end of each scheduling cycle, thus realizing online real-time task planning until the entire observation of the dynamic target is completed.

[0096] In this invention, to verify the effectiveness of the proposed dynamic task decision-making algorithm, a simulation scenario was designed involving 24 observation satellites observing 10 dynamic targets. The algorithm was then tested through simulation. The 24 heterogeneous observation satellites were distributed according to a Walker constellation configuration, with four evenly distributed orbital planes containing six satellites each. The orbital altitude was 1000 km, and the orbital inclinations were 57° and 100°. Both electronic reconnaissance and optical satellites were included. The 16 dynamic observation targets were randomly generated globally. The constellation multi-target collaborative observation algorithm performed task planning for all observation satellites with a 50-second scheduling cycle.

[0097] In a task decision-making scenario involving 40 dynamic targets, to verify the convergence of the task decision-making algorithm proposed in this patent, the algorithm was run to obtain task allocation results. The satellite constellation's allocation of observation time windows for the 40 dynamic flying targets is as follows: Figures 6-9 As shown, the algorithm first groups the target satellites, with a large scheduling cycle of 250 seconds. It focuses on one group of target satellites and performs a task allocation for the observation satellites in a scheduling cycle of 50 seconds.

[0098] Furthermore, the algorithm proposed in this patent also demonstrates good scheduling performance in terms of double observation coverage for dynamic targets. The double observation coverage for 40 observed targets is as follows: Figure 10 As shown, the average double observation coverage rate is 91.6%.

[0099] The above description is merely a preferred embodiment of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described or explained in this invention are implemented according to conventional methods in the art unless otherwise specified or limited.

Claims

1. A heterogeneous constellation intelligent task decision method for space anomaly target observation, characterized in that: The method comprises the following steps: Step 1: a multi-constraint multi-objective optimization model for dynamic target observation is established, and a Markov dynamic decision model is designed, that is, the elements of the state set of the task decision model are determined: First, according to the global distribution characteristics of the heterogeneous star cluster and the dynamic target, a task decomposition framework is determined, the relay observation process of multiple dynamic targets is divided into multiple continuous sub-tasks, each sub-task is solved respectively, and the relative motion relationship between the target and the observation satellite, the unobservable blind area, the double coverage, and the communication link constraint condition are considered, so that the continuous relay observation of the target is maximized as the objective function, a dynamic star cluster task decision model is established, and the observation sequence of each satellite and the handover relationship are determined; Secondly, an optimization model for the dynamic target observation task is established based on the double coverage rate of the overall operation process of the dynamic target, the switching times of the heterogeneous satellites and the comprehensive optimization target of the observation quality, and a Markov decision model for the relay observation task is established on this basis; Step 2: an intelligent decision algorithm architecture of the heterogeneous star cluster based on experience learning-Mask mechanism-iterative improvement is designed, that is, the offline training of the heterogeneous star cluster task decision algorithm is realized: Based on the established star cluster task allocation model, the observation resources and satellite communication capabilities of different satellites are considered, and an intelligent allocation solution strategy based on experience learning-Mask mechanism-iterative improvement-online decision is designed, wherein the Mask mechanism is used to process the different selection sequences of the electronic detection and optical heterogeneous satellites, the historical experience data provided by the simulation platform and the online network adjustment are combined, and a dynamic target continuous relay observation offline training strategy considering real-time performance and optimality is obtained; Step 3: online decision, that is, the trained network is used for real-time relay observation of the dynamic target: After the offline learning process in steps 1 and 2 is completed, the parameters of the task decision network for the dynamic target observation are determined, at this time, in each scheduling period, the Actor network estimates the future cumulative benefits of the allocation strategy according to the visibility information of different types of satellites for the space target to be observed, determines the observation satellite type and the observation time window, and determines the relay satellite and the observation window in the next period at the end of each scheduling period, realizes online real-time task planning, and completes the whole process observation of the dynamic target. 2.The heterogeneous constellation intelligent task decision method for space moving target observation according to claim 1, characterized in that: In step 1, during the optimization model establishment process: the observation task is decomposed, the life cycle of the dynamic target is divided into a series of continuous time periods, that is, scheduling periods; then in each scheduling period, the scheduling result of the observation satellite is given, and the observation satellite in the scheduling result of the next scheduling period performs relay observation on the dynamic target when the current scheduling period ends.

3. The heterogeneous constellation intelligent task decision method for space moving target observation according to claim 2, characterized in that: In step 1, first, according to the observation characteristics of the constellation satellites on the space dynamic target in each scheduling period, the following six constraint conditions are established: a. Deep space background observation constraint: low-orbit agile satellite adopts deep space background observation to space non-cooperative dynamic target, that is, the observation line of sight of the satellite to the target is higher than the height of the atmosphere of the earth, that is, the observation angle θ formed by the observation line of sight of the satellite to the target and the line connecting the satellite to the earth tar is greater than the limb observation angle θ ob That is: θ tar ≥ θ ob (1-1) The calculation method of the edge observation angle is as follows: wherein R e represents the radius of the Earth, H a represents the height of the atmosphere, H sat represents the altitude of the satellite; b. Sun avoidance constraint: there is a sun avoidance angle constraint, and the target cannot be observed in this period when the constraint is not satisfied; α sun >α ob (1-3) where α sun is the line-of-sight solar angle, α ob is the solar avoidance angle; c. Maximum detection distance constraint: the maximum detection distance depends on the detection capability of the load, the target intensity and the background intensity; d. Heterogeneous satellite observation constraint The optical satellite or SAR satellite must meet the dual observation constraints; e. Relay observation constraints When the same non-cooperative dynamic target is observed by the satellite combination, the next group of observation satellites must complete the preliminary positioning of the target to be observed, capture the target to be observed, and be ready to start relay observation before the window of the previous group of observation satellites releases the observation resources. After the previous group of observation satellites releases the observation resources, the relay observation combination can immediately enter the monitoring state to complete the positioning of the target to be observed; wherein, denotes the start time of dynamic target j at the beginning of the kth scheduling period, denotes the length of the double observation window of dynamic target j, Targets denotes the set of targets; f. Communication link constraints On the basis of meeting the detection capability, a communication link constraint mechanism is introduced to comprehensively evaluate the communication performance of various types of payload satellites, dynamically adjust the task allocation strategy, and ensure that the target information can be effectively transmitted within the time effectiveness requirement; Secondly, the target function of dynamic target observation is determined: The goal of the star cluster dynamic task planning for threat monitoring tasks is to complete the dual coverage observation of space non-cooperative dynamic targets and ensure that the system has high work efficiency and observation conditions while tracking and positioning more dynamic targets; In the formula, C j (T k ) represents the double coverage of the dynamic target j in the kth scheduling period, represents the double coverage of all J dynamic targets in the cumulative K scheduling periods, represents the cumulative switching times of the N observation satellites, represents the cumulative observation factors of the N observation satellites, which are related to the relative motion trend of the target to be observed and the observation satellite, and reflect the observation condition.

4. The heterogeneous constellation intelligent task decision method for space moving target observation according to claim 3, characterized in that, A Markov decision model is established for the observation of dynamic targets by heterogeneous satellites, including the following steps: Firstly, the state set S is established; in the tracking relay observation process of dynamic targets, the visibility of satellites to targets and the relative motion characteristics of targets will affect the observation effect of targets; therefore, the distance between the observation satellite and the target, the relative motion between the satellite and the target, the observation time window of the satellite and the switching number are taken as the state s T Therefore, the state set can be expressed as follows: wherein, represents the state information of the i-th satellite, dis i,j , rela i,j represent the distance and relative motion relationship between the target and the satellite, cover i,j represents the visible time window of the satellite i to the target j, change i represents the number of switching of the i-th satellite; Secondly, the action set A is established; according to the observation constraints of heterogeneous satellites, the decision center is planned for J dynamic targets to be observed. First, select J electronic reconnaissance satellites for preliminary reconnaissance, and further select 2*J optical observation satellites that can communicate with the confirmed electronic reconnaissance satellites according to the satellite communication link, and assign the corresponding visible time window type as the action a; a = {dian1, dian2,..., dian J , guang1, guang2,..., guang 2*J}, a e A (1-7) where dian j represents an electro-optical satellite observing the jthtarget selection, guang 2*j , guang 2*j-1 represents an optical satellite observing the jthtarget selection; Then, the immediate reward R; in the process of learning the task of cooperative observation of dynamic targets, the optimization target defined above is considered comprehensively. If the coverage rate of dynamic targets is higher while maintaining high system efficiency and observation conditions, the comprehensive observation reward is greater, and the reward is greater. Therefore, the reward is designed as follows: Finally, the discount factor γ is similar to the static task planning, γ represents the importance of future reward value relative to current reward value, 0 < γ < 1; when γ = 0, it is equivalent to only considering the current reward without considering the future reward, and when γ = 1, it means that the future reward and the current reward are equally important.

5. The heterogeneous constellation intelligent task decision method for space moving target observation according to claim 4, characterized in that: The step 2; First, experience learning: data sampling based on expert experience During data collection, the state s is obtained according to the dynamic target observation task demand and is input into the task decision network to output the corresponding operation a, i.e. the satellite allocated for the current task and the corresponding available time window, to obtain the reward r. After executing the above operation, each iteration round will store data in the following form: (s, a, r, s′); Then, Mask mechanism: self-adaptive change according to the selection requirements of the characteristics of heterogeneous satellites; In each scheduling period for dynamic target observation, appropriate satellites need to be selected for observation tasks according to the established constraint conditions and target function. A deep pointer network task decision network architecture based on the Mask mechanism is designed to realize the selection of different satellites and corresponding time windows to ensure real-time relay tracking of dynamic targets. The model of constellation dynamic task planning for dynamic target observation task is based on the Actor-Critic network architecture design; Finally, iterative improvement: parameter optimization based on policy gradient method: The reinforcement learning algorithm based on policy gradient is used to optimize the parameters of Actor network and Critic network. In the training process, the reinforcement learning method based on policy gradient is used to estimate the gradient of the cumulative expected return based on the satellite state and policy parameter, and the scheduling policy is iteratively improved. In the proposed policy gradient method, two Critic networks are used, and the minimum value is used to estimate the expected reward of any given state. Let θ, φ1, φ2 be the parameters of the Actor network and the two Critic networks, respectively, J(π θ ) denotes the expected undiscounted return of the scheduling policy π θ , D k denotes the set of scheduling decision trajectories; the gradient of the expected return J(π θ ) is computed as follows: where τ denotes the trajectory of scheduling decisions in the decomposition subtask, The advantage function of a scheduling policy π θ represents the advantage of the scheduling policy over the average given the state, and is computed as follows: wherein r k denotes the kth scheduling period, denotes the output of the ith Critic network; The Critic network is used to approximate the state value V π (s T,k ), i.e., the average return under the scheduling policy π T,k in state s θ ; the update target of the double Critic network is to minimize the mean square error of the approximated target: wherein, represents the future cumulative return, and the calculation formula is as follows:

6. The heterogeneous constellation intelligent task decision method for space moving target observation according to claim 5, characterized in that: The Actor-Critic network architecture design is introduced as follows: Actor network: Based on the essence of dynamic target observation and the selection requirements of heterogeneous satellites, the Actor network is designed as a deep pointer network architecture based on the Mask mechanism, responsible for real-time calculation of the collaborative task planning results of each scheduling period, and its goal is to approximate the optimal task planning strategy π * is to maximize the cumulative task revenue: where θ * The input of the deep pointer network based on the Mask mechanism is the state sequence of each satellite observing the target in the kth scheduling period, and the output is the observation combination of each target in the current scheduling period, including the selection combination of one electro-optical satellite and two optical or SAR satellites. The overall network structure of the pointer network is composed of an encoder and a decoder, wherein the encoder is composed of a one-dimensional convolutional neural network layer, responsible for feature extraction of the input sequence information s T,k to obtain embedded information The decoder is the core component of the pointer network, which calculates the selection probability of each satellite in the input sequence through the attention mechanism. In the online scheduling process, the satellite with the highest probability value is selected as the observation resource using the greedy strategy until the J observation satellites for the dynamic target in the current scheduling period are selected. In each decoding process, a dynamic Mask mechanism is designed, which sets the probability of the selected satellites and the unselected satellite types to 0 when calculating the selection probability of each satellite in the input sequence. For example, if one electrical reconnaissance satellite is selected, the probabilities of optical satellites or SAR satellites are set to 0. Critic network: In the offline training phase of the constellation dynamic task planning algorithm, a double Critic network structure is used for training. The input is the current state information sequence of the constellation, and the output is the future cumulative expected return under the current state, which is used to evaluate the goodness of the current state.

Citation Information

Patent Citations

  • Civil aviation Internet business management and access resource allocation method based on low-orbit giant satellite base

    CN114900225A

  • Multi-satellite autonomous cooperative scheduling method based on distributed multi-agent reinforcement learning

    CN119623910A