Heterogeneous constellation intelligent task decision-making method for space transaction target observation

By establishing a multi-constraint, multi-objective optimization model and a Markov decision model, and combining it with an intelligent decision-making algorithm based on empirical learning and iterative improvement, the problem that traditional scheduling methods are difficult to meet the needs of observing space-moving targets has been solved. This has enabled efficient relay tracking and observation of dynamic targets and improved space early warning capabilities.

CN120974902AActive Publication Date: 2025-11-18TIANJIN UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511081819.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-18
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

Traditional fixed observation and scheduling methods are difficult to meet the timeliness and accuracy requirements of space-moving targets. How to carry out efficient task allocation and coordination decisions in complex dynamic environments, especially multi-satellite relay observation of high-speed moving targets to improve space early warning capabilities.

Method used

A multi-constraint, multi-objective optimization model was established, a Markov decision model was designed, and an intelligent decision-making algorithm based on experience learning, mask mechanism, and iterative improvement was adopted. Through offline training and online decision-making, real-time mission planning for heterogeneous constellations was achieved, optimizing double coverage, system switching frequency, and observation quality.

Benefits of technology

It enables efficient relay tracking and observation of dynamic targets, improves space early warning capabilities, meets the requirements of real-time performance and optimality, and ensures observation quality and system efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120974902A_ABST
    Figure CN120974902A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of intelligent algorithms, heterogeneous satellite planning and dynamic target monitoring, and particularly relates to a heterogeneous constellation intelligent task decision-making method for spatial transaction target observation, which comprises the following steps of: 1, establishing a multi-constraint multi-target optimization model for dynamic target observation, designing a Markov dynamic decision-making model, and establishing a multi-constraint multi-target optimization model for dynamic target observation; the elements of the state set of the task decision model are determined; 2, designing a heterogeneous star group intelligent decision algorithm architecture based on empirical learning-Mask mechanism-iteration improvement, namely realizing offline training of a heterogeneous star group task decision algorithm; and step 3, performing online decision making, namely performing dynamic target real-time relay observation by using the trained network. The invention provides a universal relay observation task scheduling framework which comprehensively considers the double coverage rate, the system switching frequency and the observation quality, so that the optimal multiple satellites in the heterogeneous star group are selected to perform relay observation on the dynamic target to maintain target positioning, and the space early warning capability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of intelligent algorithm, heterogeneous satellite planning and dynamic target monitoring, and particularly relates to a heterogeneous constellation intelligent task decision method for space dynamic target observation. BACKGROUND

[0002] In the face of complex and changeable space environment and increasingly frequent space dynamic targets, such as rapid debris drift, non-cooperative spacecraft maneuvering, near-earth orbit abnormal targets, etc., the traditional fixed observation scheduling mode has been difficult to meet the timeliness and accuracy requirements of task response. Heterogeneous constellation is composed of multiple satellites with different load types such as optical, radar, electromagnetic, etc. and orbit characteristics, and has stronger all-time and all-space coverage capability and multi-modal observation advantage. Using heterogeneous constellation to realize efficient perception and continuous tracking of space dynamic targets has become an important development direction of current space situation awareness. However, different satellites have differences in observation ability, communication link, energy consumption budget and maneuvering ability, etc. How to make efficient task allocation and coordination decision in complex dynamic environment becomes a key technical difficulty. In addition, compared with the static task planning of ground targets, the dynamic task planning of star group has two very obvious characteristics. The first is that the non-cooperative target is usually a high-speed aircraft or a dynamic satellite, which is in a high-speed motion state, and the motion trend of the target can almost be ignored in the static task planning of ground targets. The second is that multiple observation targets and our observation satellites have high space-time coupling characteristics. In the time scale, our satellites need to complete continuous relay tracking observation of the target. For the same non-cooperative dynamic target, multiple different observation combinations are usually needed to complete relay, and frequent switching of observation combinations will reduce the efficiency of the system. In the space scale, there are multiple visible observation satellites for the same target in a scheduling period, and the combination decision of observation resources is needed. At the same time, due to the high-speed motion characteristics of the target, the relative motion trend of different satellites and the target is different, which further leads to the difference of observation quality, and the combination decision of observation resources is also needed to improve the observation quality.

[0003] Therefore, the application provides a heterogeneous constellation intelligent task decision method for space dynamic target observation. The application focuses on the star group task decision problem for dynamic target observation task, proposes a general relay observation task scheduling framework which comprehensively considers double coverage rate, system switching frequency and observation quality, and realizes relay observation of dynamic target by selecting the optimal multiple satellites in the heterogeneous star group to maintain target positioning and improve space early warning capability. SUMMARY

[0004] The application aims to provide a heterogeneous constellation intelligent task decision method for space dynamic target observation, and considers the relay observation problem of globally distributed and high-speed moving dynamic targets.

[0005] The technical scheme adopted by the application is specifically as follows:

[0006] A heterogeneous constellation intelligent task decision method for space dynamic target observation comprises the following steps:

[0007] Step 1: a multi-constraint multi-objective optimization model for dynamic target observation is established, and a Markov dynamic decision model is designed, that is, the state set elements of the task decision model are determined.

[0008] First, the task decomposition framework is determined according to the heterogeneous constellation and the global distribution characteristics of the dynamic target, the relay observation process for multiple dynamic targets is divided into multiple continuous subtasks, each subtask is solved, and the relative motion relationship between the target and the observation satellite, the unobservable blind area, the double coverage, the communication link constraint condition are considered, the maximum completion of the continuous relay observation of the target is taken as the objective function, the dynamic constellation task decision model is established, and the order and the handover relationship of the target observed by each satellite are determined.

[0009] Secondly, the double coverage rate of the whole operation process of the dynamic target, the switching times of the heterogeneous satellites and the comprehensive optimization target of the observation quality are taken as the optimization model for the dynamic target observation task, and the Markov decision model for the relay observation task is established.

[0010] Step 2: the intelligent decision algorithm architecture of the heterogeneous constellation based on experience learning-Mask mechanism-iterative improvement is designed, that is, the offline training of the heterogeneous constellation task decision algorithm is realized.

[0011] Based on the established star cluster task allocation model, considering different satellite observation resources and satellite communication capabilities, an intelligent allocation solving strategy based on experience learning-mask mechanism-iterative improvement-online decision is designed, wherein the mask mechanism is used for processing different selection orders of heterogeneous satellites such as electronic reconnaissance and optical satellites, and combined with historical experience data provided by the simulation platform and online network adjustment, an offline training strategy for dynamic target continuous relay observation considering real-time and optimality is obtained.

[0012] Step 3: online decision, that is, using the trained network to perform dynamic target real-time relay observation:

[0013] After completing the offline learning process through steps 1 and 2, the task decision network parameters for dynamic target observation are determined, at this time, in each scheduling period, the actor network estimates the future cumulative income of the allocation strategy according to the visibility information of different types of satellites for the space target to be observed, determines the observation satellite type and the observation time window, and determines the relay satellite and the observation window in the next period at the end of each scheduling period, realizes online real-time task planning, and completes the whole process observation of the dynamic target.

[0014] The technical effects obtained by the application are:

[0015] The application proposes a general relay observation task scheduling framework considering double coverage, system switching times, heterogeneous satellite characteristics, communication links and observation quality for the star cluster task decision problem of space dynamic target observation task, divides the dynamic target relay observation task into a series of subtasks for solving.

[0016] The application proposes a flexible and adaptive experience learning-mask mechanism-iterative improvement heterogeneous star cluster intelligent task decision algorithm for the dynamic target continuous relay observation task solving problem, designs a dynamic mask mechanism, which can better handle the actual engineering problem of action space change caused by different selection quantities of electronic reconnaissance, optical and other heterogeneous satellites in the relay observation process, and combines the deep pointer network to realize real-time dynamic selection of heterogeneous satellites, meets the demand of "plug and play", and can guarantee real-time optimal satellite selection. BRIEF DESCRIPTION OF DRAWINGS

[0017] Figure 1 is the overall technical scheme in the application;

[0018] Figure 2 is a sun avoidance schematic diagram in the application;

[0019] Figure 3 is a heterogeneous satellite observation schematic diagram in the application;

[0020] Figure 4is a deep pointer network calculation process diagram in the present application;

[0021] Figure 5 is a Critic network structure diagram in the present application;

[0022] Figure 6 is a 0-250 second task allocation result diagram in the present application;

[0023] Figure 7 is a 250-500 second task allocation result diagram in the present application;

[0024] Figure 8 is a 500-750 second task allocation result diagram in the present application;

[0025] Figure 9 is a 750-1000 second task allocation result diagram in the present application;

[0026] Figure 10 is a double coverage diagram of an observed target in the present application. DETAILED DESCRIPTION

[0027] In order to make the purpose and advantages of the present application more clear and obvious, the present application is specifically described below in combination with embodiments. It should be understood that the following text is only used to describe one or several specific embodiments of the present application, and does not strictly limit the specific protection scope requested by the present application.

[0028] As shown in Figure 1 A heterogeneous constellation intelligent task decision method for space moving target observation, the present application relates to the fields of intelligent algorithm, heterogeneous satellite planning and dynamic target monitoring, and particularly relates to a heterogeneous constellation intelligent task decision method for space moving target observation, and specifically relates to an intelligent decision algorithm architecture adopting experience learning-Mask mechanism-iterative improvement-online decision to realize real-time task planning of heterogeneous constellation satellites, including the following steps:

[0029] Step 1: Establish a multi-constraint multi-objective optimization model for dynamic target observation, design a Markov dynamic decision model, i.e. determine the elements of the state set of the task decision model:

[0030] Firstly, according to the global distribution characteristics of the heterogeneous star group and the dynamic target, the task decomposition framework is determined, the relay observation process of multiple dynamic targets is divided into multiple continuous sub-tasks, each sub-task is solved respectively, and the relative motion relationship between the target and the observation satellite, the unobservable blind area, the double coverage, the communication link constraint condition are considered, the maximum completion of the continuous relay observation of the target is taken as the objective function, the dynamic star group task decision model is established, and the order and the handover relationship of the target observed by each satellite are determined;

[0031] Secondly, an optimization model for dynamic target observation task is established with the comprehensive optimization target of double coverage of the whole operation process of dynamic target, heterogeneous satellite switching times and observation quality, and a Markov decision model for relay observation task is established on this basis.

[0032] Preferably, in the step 1, in the optimization model establishment process: in the tracking observation process of multiple targets, since the targets have the characteristics of high-speed motion, the spatial positions of the targets and the observation satellites are constantly changing, and the visible windows are also constantly changing, therefore, it is very complex to directly give the observation results of the whole life cycle of multiple dynamic targets, and it is necessary to decompose the observation task, divide the life cycle of the dynamic target into a series of continuous time periods, i.e. scheduling periods; then in each scheduling period, the scheduling result of the observation satellite is given, and when the current scheduling period ends, the observation satellite in the scheduling result of the next scheduling period performs relay observation on the dynamic target.

[0033] Preferably, in the step 1, first, the observation characteristics of the constellation satellites on the space dynamic target in each scheduling period are established, and the following six constraint conditions are established:

[0034] a. Deep space background observation constraint: the low-orbit agile satellite adopts deep space background observation on the space non-cooperative dynamic target, i.e. the observation line of sight of the satellite to the target is higher than the height of the earth's atmosphere, i.e. the observation angle θ formed by the observation line of sight of the satellite to the target and the line connecting the satellite to the center of the earth is greater than the limb observation angle θ tar ob , i.e.

[0035] θ tar ≥ θ ob (1-1)

[0036] The calculation method of the limb observation angle is as follows:

[0037]

[0038] Wherein, R e represents the radius of the earth, H a represents the height of the atmosphere, and H sat represents the height of the satellite.

[0039] b. Sun avoidance constraint: the satellite needs to avoid the influence of sunlight on the detection process when performing relay observation, and there is a sun avoidance angle constraint, when the constraint is not met, the target cannot be observed in this period of time;

[0040] α sun > α ob (1-3)

[0041] Wherein, α sun is the line-of-sight solar angle, and α ob ​The sun avoidance angle; see definition Figure 2 as shown;

[0042] c. Maximum detection distance constraint: When the non-cooperative dynamic target satisfies the deep space background observation constraint for the observation satellite, the maximum detection distance constraint of the optical load for the non-cooperative dynamic target also needs to be satisfied. The maximum detection distance depends on the detection capability of the load, the target intensity, and the background intensity.

[0043] d. Observation constraint of heterogeneous satellites

[0044] In order to obtain the position and velocity information of the dynamic target, first, a suitable electro-satellite needs to be determined according to the observation characteristics of the heterogeneous satellites to determine the approximate position information of the dynamic target. Further, at least two optical satellites or SAR satellites need to simultaneously perform stereoscopic positioning on the moving target to realize relay observation in the scheduling period, that is, the optical satellites or SAR satellites need to satisfy the double observation constraint; as shown, Figure 3 ,

[0045] e. Relay observation constraint

[0046] When the same non-cooperative dynamic target needs to be observed by the switching of the observation satellite combination, it is necessary to ensure that the next group of observation satellites has completed the preliminary positioning of the to-be-observed target and captured the to-be-observed target before the window release of the observation resources of the previous group of observation satellites, and is ready to start relay observation. After the observation resources of the previous group of observation satellites are released, the relay observation combination can immediately enter the monitoring state to complete the positioning of the to-be-observed target.

[0047]

[0048] wherein, denotes the start time of the dynamic target j in the kth scheduling period, denotes the time length of the double observation window of the dynamic target j, and Targets denotes the target set.

[0049] f. Communication link constraint

[0050] For heterogeneous satellites with different loads, such as electro-satellites, optical satellites, or SAR satellites, their communication capabilities, data transmission rates, and link durations differ significantly. In task scheduling, satellites with stable communication links that can realize real-time or quasi-real-time backhaul of observation data should be preferentially selected for observation tasks. In particular, in the scenario of rapid response observation of non-cooperative dynamic targets, if some optical satellites have observation conditions but the communication link is interrupted or delayed seriously, it will affect the data transmission efficiency and subsequent decision-making effectiveness. Therefore, on the basis of satisfying the detection capability, a communication link constraint mechanism needs to be introduced to comprehensively evaluate the communication performance of various load satellites and dynamically adjust the task allocation strategy to ensure that the target information can be effectively transmitted within the time effectiveness requirement.

[0051] Secondly, the target function of dynamic target observation is determined:

[0052] The goal of constellation dynamic task planning for threat monitoring task is to complete double coverage observation of space non-cooperative dynamic targets, and ensure that the system has high work efficiency and observation conditions at the same time to track and locate more dynamic targets;

[0053]

[0054] In the formula, C j (T k ) represents the double coverage rate of dynamic target j in the kth scheduling period, represents the double coverage rate of all J dynamic targets in the cumulative K scheduling periods, represents the cumulative switching times of N observation satellites, represents the cumulative observation factor of N observation satellites, which is related to the relative motion trend of the target to be observed and the observation satellite, and reflects the observation condition, and α1, α2, α3 represent the weight coefficients.

[0055] Preferably, a Markov decision model for dynamic target observation by heterogeneous satellites is established, including the following steps:

[0056] Firstly, a state set S is established; in the tracking relay observation process of dynamic targets, the visibility and relative motion characteristics of the satellite to the target will affect the target observation effect; therefore, the distance between the observation satellite and the target, the relative motion between the satellite and the target, the observable time window of the satellite and the switching times are taken as the state s T Therefore, the state set can be expressed as follows:

[0057]

[0058] In the formula, represents the state information of the ith satellite, dis i,j ,rela i,j respectively represent the distance and relative motion relationship between the target and the satellite, cover i,j represents the visible time window of satellite i to target j, change i represents the switching times of the ith satellite;

[0059] Secondly, the action set A is established; according to the observation constraints of heterogeneous satellites, the planning decision center is J dynamic targets to be observed, first select J electronic reconnaissance satellites for preliminary reconnaissance, and further select 2*J optical observation satellites which can communicate with the confirmed electronic reconnaissance satellites according to the satellite communication link, and assign the corresponding visible time window category as the action a;

[0060] a = {dian1, dian2,..., dian J , guang1, guang2,..., guang 2*J}, a e A (1-7)

[0061] where dian j represents the observation of the jth target selection of the electronic reconnaissance satellite, guang 2*j , guang 2*j-1 represents the observation of the jth target selection of the optical satellite;

[0062] Then, the immediate reward R; in the process of learning task for dynamic target cooperative observation, the optimization objective defined above is considered comprehensively, if the coverage of dynamic target is higher while maintaining a high system efficiency and observation conditions, the comprehensive observation reward is greater, the reward is greater, therefore the reward design is as follows:

[0063]

[0064] Finally, the discount factor γ is similar to the static task planning, γ represents the importance of future income value relative to the current income value, 0 < γ < 1; when γ = 0, it is equivalent to only considering the current income without considering the future income, γ = 1, then it means that the future income and the current income are equally important.

[0065] Step 2: design an intelligent decision algorithm architecture of heterogeneous star group based on experience learning-Mask mechanism-iterative improvement, that is, realize the offline training of the heterogeneous star group task decision algorithm:

[0066] Based on the established star group task allocation model, considering different satellite observation resources and satellite communication capabilities, an intelligent allocation solution strategy based on experience learning-Mask mechanism-iterative improvement-online decision is designed, wherein the Mask mechanism is used to process different selection orders of electronic reconnaissance, optical and other heterogeneous satellites, combined with the historical experience data provided by the simulation platform and online network adjustment, the dynamic target continuous relay observation offline training strategy considering real-time and optimality is obtained;

[0067] Preferably, in step 2, considering the different observation functions of the heterogeneous star group, such as electronic reconnaissance satellites, optical satellites, etc., in the process of observing dynamic targets, due to the dynamic running nature of the observed targets, it is necessary to jointly relay observation by electronic reconnaissance satellites, optical satellites and SAR satellites, etc. According to the observation constraint conditions of the heterogeneous satellites established in the foregoing, an electronic reconnaissance satellite needs to be selected in each scheduling observation period, and optical satellites and SAR satellites need to meet the double observation requirements. Therefore, this will cause the action space of the intelligent task decision network to change according to the selection of satellites, and the traditional intelligent decision network is difficult to solve such problems. Therefore, the patent proposes a heterogeneous star group intelligent decision algorithm based on experience learning-Mask mechanism-iterative improvement;

[0068] First, experience learning: data sampling based on expert experience

[0069] In the data collection process, the state s is obtained according to the observation task demand of the dynamic target, and is sent into the task decision network to output the corresponding operation a, that is, the satellite allocated for the current task including an electronic reconnaissance satellite or an optical satellite or a SAR satellite and the corresponding available time window, to obtain the reward r; After executing the above operation, each iteration round will store the data in the following form: (s, a, r, s');

[0070] This process needs a large number of rounds to obtain enough experience data, which is all stored in the replay experience pool itself; Therefore, the patent adopts the method of collecting and training the network at the same time, and collects the size of the batch size of the self decision data in the experience pool itself with a probability p at the beginning of training, and samples in the expert experience pool with a probability of 1-p, The decision data in the expert experience pool has high reference value, thereby improving the learning speed and stability of the training of the algorithm; With the increase of the number of training times, further increase p gradually, gradually use the data of itself for learning, to improve the adaptability of the network itself;

[0071] Then, Mask mechanism: self-adaptive change according to the selection requirements of the characteristics of heterogeneous satellites;

[0072] In each scheduling period of observing dynamic targets, appropriate satellites need to be selected to execute observation tasks according to the established constraint conditions and objective functions, which is ultimately a permutation and combination problem. Therefore, a star group task decision network architecture based on a deep pointer network is designed. Further, considering the engineering requirement that the action space changes over time due to the different characteristics of heterogeneous satellites, the deep pointer network is further improved, and a deep pointer network star group task decision network architecture based on Mask mechanism is designed to realize the selection of different satellites and corresponding time windows, so as to ensure real-time relay tracking of dynamic targets;

[0073] The model of constellation dynamic task planning for dynamic target observation task is based on the Actor-Critic network architecture, which will be introduced below.

[0074] The Actor-Critic network architecture is based on the Actor-Critic network architecture, which will be introduced below.

[0075] Actor network:

[0076] Based on the essence of dynamic target observation and the selection requirements of heterogeneous satellites, the Actor network is designed as a deep pointer network architecture based on Mask mechanism, which is responsible for real-time calculation of the cooperative task planning results of each scheduling period. Its goal is to approximate the optimal task planning strategy π * which maximizes the cumulative task revenue:

[0077]

[0078] where θ * represents the network parameters of the optimal strategy. The input of the deep pointer network based on Mask mechanism is the state sequence of each type of satellite for target observation in the kth scheduling period, and the output is the observation combination for each target in the current scheduling period, including the selection combination of one electro-optical satellite and two optical or SAR satellites.

[0079] The overall network structure of the pointer network is shown in Figure 4 The overall network structure of the pointer network is composed of an encoder and a decoder. The encoder is composed of a one-dimensional convolutional neural network, which is responsible for feature extraction of the input sequence information s T,k to obtain embedded information

[0080] The decoder is the core component of the pointer network, which calculates the selection probability of each satellite in the input sequence through the attention mechanism. In the online scheduling link, the satellite with the maximum probability value is selected as the observation resource at each time by using the greedy strategy, until the observation satellites corresponding to the J dynamic targets in the current scheduling period are selected. Considering the uniqueness observation constraint and the observation constraint of heterogeneous satellites, i.e. each satellite can only observe one target at the same time, and the different characteristics of heterogeneous satellites require the action space to change with time, a dynamic Mask mechanism is designed in each decoding process, i.e. the probability of the selected satellite and the unselected satellite type is set to 0 when calculating the selection probability of each satellite in the input sequence. For example, if one electro-optical satellite is selected, the probabilities of optical satellites or SAR satellites are set to 0, so as to avoid repeated selection of the same satellite and meet the selection requirements of the characteristics of heterogeneous satellites.

[0081] Critic network:

[0082] Critic network structure Figure 5 As shown, in the offline training phase of the constellation dynamic task planning algorithm, a double Critic network structure is used for training, the input is the current state information sequence of the constellation, and the output is the future cumulative expected return under the current state, which is used to evaluate the goodness of the current state. The double Critic network architecture is completely the same, both have an input layer, four layers of one-dimensional convolutional neural network, and an output layer composition. The initialized network parameters are not the same.

[0083] Finally, iterative improvement: parameter optimization based on policy gradient method:

[0084] The policy gradient reinforcement learning algorithm is used to optimize the parameters of the Actor network and the Critic network. In the training process, the policy gradient reinforcement learning method is based on the gradient of the estimated cumulative expected return of the satellite state and the policy parameter, and iteratively improves the scheduling policy. However, due to factors such as parameter initialization and data noise, the expected return may be overestimated, resulting in gradient estimation error, which slows down the convergence speed of the scheduling policy. Therefore, two Critic networks are used in the policy gradient method proposed in this paper, and the minimum value is used to estimate the expected reward of any given state.

[0085] θ, φ1, φ2 are the parameters of the Actor network and the two Critic networks respectively, J(π θ ) represents the expected undiscounted return of the scheduling policy π θ , D k represents the set of scheduling decision trajectories; The gradient of the expected return J(π θ ) is calculated as follows:

[0086]

[0087] Where τ represents the scheduling decision trajectory in the decomposition subtask, is the advantage function of the scheduling policy π θ ; The advantage function represents the advantage of the scheduling policy relative to the average value under a given state, and the calculation formula is as follows:

[0088]

[0089] Where r k represents the kth scheduling period, represents the output of the ith Critic network;

[0090] The Critic network is used to approximate the state value V π (s T,k ), that is, according to the scheduling policy π T,k under the state s θThe average return of the double Critic network is the update target of minimizing the mean square error of the approximation target:

[0091]

[0092] wherein, represents the future cumulative return, and the calculation formula is as follows:

[0093]

[0094] Step 3: online decision, that is, using the trained network to perform real-time relay observation of dynamic targets:

[0095] After the offline learning process is completed through steps 1 and 2, the entire task decision network parameter facing dynamic target observation is determined, at this time, in each scheduling period, the Actor network estimates the future cumulative return of the allocation strategy according to the visibility information of different types of satellites for the space target to be observed, determines the observation satellite type and the observation time window, and determines the relay satellite and the observation window in the next period at the end of each scheduling period, realizes online real-time task planning, and until the whole observation of the dynamic target is completed.

[0096] In the application, in order to verify the effectiveness of the dynamic task decision algorithm proposed in the application, a simulation scene of 24 observation satellites observing 10 dynamic targets is designed, and simulation tests are performed on the algorithm, wherein the 24 heterogeneous observation satellites are distributed according to the Walker constellation configuration, there are 4 orbit planes with 6 evenly distributed satellites, the orbit height is 1000km, the orbit inclination is 57° and 100°, and there are two types of satellites, namely, electronic reconnaissance satellites and optical satellites; 16 dynamic observation targets are randomly generated in the global range. The star cluster multi-target cooperative observation algorithm performs task planning on all observation satellites with a scheduling period of 50 seconds.

[0097] In the task decision scene of 40 dynamic targets, in order to verify the convergence of the task decision algorithm proposed in the application, the task allocation result is obtained by running the algorithm, and the observation time window allocation result of the satellite cluster to the 40 dynamic flying targets is as shown in Figure 6-9 The algorithm first groups the target satellites, and every 250 seconds is a large scheduling period, and a group of target satellites is focused on, and every 50 seconds is a scheduling period, and the observation satellites are allocated once.

[0098] In addition, the algorithm proposed in the application also has good scheduling effect in the double observation coverage of dynamic targets. The double observation coverage of the 40 observation targets is as shown in Figure 10 The average double observation coverage is 91.6%.

[0099] The above merely describes the preferred embodiments of the present application, and it should be pointed out that those skilled in the art can make several improvements and refinements without departing from the principles of the present application, and these improvements and refinements should also be considered as the protection scope of the present application. The structures, devices and operation methods not specifically described and explained in the present application are implemented according to the conventional means in the art, unless specifically described and limited.

Claims

1. A heterogeneous constellation intelligent mission decision-making method for observing space-moving targets, characterized in that: Includes the following steps: Step 1: Establish a multi-constraint, multi-objective optimization model for dynamic target observation, and design a Markov dynamic decision model, i.e., determine the various elements of the state set of the task decision model: First, considering the global distribution characteristics of heterogeneous satellite constellations and dynamic targets, a task decomposition framework is determined. The relay observation process of multiple dynamic targets is divided into multiple consecutive sub-tasks, and solutions are performed in each sub-task. The relative motion relationship between the target and the observation satellite, the unobservable blind zone, double coverage, and communication link constraints are considered. The objective function is to maximize the completion of continuous relay observation of the target. A dynamic satellite constellation task decision model is established to determine the order of observation of the target by each satellite and the handover relationship. Secondly, with the comprehensive optimization objectives of double coverage, number of heterogeneous satellite switching and observation quality for the overall operation of dynamic targets, an optimization model for dynamic target observation tasks was established, and on this basis, a Markov decision model for relay observation tasks was established. Step 2: Design a heterogeneous constellation intelligent decision-making algorithm architecture based on experience learning, mask mechanism, and iterative improvement, i.e., realize offline training of the heterogeneous constellation task decision-making algorithm: Based on the established constellation task allocation model, considering different satellite observation resources and satellite communication capabilities, an intelligent allocation solution strategy based on experience learning, Mask mechanism, iterative improvement, and online decision-making is designed. The Mask mechanism is used to handle different selection orders of heterogeneous satellites such as electronic reconnaissance and optical satellites. Combined with historical experience data provided by the simulation platform and online network adjustment, a dynamic target continuous relay observation offline training strategy that balances real-time performance and optimality is obtained. Step 3: Online decision-making, which involves using a trained network for real-time relay observation of dynamic targets: After completing the offline learning process through steps 1 and 2, the parameters of the entire task decision network for dynamic target observation are determined. At this point, in each scheduling cycle, the Actor network estimates the future cumulative revenue of the allocation strategy based on the visibility information of different types of satellites for the space target to be observed, determines the type of observation satellite and the observation time window, and determines the next cycle's relay satellite and observation window at the end of each scheduling cycle, thus realizing online real-time task planning until the entire observation of the dynamic target is completed.

2. The heterogeneous constellation intelligent mission decision-making method for observing space-moving targets according to claim 1, characterized in that: In step 1, during the optimization model establishment process: the observation task is decomposed, and the life cycle of the dynamic target is divided into a series of continuous time periods, i.e., scheduling cycles; then, within each scheduling cycle, the scheduling results of the observation satellites are given, and when the current scheduling cycle ends, the observation satellites in the scheduling results of the next scheduling cycle take over the observation of the dynamic target.

3. The heterogeneous constellation intelligent mission decision-making method for observing space-moving targets according to claim 2, characterized in that: In step 1, the following six constraints are first established based on the observation characteristics of constellation satellites of dynamic space targets in each scheduling cycle: a. Constraints of Deep Space Background Observation: Low-Earth Orbit agile satellites employ deep space background observation for non-cooperative dynamic targets in space. This means that the satellite's line of sight to the target must be higher than the Earth's atmospheric altitude, specifically the observation angle θ formed by the satellite's line of sight and the line connecting the satellite to the Earth's center. tar It must be greater than the adjacent observation angle θ ob ,Right now: i tar ≥ θ ob (1-1) The calculation method for the near-edge observation angle is as follows: Among them, R e H represents the Earth's radius. a H represents the altitude of the atmosphere. sat Indicates the satellite's altitude; b. Solar avoidance constraint: There is a solar avoidance angle constraint. When this constraint is not met, the target cannot be observed during this period. α sun >α ob (1-3) Where α sun For the line-of-sight solar angle, α ob For solar avoidance angle; c. Maximum detection range constraint: The maximum detection range depends on the detection capability of the payload, the target strength, and the background strength; d Heterogeneous satellite observation constraints Optical satellites or SAR satellites must meet dual observation constraints; e. Relay observation constraints When switching observation satellite combinations for the same non-cooperative dynamic target, it is necessary to ensure that before the observation window of the previous group of observation satellites releases the observation resources, the next group of observation satellites has completed the preliminary positioning of the target to be observed, captured the target to be observed, and is ready to start relay observation at any time. After the previous group of observation satellites releases the observation resources, the relay observation combination can immediately enter the monitoring state and complete the positioning of the target to be observed. in, This represents the start time of dynamic objective j in the k-th scheduling cycle. The duration of the double observation window for dynamic target j is represented by Targets, which represents the set of targets. f. Communication link constraints On the basis of meeting the detection capabilities, a communication link constraint mechanism needs to be introduced to comprehensively evaluate the communication performance of various payload satellites, dynamically adjust the task allocation strategy, and ensure that target information can be effectively transmitted within the time requirements. Secondly, determine the objective function for dynamic target observation: The goal of constellation dynamic mission planning for threat monitoring is to achieve double-coverage observation of non-cooperative dynamic targets in space, ensuring that the system has high working efficiency and observation conditions while tracking and locating more dynamic targets. In the formula, C j (T k () represents the double coverage rate of dynamic target j in the k-th scheduling period. This represents the double coverage rate of all J dynamic targets over a cumulative K scheduling cycles. This represents the cumulative number of switches among N observation satellites. This represents the cumulative observation factor of N observation satellites, which is related to the relative motion trend of the target to be observed and the observation satellites, and reflects the quality of the observation conditions. α1, α2, and α3 represent weighting coefficients.

4. The heterogeneous constellation intelligent mission decision-making method for observing space-moving targets according to claim 3, characterized in that, The establishment of a Markov decision model for observing dynamic targets from heterogeneous satellites includes the following steps: First: Establish a state set S; during the relay observation of a dynamic target, the satellite's visibility to the target and its relative motion characteristics both affect the target observation effect; therefore, the distance between the observing satellite and the target, the relative motion between the satellite and the target, the observable time window of the satellite, and the number of switching are used as state s. T Then the state set can be represented as follows: In the formula, Dis represents the status information of the i-th satellite. i,j ,rela i,j These represent the distance between the target and the satellite, and their relative motion, respectively. i,j This represents the visible time window of satellite i relative to target j, change i This represents the number of times the i-th satellite has been switched. Next, establish: establish action set A; based on the heterogeneous satellite observation constraints, plan the decision center as J dynamic targets to be observed, first select J electronic reconnaissance satellites for preliminary reconnaissance, and further select 2*J optical observation satellites that can communicate with the confirmed electronic reconnaissance satellites according to the satellite communication link, and assign the corresponding visible time window type as action a. a = {dian1, dian2, ..., dian} J ,guang1,guang2,...,guang 2*J }, a∈A (1-7) Among them dian j This represents the electronic reconnaissance satellite selected to observe the j-th target. 2*j ,guang 2*j-1 This indicates the optical satellite selected for observing the j-th target; Then, the immediate benefit is R; in the learning process of collaborative observation tasks for dynamic targets, considering the optimization objectives defined above, if the coverage of the dynamic target is higher while maintaining high system efficiency and observation conditions, the overall observation benefit will be greater, and the reward will be larger. Therefore, the reward design is as follows: Finally, the discount factor γ is similar to that in static task planning. γ represents the importance of future returns relative to current returns, with 0 < γ < 1. When γ = 0, it is equivalent to considering only current returns and not future returns. When γ = 1, it means that future returns and current returns are considered equally important.

5. A heterogeneous constellation intelligent mission decision-making method for observing space-moving targets according to claim 4, characterized in that: In step 2; First, experience-based learning: data sampling based on expert experience. During data acquisition, the state s is obtained according to the dynamic target observation task requirements and sent to the task decision network to output the corresponding operation a, that is, for the satellite allocated to the current task and the corresponding available time window, the reward r is obtained; after the above operation is executed, the data will be stored in the following form in each iteration round: (s,a,r,s′). Then, the Mask mechanism: adaptively changes the selection of requirements based on the characteristics of heterogeneous satellites; In each scheduling cycle of dynamic target observation, it is necessary to select appropriate satellites to perform observation tasks based on the established constraints and objective functions. A heterogeneous satellite constellation task decision network architecture based on the Mask mechanism is designed to realize the selection of different satellites and corresponding time windows to ensure real-time relay tracking of dynamic targets. The constellation dynamic mission planning model for dynamic target observation missions is based on the actor-critic network architecture design. Finally, iterative improvement: parameter optimization based on the policy gradient method: A policy gradient-based reinforcement learning algorithm is used to optimize the parameters of the Actor network and the Critic network. During training, the policy gradient-based reinforcement learning method iteratively improves the scheduling policy based on the gradient of the cumulative expected reward estimated by the satellite state and policy parameters. The proposed policy gradient method uses two Critic networks, and their minimum value is used to estimate the expected reward for any given state. Let θ, φ1, φ2 be the parameters of the Actor network and the two Critic networks, respectively, and J(π θ ) represents the scheduling strategy π θ Expected return without discount, D k Represents the set of scheduling decision trajectories; expected return J(π) θ The gradient of ) is calculated as follows: Where τ represents the scheduling decision trajectory in the decomposed subtask. For scheduling strategy π θ The advantage function represents the advantage of a scheduling strategy relative to the average value given a given state, and is calculated using the following formula: Where, r k This indicates the k-th scheduling cycle. This represents the output of the i-th Critic network; Critic networks are used to approximate the state value V. π (s T,k ), that is, in state s T,k The following is based on the scheduling strategy π θ The average return; the update objective of the dual-Critic network is to minimize the mean square error of the approximate objective: in, The formula for calculating the cumulative future return is as follows:

6. The heterogeneous constellation intelligent mission decision-making method for observing space-moving targets according to claim 5, characterized in that: Based on the Actor-Critic network architecture design, the following sections will introduce each aspect: Actor Network: Based on the dynamic nature of target observation and the requirements for selecting heterogeneous satellites, the Actor network is designed as a deep pointer network architecture based on the Mask mechanism. It is responsible for calculating the collaborative task planning results for each scheduling cycle in real time, and its objective is to approximate the optimal task planning strategy π. * It is about maximizing the cumulative task rewards: Where, θ * The network parameters represent the optimal strategy; the input of the depth pointer network based on the Mask mechanism is the state sequence of each type of satellite's observation of the target in the k-th scheduling period, and the output of the current scheduling period is the observation combination corresponding to each target, including the selection combination of one electronic reconnaissance satellite and two optical or SAR satellites; The overall network structure of the pointer network consists of an encoder and a decoder. The encoder is composed of a one-layer one-dimensional convolutional neural network, which is responsible for processing the input sequence information s. T,k Feature extraction is performed to obtain embedded information. The decoder is the core component of the pointer network. It calculates the selection probability of each satellite in the input sequence through an attention mechanism. In the online scheduling stage, a greedy strategy is adopted to select the satellite with the highest probability value as the observation resource each time, until the observation satellites corresponding to J dynamic targets in the current scheduling period are selected. In each decoding process, a dynamic mask mechanism is designed. That is, when calculating the selection probability of each satellite in the input sequence, the probability of the selected satellite and the non-selected satellite type is set to 0. For example, if an electronic reconnaissance satellite is to be selected, the probability of optical satellite or SAR satellite is set to 0. Critic Network: In the offline training phase of the swarm dynamic task planning algorithm, a dual-Critic network structure is used for training. The input is the current state information sequence of the swarm, and the output is the future cumulative expected reward under the current state, which is used to evaluate the quality of the current state.

Citation Information

Patent Citations

  • Civil aviation Internet business management and access resource allocation method based on low-orbit giant satellite base

    CN114900225A

  • Star group collaborative task planning method based on mixed expert experience playback

    CN117068393A

  • Deep reinforcement learning scheduling method and device for satellite multi-point target imaging

    CN118709748A

  • Multi-satellite autonomous cooperative scheduling method based on distributed multi-agent reinforcement learning

    CN119623910A

  • SAR imaging satellite task planning method based on deep reinforcement learning

    CN120163365A