Multi-satellite cooperative multi-target observation planning method based on state estimation coupled reinforcement learning

CN122885318APending Publication Date: 2026-10-09INNOVATION ACAD FOR MICROSATELLITES OF CAS +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611023659.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

该类技术未针对多星协同观测的任务分配逻辑进行优化,也未解决纯测角观测条件下,多星观测几何构型对空间目标三维定位精度、滤波稳定性的直接影响问题

Benefits of technology

本发明的多星协同多目标观测规划方法面向多颗观测卫星、多个非合作目标和星载光学仅测角观测场景,将扩展卡尔曼滤波嵌入到强化学习的每个决策步长中,将EKF估计值及协方差特征作为决策网络的输入,并使卫星目标选择动作直接触发相应目标的EKF测量更新。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122885318A_ABST
    Figure CN122885318A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of multi-satellite cooperative multi-target observation planning methods based on state estimation coupling reinforcement learning, comprising the following steps: at each decision-making time, each satellite performs EKF time update to each target, obtains prior estimation state and prior covariance matrix, extracts target uncertainty feature from prior covariance matrix, and constructs decision network input in combination with the visibility flag of each satellite relative to each target;Each decision network outputs observation action, each satellite observes the selected target, and generates effective measurement set;When one or more satellites effectively observe the same target, construct multi-satellite joint measurement equation, and perform extended Kalman filter measurement update;At update interval, the parameters of decision network and value network are updated;And the trained multiple strategy networks are respectively deployed to corresponding satellite, and in the execution phase, the strategy network on each satellite independently generates observation target instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of satellite observation technology, and in particular to a multi-satellite collaborative multi-objective observation planning method based on state estimation coupled reinforcement learning. Background Technology

[0002] Current mainstream satellite mission planning technologies generally separate mission planning and state estimation into two independent execution stages. The planning algorithm typically assumes the target state is known or relies solely on offline prior information to generate satellite observation commands, failing to consider the dynamic impact of real-time observation geometry on orbit determination accuracy during the decision-making process. Traditional state estimation techniques can only passively receive and process observation data output from the planning system, unable to actively drive the mission planning system to obtain high-quality observation data that effectively reduces target state uncertainty. This open-loop working mode, where planning and estimation are separated, means that satellite observation resource allocation focuses solely on achieving target observation coverage, rather than maximizing system information gain and improving target observation accuracy, making it difficult to meet the application requirements of high-precision positioning and tracking of space targets. In the field of multi-satellite collaborative mission planning, existing distributed mission allocation methods mostly rely on explicit negotiation mechanisms such as contract network protocols and auction algorithms. During mission execution, these methods require satellite nodes to conduct multiple rounds of high-frequency information exchange and iterative negotiation to reach a consensus on mission allocation. However, the real-world onboard working environment has significant constraints: inter-satellite link bandwidth resources are limited, and inter-satellite communication suffers from unavoidable latency and data loss. High-frequency inter-satellite information exchange not only causes severe system scheduling delays, resulting in planning results lagging behind the real-time status of high-speed moving targets, but also easily leads to inter-satellite communication congestion, failing to meet the requirements of scenarios requiring real-time autonomous satellite decision-making. For observation targets such as space debris and non-cooperative space objects, their orbital states are highly unpredictable and maneuverable, exhibiting a highly dynamic evolutionary characteristic. Traditional satellite planning methods are mostly based on preset rules or deterministic optimization methods, lacking the ability to adaptively adjust to the uncertainty and dynamics of the target's orbit. Especially in scenarios where observation is conducted solely using optical sensors, the system can only acquire two-dimensional direction-finding information about the target's line-of-sight angle, resulting in inherent partial observability. Adopting a single, fixed observation strategy is highly susceptible to problems such as filter divergence and getting trapped in local optima, making it impossible to achieve stable convergence and accurate calculation of the target's three-dimensional position and velocity parameters. Deep reinforcement learning technology has been gradually applied to the field of satellite scheduling optimization. However, existing research focuses primarily on macroscopic discrete evaluation indicators such as mission completion rate and observation coverage, without deeply integrating the core technical requirements of underlying signal processing and precise orbit determination. It also neglects core quantitative indicators characterizing orbit determination accuracy, such as eigenvalues, traces, and determinants of the covariance matrix. This results in existing reinforcement learning scheduling strategies only achieving effective satellite observation coverage of the target, failing to construct a "three-dimensional observation configuration" suitable for high-precision orbit determination. Ultimately, this makes it difficult for observation planning strategies to achieve optimal results in terms of target positioning accuracy and tracking stability. Furthermore, most existing space target behavior monitoring technologies employ Kalman filtering to track target trajectories and combine it with reinforcement learning agents to achieve filter state switching and dynamic adjustment of filter parameters. Research focuses primarily on whether target behavior matches expectations, behavior pattern modeling, and future behavior prediction. These technologies do not optimize the task allocation logic for multi-satellite collaborative observations, nor do they address the direct impact of multi-satellite observation geometry on the 3D positioning accuracy and filtering stability of space targets under pure angle measurement conditions. Summary of the Invention To address at least some of the problems mentioned above in the prior art, this invention provides a multi-satellite collaborative multi-objective observation planning method based on state estimation coupled reinforcement learning, comprising the following steps: At each decision time, each satellite performs an extended Kalman filter time update for each target to obtain the prior estimated state and the prior covariance matrix. The target uncertainty features are extracted from the prior covariance matrix and combined with the visibility flags of each satellite relative to each target to construct the input of the decision network. Each decision network outputs an observation action, which is to select the current observation target; Each satellite observes the selected target and generates a valid set of measurements; When one or more satellites effectively observe the same target, construct a multi-satellite joint measurement equation and perform extended Kalman filter measurement updates; At the update interval, update the parameters of the decision network and the value network; and Multiple trained policy networks are deployed to corresponding satellites. During the execution phase, each policy network on a satellite independently generates observation target instructions based on local observations, estimated states from local extended Kalman filters, target uncertainty characteristics, and visibility flags.

[0003] Furthermore, before performing the extended Kalman filter time update, the process also includes: constructing a decentralized partially observable Markov decision process model, including: defining the global state space, the local observation space for each satellite, the target selection action, the state transition probability, the reward function, and the discount factor, wherein: The global state space includes the actual orbital state of the observed satellite, the actual state of the target, the estimated state of the target, the covariance matrix, and historical mission allocation information; The local observation space of each satellite contains the satellite's own orbital state information, the target's real-time state estimation information calculated by the extended Kalman filter, the target's real-time state estimation information includes the target's relative position estimate and relative velocity estimate, and the target's uncertainty characteristics.

[0004] Furthermore, before constructing a decentralized partially observable Markov decision process model, the following is also included: Initialize system parameters and mission scenario, including: setting the observation satellite set as follows: The target set is Input the orbital elements of each observation satellite, the initial orbital state of the target, the parameters of the onboard optical sensor, the initial covariance of the extended Kalman filter, the process noise, the measurement noise, and the reinforcement learning training parameters, and initialize the true state, estimated state, and covariance matrix of each target; Establish orbital dynamics model and angle-only measurement model; Construct observation visibility constraints and sensor pointing constraints. The observation visibility constraints include: for each satellite-target combination, the target needs to be within the sensor's half field of view, not be blocked by the Earth, meet the requirements of solar illumination and solar avoidance angle, and the distance between the satellite and the target does not exceed the maximum detection distance. The sensor pointing constraint includes the requirement that the time required for the sensor to turn from the current optical axis to the target direction is less than the current decision step size.

[0005] Furthermore, the target uncertainty features include diagonal elements, traces, determinants, or eigenvalues; The prior estimated state includes the target's relative position and velocity estimates.

[0006] Furthermore, each decision network receives prior estimated state, target uncertainty characteristics, and visibility flags, outputs a selection probability distribution for multiple targets, and samples or selects the currently observed target based on the selection probability distribution.

[0007] Furthermore, when each satellite observes a selected target, an action validity judgment is performed to determine whether the current satellite observation action is a valid observation; For each target, a set of satellites that are selected for effective observation is generated. If the set is empty, no measurement update is performed for the target.

[0008] Furthermore, the multi-satellite joint measurement equation is formed by superimposing satellite measurement vectors, observation matrices, and measurement noise matrices; Perform extended Kalman filter measurement update: calculate the Kalman gain and update the target posterior estimated state and posterior covariance matrix.

[0009] Furthermore, before updating the parameters of the decision network and the value network, a piecewise guided reward function is constructed based on the estimation error, including: like , , This represents the reward scaling factor. This represents the normalized positioning error; like , , A fixed positive reward is a pre-set, unchanging positive reward value. Total Rewards , This indicates the terminal reward item.

[0010] Furthermore, during the training phase, the input to the value network includes the actual satellite orbital state, the actual target state, the estimated target state, the covariance matrix, and historical cooperative observation state information, and the value network outputs the state value.

[0011] The present invention also provides a computer-readable storage medium having a computer program stored thereon, the computer program performing the steps of the above method when executed by a processor.

[0012] The present invention has at least the following beneficial effects: The multi-satellite collaborative multi-target observation planning method of the present invention is designed for scenarios involving multiple observation satellites, multiple non-cooperative targets, and spaceborne optical angle-only observation. It embeds extended Kalman filtering into each decision step of reinforcement learning, uses EKF estimates and covariance features as inputs to the decision network, and enables satellite target selection actions to directly trigger the EKF measurement update of the corresponding target.

[0013] This invention adopts a "centralized training, decentralized execution" architecture, independently deploying the trained decision-making strategy network to each satellite in the constellation, achieving fully decentralized autonomous collaboration. Each satellite only needs to rely on its own sensor measurement data, locally maintained target EKF estimation state, covariance characteristics, and target visibility flags to directly output the current optimal observation target number. There is no need to exchange bidding, marginal revenue, or task allocation information at each decision moment, nor is there a need to access the global real state. Relying on the implicit collaborative strategies learned by each agent during the training phase, it can spontaneously form multi-satellite synchronous observation and alternating observation configurations in mission scenarios with limited inter-satellite links, large communication latency, or communication silence requirements. This enables high-precision and high-efficiency collaborative orbit determination of non-cooperative targets, significantly improving the positioning accuracy, filtering consistency, and tracking stability of non-cooperative targets.

[0014] This scheme combines the measurement equations of each satellite into a superimposed measurement equation when multiple satellites simultaneously observe the same target, and constructs a joint observation matrix and a joint observation noise matrix. This mechanism enables line-of-sight measurements from satellites at different spatial locations to jointly contribute to the EKF posterior update, improving the problem of unobservable distances in angle-only positioning.

[0015] This scheme not only selects the observation target, but also simulates the continuous rotation process of the sensor pointing from the current target to the new target, and takes into account constraints such as field of view, solar avoidance, earth occlusion, and detection distance. This setting avoids the algorithm from generating observation sequences that are theoretically advantageous but physically unexecutable. Attached Figure Description

[0016] To further illustrate the above and other advantages and features of the various embodiments of the present invention, a more specific description of the embodiments of the invention will be presented with reference to the accompanying drawings. It is to be understood that these drawings depict only typical embodiments of the invention and are therefore not intended to limit its scope. In the drawings, identical or corresponding parts will be indicated by identical or similar reference numerals for clarity.

[0017] Figure 1A The flowchart of a multi-satellite collaborative multi-objective observation planning method based on state estimation coupled reinforcement learning according to an embodiment of the present invention is shown.

[0018] Figure 1B The flow of multi-agent proximal policy optimization (PEC-MAPPO) according to an embodiment of the present invention is illustrated.

[0019] Figure 2 Pseudocode for implementing a multi-satellite collaborative multi-target observation planning method according to an embodiment of the present invention is shown.

[0020] Figure 3A The positioning accuracy convergence curve of a baseline algorithm according to an embodiment of the present invention is shown.

[0021] Figure 3B The PEC-MAPPO positioning accuracy convergence curve according to an embodiment of the present invention is shown.

[0022] Figure 3C A comparison chart of the actual error reduction effect according to an embodiment of the present invention is shown.

[0023] Figure 4A The root mean square error curve of filter estimation using a baseline strategy according to an embodiment of the present invention is shown.

[0024] Figure 4B The root mean square error curve of filter estimation using PEC-MAPPO according to an embodiment of the present invention is shown.

[0025] Figure 4C A comparison chart showing the reduction effect of the baseline strategy and the PEC-MAPPO filter estimation root mean square error according to an embodiment of the present invention is presented. Detailed Implementation

[0026] It should be noted that the components in the accompanying drawings may be shown exaggerated for illustrative purposes and may not be to scale.

[0027] In this invention, the various embodiments are merely intended to illustrate the solutions of the invention and should not be construed as limiting.

[0028] In this invention, unless otherwise specified, the quantifiers “a” and “one” do not exclude scenarios involving multiple elements.

[0029] It should also be noted that, in the embodiments of the present invention, only a portion of the parts or components may be shown for clarity and simplicity. However, those skilled in the art will understand that, under the teachings of the present invention, the required parts or components can be added as needed for specific scenarios.

[0030] It should also be noted that within the scope of this invention, the terms "same", "equal", and "equal to" do not mean that the two values ​​are absolutely equal, but allow for a certain reasonable error. In other words, the terms also cover "substantially the same", "substantially equal", and "substantially equal to".

[0031] It should also be noted that in the description of this invention, the terms "center," "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not explicitly or implicitly suggest that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0032] Furthermore, the embodiments of the present invention describe the process steps in a specific order; however, this is only for the convenience of distinguishing each step, and is not intended to limit the order of the steps. In different embodiments of the present invention, the order of each step can be adjusted according to the process.

[0033] This invention addresses the challenge of angle-only observation and positioning of non-cooperative targets in high orbits. It models the multi-satellite collaborative observation planning problem as a decentralized partially observable Markov decision process, namely Dec-POMDP. The state estimation results and covariance uncertainty information of the Extended Kalman Filter (EKF) are embedded into the observation space and reward function of multi-agent reinforcement learning. The PEC-MAPPO algorithm is used to learn multi-satellite collaborative observation strategies within a centralized training and decentralized execution CTDE framework. During the execution phase, decision networks are deployed on each satellite, independently selecting observation targets based solely on local observations and local EKF information, eliminating the need for frequent inter-satellite communication. By constructing geometric observation configurations favorable for angle-only positioning through synchronous or alternating observations by multiple satellites, the accuracy of multi-target positioning and the stability of filtering are improved.

[0034] In this invention, the EKF state estimation process is embedded into each decision step of reinforcement learning: satellite observation actions determine which targets can receive measurement updates; measurement updates change the posterior estimated state and covariance matrix; the posterior estimated state and covariance information then serve as the input for the next agent observation and the basis for reward calculation. This closed-loop mechanism enables the reinforcement learning strategy to learn "which observation actions best reduce future positioning uncertainty," representing a deep coupling of planning and estimation.

[0035] The core technologies of this invention lie in multi-satellite collaborative multi-target observation planning, improved positioning accuracy under angle-only conditions, closed-loop coupling of EKF and multi-agent reinforcement learning, and decentralized execution. Instead of introducing behavioral pattern clustering, it directly optimizes the contribution of observation resource allocation to positioning error, covariance reduction, and multi-satellite observation geometry.

[0036] Figure 1A The flowchart of a multi-satellite collaborative multi-objective observation planning method based on state estimation coupled reinforcement learning according to an embodiment of the present invention is shown. Figure 1B The flow of multi-agent proximal policy optimization (PEC-MAPPO) according to an embodiment of the present invention is illustrated.

[0037] like Figure 1A and 1B As shown, a multi-satellite collaborative multi-objective observation planning method based on state estimation coupled reinforcement learning includes the following steps: Step S1: Initialize system parameters and mission scenario. Let the set of observation satellites be... The target set is Input the orbital elements of each observation satellite, the initial orbital state of the target, the parameters of the onboard optical sensor, the initial covariance of EKF (Extended Kalman Filter), the process noise, the measurement noise, and the reinforcement learning training parameters (MAPPO parameters, see Table 1) to initialize the true state, estimated state, and covariance matrix of each target.

[0038] Step S2: Establish the orbital dynamics model and the angle-only measurement model.

[0039] The orbital dynamics model describes the relative motion between the satellite and the target. It is used to propagate the actual orbital states of the observed satellite and target.

[0040] The angle-only measurement model is used to generate line-of-sight, azimuth, or elevation angle observations based on the relative position vector between the satellite and the target. The measurement noise is set to zero-mean Gaussian white noise. The angle-only measurement model is constrained by the field of view, Earth occlusion, solar avoidance angle, maximum detection range, and maximum pointing angular velocity of the sensor.

[0041] Step S3: Construct observation visibility constraints and sensor pointing constraints.

[0042] Observation visibility constraints: For each satellite-target combination, the target must be located within half the sensor's field of view, not be obscured by the Earth, meet the requirements for solar illumination and solar avoidance angle, and the distance between the satellite and the target must not exceed the maximum detection range.

[0043] Sensor pointing constraint: The time required for the sensor to turn from the current optical axis to the target direction is less than the current decision step size.

[0044] It is possible to simulate the continuous and smooth rotation of the sensor's optical axis in three-dimensional space based on the Slerp algorithm.

[0045] Step S4: Construct the Dec-POMDP (Decentralized Partially Observable Markov Decision Process) model. Define the global state space, the local observation space for each satellite, the target selection action (action space), the state transition probability, the reward function, and the discount factor.

[0046] The Dec-POMDP model can be expressed as follows: Here, S represents the global state space, which includes the actual orbital state of the observed satellites, the actual state of the target, the estimated state of the target, the covariance matrix, and historical task allocation information. The historical task allocation information in the global state space is the cooperative state derived from the joint actions or effective observation results of previous moments, used to enable the value network to perceive multi-satellite cooperative relationships during the training phase. Historical task allocation information may include, for example, the number of satellites that effectively observed each target at the previous moment, the duration of continuous observation of each target by each satellite, and the time since the last EKF measurement update of the target.

[0047] For the first The local observation space of a satellite contains the satellite's own orbital state information, the target's real-time state estimation information calculated by the extended Kalman filter (EKF), which includes the target's relative position and velocity estimates, and the eigenvalues ​​of the EKF covariance matrix (target uncertainty characteristics). For the first The action space (target selection action) of each satellite means selecting an observation target from N targets at the current decision moment, rather than switching or adjusting filter parameters. The actions of each satellite are combined to form a joint action. This joint action is the result of the multi-satellite collaborative observation mission planning at the current moment; P is the state transition probability, and R is the reward function based on positioning error and filtering stability. This is the discount factor.

[0048] This invention models the multi-satellite collaborative mission planning problem as a decentralized partially observable Markov decision process (Dec-POMDP).

[0049] Step S5: At each decision time, each satellite performs an EKF time update for each target, obtaining the prior estimated state and the prior covariance matrix. Target uncertainty features are extracted from the prior covariance matrix and combined with the visibility indicators of each satellite relative to each target to construct the decision network input. The decision network input includes the prior state estimate obtained from each satellite's time update for each target, the target uncertainty features, and the visibility indicators.

[0050] At each decision time, a priori recursion (EKF time update) is first performed on the estimated state and covariance of each target to obtain the prior estimated state and prior covariance matrix. Then, the diagonal elements, trace, determinant, or eigenvalues ​​are extracted from the prior covariance matrix as target uncertainty features, and input together with the visibility flag into the Actor (decision) network. The prior estimated state includes the target's relative position estimate and relative velocity estimate.

[0051] The visibility flag is a binary variable defined for the "satellite i-target j" combination, and can be denoted as v_{i,j,k}. When target j satisfies the observation visibility constraint and sensor pointing constraint for satellite i at time k, v_{i,j,k}=1; otherwise, v_{i,j,k}=0.

[0052] For example, the decision network input uses geometric visibility flags. After the observation action is executed, it will also determine whether a valid observation has been formed based on the angle between the current sensor optical axis and the true direction of the target.

[0053] This invention designs a "planning-estimation coupled" state-space structure: defining the local observation space (local observation space) for each satellite. In addition to including the satellite's own orbital state, sensor pointing deviation, operating mode, and its own identification number, the system also embeds real-time target state estimation information calculated through local parallel EKF. This estimation information includes estimates of the target's relative position and velocity, the diagonal elements, trace, determinant, or eigenvalues ​​of the EKF covariance matrix, and target visibility indicators. This design allows the input of the observation strategy to directly include features characterizing the size of the target uncertainty ellipsoid and the observation geometry quality, enabling the observation satellite to perceive the "current observation quality" and "positioning accuracy requirements," thus laying a data foundation for closed-loop planning.

[0054] Step S6: The Actor network outputs the observation action. Each observation satellite's Actor network receives local observations (prior estimated state, target uncertainty characteristics, and visibility indicators), outputs a selection probability distribution for N targets, and samples or selects the current observation target based on this distribution. The output of the decision network is limited to a target number selection action.

[0055] Step S7: Observe the selected target for each satellite and generate a valid set of measurements.

[0056] When a satellite observes a selected target, an action validity judgment is performed to determine whether the current satellite observation action is a valid observation. Specifically, it is necessary to determine whether the observation visibility constraint and sensor pointing constraint are met.

[0057] The system drives each satellite sensor to turn towards its corresponding target based on the joint action. Each Actor (decision-making) network outputs target selection actions, forming a joint action. , This represents the target selection action for the first satellite at time k. The combined action represents the multi-satellite collaborative observation mission planning result at the current time.

[0058] Each Actor network outputs its target selection action for its satellite, and all satellite actions are combined into a joint action. The system then... The target number selected for each satellite, combined with the target position estimated by EKF prior, drives the corresponding onboard optical sensor to rotate in the direction of the target; then visibility constraints and sensor pointing constraints are used to determine whether the action produces a valid observation.

[0059] In determining the effectiveness of an action, the system checks for each satellite-target combination whether the target is within the sensor's half-field of view, whether it is obscured by the Earth, whether it meets the requirements for solar illumination and solar avoidance angles, whether the distance between the satellite and the target exceeds the maximum detection range, and whether the time required for the sensor to turn from the current optical axis to the target direction meets the current decision step size. For changes in sensor pointing, spherical linear interpolation (Slerp) can be used to simulate the continuous rotation of the optical axis. Only when the target enters the half-field of view and all the above constraints are met can the selected action produce a valid observation.

[0060] For each target, the set of satellites that selected that target and actually generated valid observations is counted; if the set is empty, no measurement update is performed for that target.

[0061] Step S8: When one or more satellites effectively observe the same target, construct a multi-satellite joint measurement equation and perform an EKF measurement update.

[0062] When one or more satellites effectively observe the same target, a joint measurement matrix is ​​constructed. The multi-satellite joint measurement equation is formed by superimposing the satellite measurement vectors, the observation matrix, and the measurement noise matrix. An EKF measurement update is performed: the Kalman gain is calculated, and the target posterior estimated state and posterior covariance matrix are updated.

[0063] When multiple satellites simultaneously and effectively observe the same target at the same decision moment In this invention, a multi-satellite joint measurement equation is constructed. Let the effective set of observation satellites be... If satellite i satisfies the effective observation conditions for target j, then the measurement vectors, observation matrices, and measurement noise matrices of each satellite are superimposed as follows: , Represents satellite measurement vectors. Indicates the first i The measurement vector obtained by one effective observation satellite for target j at time k. Represents the observation matrix. Indicates the first i The observation matrix corresponding to one satellite Represents the measurement noise matrix. Indicates the first i The noise covariance matrix is ​​measured by one satellite.

[0064] Calculation of innovation covariance and Kalman gain based on multi-satellite joint measurement equations. Formula for calculating innovation covariance: ,in This represents a joint observation matrix formed by stacking the observation matrices of multiple effective observation satellites row by row. Let R_aug represent the prior covariance matrix of target j at time k, and let R_aug represent the joint measurement noise matrix composed of the measurement noise covariance matrices of each effective observation satellite. This represents the covariance matrix of the joint measurement information.

[0065] Kalman gain calculation formula: .

[0066] First calculate the joint measurement information Then update the posterior estimated state of the target. And update the posterior covariance matrix. .

[0067] If the effective observation satellite set If the value is empty, then target j only undergoes time updates and not measurement updates, and its covariance changes with dynamic propagation. Therefore, the observation planning action directly determines the degree of contraction of the EKF posterior uncertainty.

[0068] For angle-only measurement models, spaceborne optical sensors cannot directly obtain the target distance; they can only form observations of the line-of-sight direction, azimuth, or elevation angle based on the relative position vectors of the satellite and the target. Single-satellite short-arc observations suffer from weak or even unobservable depth conditions. This invention improves angle-only positioning conditions by forming a spatial baseline through multi-satellite synchronous observations or by forming spatiotemporal intersection geometry through multi-satellite alternating observations, addressing both the observation matrix condition number and posterior covariance contraction.

[0069] The condition number of the observation matrix is ​​the condition number of the joint observation matrix H_aug, which is the ratio of the maximum singular value to the minimum singular value. It is used to measure the ill-conditioned nature of the observation geometry. The smaller the condition number, the more balanced the multi-satellite line-of-sight geometry, and the better the observability of the target position in all directions. Posterior covariance shrinkage refers to the decrease in the trace, determinant, or eigenvalues ​​of the posterior covariance matrix of the target position estimate relative to the prior covariance after EKF measurement updates, indicating a reduction in positioning uncertainty.

[0070] Step S9: Construct a piecewise guided reward function based on the estimation error. This is done based on the target localization error, the error threshold, and the terminal reward item. The reward is calculated for each step, enabling the strategy to converge quickly when the error is large and remain stable once the error reaches the threshold.

[0071] Segmented Guided Reward Mechanism: A two-stage reward function based on estimation error is designed. A normalized localization error is defined for target j: , in To train the simulation on the true position of target j at time k, For EKF posterior location estimation, E norm This serves as the normalization scale for positioning errors. When the error exceeds a set error threshold... During the convergence phase, a guided reward related to error descent is used to provide a continuous gradient to accelerate error convergence; when the error is less than or equal to the error threshold... During the maintenance phase, a fixed positive reward is given. (A pre-set, constant positive reward value) is used to suppress the agent from excessively chasing random measurement noise and to maintain policy stability.

[0072] The reward function can be expressed as: if , like , Total Rewards .in, Indicates the reward scaling factor; This represents the terminal reward, used to penalize the terminal when the average positioning error exceeds a terminal threshold within the evaluation window, and to reward the terminal upon task completion. This reward directly corresponds to positioning accuracy and filtering stability, rather than general task coverage or behavior monitoring results.

[0073] Step S10: Update the parameters of the decision network and the value network at the update interval.

[0074] Local observation information (target estimated state, target uncertainty characteristics, and visibility indicators), global state, joint actions, rewards, and next state are stored in the trajectory cache module. After the update interval is reached, the advantage function is calculated based on the reward and the state value output by the Critic. The advantage function is substituted into the PPO-Clip objective function, the Actor network is updated using the PPO-Clip objective function, and the Critic network is updated using the value loss function.

[0075] The estimated state and the true state are used to calculate the PPO-Clip objective function, while other information serves as the decision criteria.

[0076] The next state is the state at the next time step in a reinforcement learning sample; the training phase may include the next global state. and the next local observation of each satellite The next global state, during simulation training, can include the orbital states of all satellites, the true state of the target, the estimated state of the target, and covariance characteristics (target uncertainty characteristics). The next local observation includes the local observations available for each satellite at the next time step, the EKF estimation results, the covariance characteristics (target uncertainty characteristics), and the visibility flags. The true state of the target is not used during the execution phase.

[0077] This invention employs a multi-agent proximal strategy optimization MAPPO for centralized training of the Actor network and the Critic (value) network.

[0078] This invention employs a centralized training, decentralized execution (CTDE) architecture for model training. During the training phase, the Critic value network receives inputs including the satellite's true orbital state, the target's true state, the estimated target state, the covariance matrix, and historical cooperative observation state information, used to estimate the state value. The Actor network receives inputs from the local observation information of each satellite (the target's prior estimated state, target uncertainty features, and visibility indicators), and outputs the target selection probability. During training, the parameters of the Actor network are updated based on the PPO-Clip (Proximal Policy Optimization Clipping) objective function, and the parameters of the Critic network are updated based on the value loss function. GAE (Generalized Advantage Estimation) is used to estimate the value (advantage) of the advantage function using state value and reward, reducing the training variance of the multi-agent policy and thus learning a cooperative observation strategy that maximizes positioning accuracy and filtering consistency. Table 1 shows the parameters involved in the Multi-Agent Proximal Policy Optimization (MAPPO).

[0079] This invention employs a neural network structure comprising an Actor network and a Critic network. The Actor network is a Multi-Layer Perceptron (MLP). It consists of an input layer, three fully connected hidden layers, and an output layer. Each hidden layer contains 512 neurons and uses the Corrected Linear Unit (ReLU) activation function, which helps accelerate network convergence and alleviate the gradient vanishing problem. The number of neurons in the output layer equals the dimension of the action space, i.e., the number of targets N. The output layer uses the Softmax activation function and outputs a probability distribution vector of dimension N. The Critic network is also an MLP with three fully connected hidden layers, with parameters similar to the Actor network. Each hidden layer also has 512 neurons and uses ReLU as the activation function. It is used to extract high-level features from global information. The output layer is a linear layer with the number of neurons equal to the number of agents M, and it does not use a non-linear activation function.

[0080] Step S11: The trained policy networks are deployed to the corresponding satellites. During the execution phase, the policy network on each satellite independently generates observation target instructions based on local observations, local EKF estimation results (estimated state), covariance features (target uncertainty features), and visibility flags.

[0081] During the mission execution phase, the trained Actor policy network is independently deployed to each satellite in the constellation, achieving fully decentralized autonomous collaboration. Each satellite only needs to rely on its own sensor measurement data, locally maintained target EKF estimated state (target estimated state calculated by EKF), covariance features (diagonal elements, trace, determinant, or eigenvalues ​​of the EKF covariance matrix), and target visibility flags to directly output the current optimal observed target number.

[0082] The local observations of each satellite during the execution phase may include: satellite orbital status, current pointing or alignment status of satellite sensors, estimated azimuth / elevation angles or relative position and velocity information of each target relative to the satellite, EKF estimated status of each target, covariance characteristics (target uncertainty characteristics) of each target, visibility flags, continuous observation duration, and time since the last measurement update. The target's EKF estimated status mainly refers to the target state estimation vector, including target position and velocity estimates; covariance characteristics refer to the diagonal elements, trace, determinant, or eigenvalues ​​extracted from the covariance matrix; and the visibility flag is a 0 / 1 indicator indicating whether the current satellite meets the observation constraints for the target.

[0083] The output of this invention includes the observation target instructions, target state estimation sequence, covariance update results, and positioning error improvement results for each observation satellite at each decision time. Its output is not a target behavior anomaly detection or future behavior pattern prediction, but rather observation mission planning instructions that can directly drive the onboard optical sensors.

[0084] The method of this invention does not require exchanging bidding, marginal revenue or task allocation information at each decision moment, nor does it require access to the global real state. Relying on the implicit cooperative strategies learned by each agent during the training phase, it can spontaneously form multi-satellite synchronous observation and alternating observation configurations in task scenarios where inter-satellite links are limited, communication latency is large or there are requirements for communication silence, thereby achieving high-precision and high-efficiency cooperative orbit determination of non-cooperative targets.

[0085] Table 1. MAPPO training parameters.

[0086] See the pseudocode for the above method implementation. Figure 2 .

[0087] This invention proposes a multi-agent proximal policy optimization (PEC-MAPPO) algorithm framework to implement the state estimation coupled method described above. The PEC-MAPPO algorithm comprises multiple Actor networks and a single Critic network. The core feature of this framework is: A dual embedding mechanism is employed: the EKF state estimation process is directly embedded into each decision step of reinforcement learning. Before each decision step, the EKF time update is performed on each target using a dynamic model to obtain the prior estimated state and prior covariance matrix; the Actor network outputs a target selection action based on local observations containing covariance features (the target's prior estimated state, target uncertainty features, and visibility indicators); after making a decision and executing the observation action, if the target selected by the satellite satisfies the visibility and pointing constraints, the action generates the corresponding measurement equation and triggers an EKF measurement update; the updated posterior covariance matrix... The feedback continues to the Actor policy network as input for the next step, forming a closed loop of "estimation-planning-observation-re-estimation".

[0088] For each satellite and each target, the prior estimated state and prior covariance matrix of the target obtained by EKF time update are used to calculate the observation geometric information of the target relative to the satellite; uncertainty features are extracted from the prior covariance matrix; then visibility / observability flags are generated according to observation visibility constraints and sensor pointing constraints; finally, these quantities are concatenated with the satellite's own state to form the local observation vector input to the satellite's Actor network.

[0089] When the same target is observed simultaneously at the same decision moment When the number of satellites is 1, the joint measurement equation degenerates into a single-satellite measurement update; when the number is greater than 1, the measurement vectors, observation matrices and noise matrices of multiple satellites are stacked to form a multi-satellite joint measurement equation for joint update.

[0090] A segmented guided reward mechanism was designed, employing a two-stage reward function based on estimation error. This has been detailed above and will not be repeated here.

[0091] The model is trained using a centralized training, decentralized execution (CTDE) architecture.

[0092] To verify the steady-state accuracy and robustness of the method of this invention at the statistical level, 200 Monte Carlo simulation experiments were conducted, and the performance differences between the state estimation coupled multi-agent proximal policy optimization (PEC-MAPPO) of this invention and the baseline algorithm were systematically compared.

[0093] Figure 3A The positioning accuracy convergence curve of a baseline algorithm according to an embodiment of the present invention is shown. Figure 3B The PEC-MAPPO positioning accuracy convergence curve according to an embodiment of the present invention is shown. Figure 3C A comparison chart of the actual error reduction effect according to an embodiment of the present invention is shown.

[0094] Figure 3AThis demonstrates the true position error using the MAPPO strategy. Figure 3B The true position error using the PEC-MAPPO strategy is shown.

[0095] like Figure 3C As shown, using the baseline algorithm, the average positioning error for all targets is 0.3746 km (approximately 0.37 km). Figure 3C As shown, PEC-MAPPO, by introducing feedback from planning and estimation, reduces the average positioning error of all targets to 0.2498 km (approximately 0.25 km). This means that, under the same observation conditions, the method of this invention improves the overall positioning accuracy of the system by 33.32%, verifying its effectiveness in improving observation performance.

[0096] Specifically, in single-target analysis, taking Target 9, which has the largest error under the baseline strategy, as an example, its average positioning error is 0.8132 km. PEC-MAPPO reduces the average positioning error to 0.3476 km, an improvement of 57.26% in positioning accuracy. Furthermore, for Target 3 and Target 5, the PEC-MAPPO of this invention achieves improvements of 64.06% and 70.02%, respectively. It also shows that for Target 1 and Target 8, the positioning error of PEC-MAPPO is slightly higher than the baseline. This is due to the different decision-making logic between the current PEC-MAPPO strategy and the baseline strategy: when the positioning error of some targets is already sufficiently low, the system allocates observation resources to targets with larger errors, thereby effectively balancing the error distribution of all targets within the system, avoiding situations where individual targets have low positioning accuracy, and improving the overall positioning accuracy.

[0097] Figure 4A The root mean square error curve of filter estimation using a baseline strategy according to an embodiment of the present invention is shown. Figure 4B The root mean square error curve of filter estimation using PEC-MAPPO according to an embodiment of the present invention is shown. Figure 4C A comparison chart showing the reduction effect of the baseline strategy and the PEC-MAPPO filter estimation root mean square error according to an embodiment of the present invention is presented.

[0098] The root mean square error not only reflects the magnitude of a single estimation deviation, but is also a key indicator for measuring the consistency of filtering and the robustness of the system. Figures 4A to 4C The root mean square error (RMSE) of filter estimation under the two algorithm frameworks was further compared, and the superiority of the method of the present invention was further verified from the perspective of error fluctuation amplitude.

[0099] like Figure 4A and 4CAs shown, the baseline strategy exhibits significant fluctuations in the observation and localization task, with a system-wide average RMSE of 0.9761 km (approximately 0.98 km), indicating a large variance in the filtering results. In contrast, as... Figure 4C As shown, PEC-MAPPO, with its closed-loop strategy, reduced the average RMSE of the entire system to 0.4982 km (approximately 0.5 km). The overall tracking stability of the system was improved by 48.97%, effectively curbing the frequency of extreme error values ​​and improving the stability of the estimation results.

[0100] Looking at the RMSE statistics for a single target, taking Target 9, which has the most unstable performance, as an example, its RMSE under the baseline strategy is 2.4534 km, indicating that its location estimation is relatively unreliable. PEC-MAPPO reduced its RMSE to 0.6159 km, an improvement of 74.90%. Figure 4C The RMSE of some targets also showed a slight increase, which further confirms the system's resource scheduling logic: sacrificing the redundancy and stability of some targets to balance the overall estimation robustness.

[0101] In summary, the experimental results show that PEC-MAPPO dynamically constructs a multi-satellite stereo view at key stages such as initial acquisition, error divergence, and high-precision maintenance, thereby maintaining a stable estimate of the target state during dynamic observation and positioning.

[0102] The reason for the aforementioned technical effects lies in the fact that this invention does not simply utilize reinforcement learning to adjust filter parameters, but rather enables the observation action to determine which targets receive single-satellite or multi-satellite joint measurement updates, and allows EKF posterior uncertainty to continue into the next agent's observation space. Line-of-sight measurements provided by multiple satellites from different spatial locations can improve the rank or condition number of the angle-only observation matrix, thereby reducing target depth direction errors and improving filtering consistency.

[0103] The output of this invention includes the observation target instructions, target state estimation sequence, covariance update results, and positioning error improvement results for each observation satellite at each decision time. Its output is not a target behavior anomaly detection or future behavior pattern prediction, but rather observation mission planning instructions that can directly drive the onboard optical sensors.

[0104] While some embodiments of the present invention have been described in this application, those skilled in the art will understand that these embodiments are merely illustrative. Numerous variations, alternatives, and improvements will arise in those skilled in the art under the teachings of this invention without departing from its scope. The appended claims are intended to define the scope of the invention and thereby cover methods and structures within the scope of the claims themselves and their equivalents.

Claims

1. A multi-satellite collaborative multi-objective observation planning method based on state estimation coupled with reinforcement learning, characterized in that, Includes the following steps: At each decision time, each satellite performs an extended Kalman filter time update for each target to obtain the prior estimated state and the prior covariance matrix. The target uncertainty features are extracted from the prior covariance matrix and combined with the visibility flags of each satellite relative to each target to construct the input of the decision network. Each decision network outputs an observation action, which is to select the current observation target; Each satellite observes the selected target and generates a valid set of measurements; When one or more satellites effectively observe the same target, construct a multi-satellite joint measurement equation and perform extended Kalman filter measurement updates; At the update interval, update the parameters of the decision network and the value network; as well as Multiple trained policy networks are deployed to corresponding satellites. During the execution phase, each policy network on a satellite independently generates observation target instructions based on local observations, estimated states from local extended Kalman filters, target uncertainty characteristics, and visibility flags.

2. The multi-satellite collaborative multi-objective observation planning method based on state estimation coupled reinforcement learning according to claim 1, characterized in that, Before performing the extended Kalman filter time update, the following steps are also included: constructing a decentralized partially observable Markov decision process model, including: defining the global state space, the local observation space for each satellite, the target selection action, the state transition probability, the reward function, and the discount factor, wherein: The global state space includes the actual orbital state of the observed satellite, the actual state of the target, the estimated state of the target, the covariance matrix, and historical mission allocation information; The local observation space of each satellite contains the satellite's own orbital state information, the target's real-time state estimation information calculated by the extended Kalman filter, the target's real-time state estimation information includes the target's relative position estimate and relative velocity estimate, and the target's uncertainty characteristics.

3. The multi-satellite collaborative multi-objective observation planning method based on state estimation coupled reinforcement learning according to claim 2, characterized in that, Before constructing a decentralized, partially observable Markov decision process model, the following is also included: Initialize system parameters and mission scenario, including: setting the observation satellite set as follows: The target set is Input the orbital elements of each observed satellite, the initial orbital state of the target, the parameters of the onboard optical sensor, the initial covariance of the extended Kalman filter, the process noise, the measurement noise, and the reinforcement learning training parameters, and initialize the true state, estimated state, and covariance matrix of each target; Establish orbital dynamics model and angle-only measurement model; Construct observation visibility constraints and sensor pointing constraints. The observation visibility constraints include: for each satellite-target combination, the target needs to be within the sensor's half field of view, not be blocked by the Earth, meet the requirements of solar illumination and solar avoidance angle, and the distance between the satellite and the target does not exceed the maximum detection distance. The sensor pointing constraint includes the requirement that the time required for the sensor to turn from the current optical axis to the target direction is less than the current decision step size.

4. The multi-satellite collaborative multi-objective observation planning method based on state estimation coupled reinforcement learning according to claim 1, characterized in that, The target uncertainty features include diagonal elements, trace, determinant, or eigenvalues; The prior estimated state includes the target's relative position and velocity estimates.

5. The multi-satellite collaborative multi-objective observation planning method based on state estimation coupled reinforcement learning according to claim 1, characterized in that, Each decision network receives prior estimated state, target uncertainty characteristics, and visibility flags, outputs a selection probability distribution for multiple targets, and samples or selects the currently observed target based on the selection probability distribution.

6. The multi-satellite collaborative multi-objective observation planning method based on state estimation coupled reinforcement learning according to claim 1, characterized in that, When a target is selected for each satellite observation, an action validity judgment is performed to determine whether the current satellite observation action is a valid observation. For each target, a set of satellites that are selected for effective observation is generated. If the set is empty, no measurement update is performed for the target.

7. The multi-satellite collaborative multi-objective observation planning method based on state estimation coupled reinforcement learning according to claim 1, characterized in that, The multi-satellite joint measurement equation is composed of the superposition of satellite measurement vectors, observation matrix, and measurement noise matrix; Perform extended Kalman filter measurement update: calculate the Kalman gain and update the target posterior estimated state and posterior covariance matrix.

8. The multi-satellite collaborative multi-objective observation planning method based on state estimation coupled reinforcement learning according to claim 1, characterized in that, Before updating the parameters of the decision network and value network, a piecewise guided reward function based on the estimation error is constructed, including: like , , This represents the reward scaling factor. This represents the normalized positioning error; like , , A fixed positive reward is a pre-set, unchanging positive reward value. Total Rewards , This indicates the terminal reward item.

9. The multi-satellite collaborative multi-objective observation planning method based on state estimation coupled reinforcement learning according to claim 1, characterized in that, During the training phase, the input to the value network includes the actual satellite orbital state, the actual target state, the estimated target state, the covariance matrix, and historical cooperative observation state information. The value network outputs the state value.

10. A computer-readable storage medium having a computer program stored thereon, the computer program performing the steps of the method according to any one of claims 1-9 when executed by a processor.