An Intelligent Cooperative Tracking Method for Aerial Moving Targets by a Remote Sensing Constellation
Through the multi-objective gaze tracking model and binary star geometric positioning model, combined with reinforcement learning methods, the tracking and positioning problems of HGVs in remote sensing constellation systems are solved, efficient and real-time multi-star collaborative tracking and positioning are achieved, and the limitations of traditional algorithms are overcome.
Patent Information
- Application Number
- CN202310106923.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-14
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-02-14
AI Technical Summary
The existing remote sensing constellation systems have large solution space, easy to fall into local optimization and difficult to deal with high maneuverability problems in tracking and positioning hypersonic vehicles (HGVs). Traditional task decision algorithms have high computing pressure, making it difficult to achieve real-time and effective multi-star collaborative tracking.
A multi-objective gaze tracking model and a binary star geometric positioning model were designed, combining constellation configuration, inter-star communication constraints and infrared sensor constraints, and intelligent collaborative tracking of remote sensing constellations was realized through reinforcement learning methods, and independent real-time decision-making was made using the MAPPO algorithm.
The efficient, real-time tracking and positioning of HGVs by remote sensing constellations is achieved, and the calculation pressure and high maneuverability challenges in multi-star collaborative decision-making are solved, ensuring continuous coverage and positioning accuracy of the target.
Smart Images

Figure CN116280270B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and is an intelligent cooperative tracking method for a remote sensing constellation to an airborne moving target. Background Art
[0002] The space-based tracking problem of hypersonic glide vehicles (HGVs) has received considerable attention in recent years. Compared with ground-based early warning systems, the advantages of space-based early warning systems are wider space coverage, better tracking continuity, and no geographical location restrictions. The cooperative positioning problem of low-Earth orbit satellite constellations for hypersonic vehicles is the research focus of space-based early warning systems. The flight path of hypersonic vehicles has high mobility and uncertainty, which poses new challenges to the cooperative positioning ability of satellite constellations. In recent years, the breakthrough of artificial intelligence technology has provided a new way for multi-satellite cooperative autonomous intelligent decision-making technology. Deep reinforcement learning is an effective method to solve sequential decision-making problems. It continuously updates its own decision-making network through the interaction between the agent and the environment, and can effectively solve the problem of difficult acquisition of sample data in satellite constellation early warning systems. Currently, the commonly used deep reinforcement learning algorithms mainly include: Deep Q-Network (DQN), Deep Deterministic Policy Gradient (DDPG), Proximal Policy Optimization (PPO), and Soft Actor-Critic (SAC), etc. The decision-making process of large-scale satellite constellations has cooperation and continuity, and the state information of satellites, sensors, and targets changes in real time. Therefore, the solution space dimension of the task decision-making algorithm is very high. Summary of the Invention
[0003] Aiming at the multi-satellite cooperative tracking problem of HGVs by remote sensing constellations, the present invention designs a multi-objective staring tracking model and a two-satellite geometric positioning model to achieve the tracking and positioning of HGVs. Based on this, the present invention provides an intelligent cooperative tracking method for a remote sensing constellation to an airborne moving target.
[0004] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0005] The present invention provides an intelligent cooperative tracking method for a remote sensing constellation to an airborne moving target, and the present invention provides the following technical solutions:
[0006] An intelligent cooperative tracking method for a remote sensing constellation to an airborne moving target, the method comprising the following steps:
[0007] Step 1: Establish a staring and tracking model of a remote sensing constellation for multiple near-space targets, including constellation configuration and constraint models;
[0008] Step 2: Establish a two-satellite geometric positioning model for near-space aircraft;
[0009] Step 3: Based on the constellation configuration in the multi-target staring and tracking model, considering constraint limitations, and taking the two-satellite geometric positioning accuracy as the optimization goal, realize the tracking and positioning of the remote sensing constellation for near-space aircraft.
[0010] Preferably, the specific content of Step 1 is as follows:
[0011] Step 1.1: Conduct inter-satellite communication constraints. Adjacent satellites adopt laser communication technology, considering the lowest link altitude H c , and the maximum inter-satellite communication distance is expressed by the following formula:
[0012]
[0013]
[0014] where H s is the satellite orbit altitude, R e is the length of the Earth's radius, is the maximum geocentric angle;
[0015] Step 1.2: Establish infrared sensor constraints. The covered airspace of the infrared sensor is determined by the field of view angle α field , the maximum detection distance the satellite orbit altitude H s and the target flight altitude H tar jointly. It is related to the payload capacity and the target infrared radiation intensity. The maximum detection distance is expressed by the following formula:
[0016]
[0017] where D0 is the sensor aperture, D * is the sensor detectivity (m·Hz 1 / 2 ·W -1 ), τ a is the environmental transmittance, τ o is the transmittance of the optical system, Δ is the sensor signal process factor, A d is the detection unit area (m 2 ), Δf is the noise equivalent bandwidth (Hz), SNR min is the minimum signal-to-noise ratio required for the sensor to detect the target, and I is the target infrared intensity captured by the sensor;
[0018] The infrared radiation intensity is expressed by the following formula:
[0019]
[0020] where A is the infrared radiation area of the target (m 2 ), λ1 and λ2 are the lower and upper limits of the infrared band, ε is the spectral emissivity of the target surface, C1 is the first radiation constant (W·m 2 ), and C2 is the second radiation constant (m·K);
[0021] Step 1.3: Determine the stagnation temperature, which is described by the following formula:
[0022]
[0023] where T0 is the ambient temperature at the location of the target, ν is the atmospheric adiabatic index, β is the heat transfer recovery coefficient, and M is the Mach number of the target.
[0024] Preferably, the specific steps of step 2 are as follows:
[0025] Step 2.1: Set the positions of two satellites for collaborative observation as [X a Y a Z a and [X c Y c Z c . Calculate the farthest detectable points [X b Y b Z b and [X d Y d Z d within the satellite line of sight through the satellite positions and the angle measurement information. First, calculate the common perpendicular of the skew lines, divide the common perpendicular proportionally, estimate the target position, and the positioning is expressed by the following formula:
[0026]
[0027]
[0028] where F and t are intermediate variables, and [X tar Y tar Z tar is the estimated target position;
[0029] Step 2.2: The image plane measurement error and Euler angle measurement error of the constellation infrared sensor are equivalent to the angle measurement error in the two-dimensional plane. Project the satellite line of sight onto the plane with the common perpendicular as the normal, and the angle measurement is expressed as:
[0030]
[0031] The projected coordinates of the target in the plane are [x, y], and (x1, y1) and (x2, y2) are the projected coordinates of two satellites in the plane respectively:
[0032]
[0033] The observation matrix H for converting the angle measurement error to the target positioning error θ is:
[0034]
[0035] where L is the distance between the two satellites;
[0036] The geometric positioning accuracy is calculated as follows:
[0037]
[0038] In the formula, σ θ is the angle measurement error.
[0039] Preferably, step 3 is specifically:
[0040] Step 3.1: Set the environmental state space. The environmental state space S is defined as S = {T num , Obs angle , Pos tar , Vel tar , L tar , L score}, where T num is the number of targets, is the observation angle in the antenna coordinate system, n is the satellite number; Pos tar is the target position, Vel tar is the target velocity; is the list of satellites tracking the target, is the number of the target pointed by the sensor of satellite sat n , is the satellite return list, is the satellite sat n the return obtained when the sensor points to the target ID tar , L tar and L score have the same dimension as the number of satellites. When satellite sat n does not observe the target
[0041] Step 3.2: Establish the behavior space. The behavior space A of the satellite is A = {Tar1, Tar2,..., Tar k} is described as the target pointed by the sensor at time t; A is a k-dimensional vector, Tar kThe initial value is -1; since in the action space, in addition to the target direction, there is also an action to sweep back to the initial position, when the number of visible targets of the satellite is less than k - 1, A contains the directions of all visible targets;
[0042] Step 3.3: Establish a reward function, which is used to evaluate the value of the satellite's actions; r t Through a t , s t It is calculated as follows:
[0043]
[0044] Among them, C m and C a are the observation angle weight and the sensor rotation weight respectively, is the maximum detection distance of the sensor, is the distance between the satellite and the target, is the rotation angle of the sensor from the current direction to the target direction, ω max is the maximum rotation angular velocity of the sensor, and the visibility parameter I vis is expressed by the following formula:
[0045]
[0046] Step 3.4: Perform network update. The gradient update equation of the decision network is expressed by the following formula:
[0047]
[0048]
[0049] Among them, the policy π θ (a t |s t ) represents the probability that the agent selects the action a t after obtaining the environmental state s t at time step t; is the average value of the samples at time step t, is the generalized advantage function at time step t, γ is the discount factor, λ is the discount coefficient of the generalized advantage estimation, r t is the current reward, and V is the evaluation value of the critic network;
[0050] Establish the objective function of the policy gradient, and the objective function is expressed by the following formula:
[0051]
[0052] Among them, is the importance weight; the update equation of the policy gradient is:
[0053]
[0054] Among them, α is the parameter update step size;
[0055] The MAPPO algorithm training and parameter update adopt an experience sharing mechanism, and use global information to train the independent decision-making strategies of each satellite.
[0056] Preferably, when the number of visible targets is greater than k - 1, A includes the directions of the top k - 1 targets in the visible targets arranged in descending order of return value. The satellite comprehensively considers the current pointing of the sensor and the environmental state, and decides the pointing of the sensor at the next moment.
[0057] Preferably, the importance weight is limited to the range of (1 - ε, 1 + ε) by the hyperparameter ε and the truncation operation CLIP.
[0058] An intelligent cooperative tracking system for air moving targets by a remote sensing constellation, the system includes:
[0059] A staring tracking module, the staring tracking module establishes a staring tracking model for multiple targets in the near space by the remote sensing constellation, including the constellation configuration and the constraint model;
[0060] A dual-star geometric positioning module, the dual-star geometric positioning module establishes a dual-star geometric positioning model for near space aircraft;
[0061] A tracking and positioning module, the tracking and positioning module is based on the constellation configuration in the multi-target staring tracking model, considers the constraint limitations, and takes the dual-star geometric positioning accuracy as the optimization goal to realize the tracking and positioning of the near space aircraft by the remote sensing constellation.
[0062] Preferably, during the tracking process of HGVs, multiple satellites cooperate in observation to ensure continuous coverage of the target. More than two satellites track the same target at the same time to meet the positioning requirements; since the satellites in the constellation early warning system are all low-orbit satellites, both the satellites and HGVs are in a high-speed motion state, and the spatio-temporal relationship between the satellites and the target has high dynamics. During the flight of the target, multiple satellites need to relay tracking to achieve continuous multiple coverage of the target. For the constellation early warning system, HGVs are non-cooperative targets, and the position and speed information are obtained by satellite observation and message sharing.
[0063] A computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to be used to implement an intelligent cooperative tracking method for air moving targets by a remote sensing constellation.
[0064] A computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, an intelligent cooperative tracking method for a remote sensing constellation to track moving targets in the air is implemented.
[0065] The present invention has the following beneficial effects:
[0066] Compared with the prior art, the present invention:
[0067] The present invention aims at the tracking and positioning of hypersonic glide vehicles (HGVs) by a remote sensing constellation, which is a key issue in space-based early warning. The cooperative positioning process of the remote sensing constellation for HGVs is a multi-to-multi dynamic task allocation process. The task allocation of early warning satellites has temporal relevance and dynamics, so the decision-making algorithm will face the problems of a large solution space and being easily trapped in local optimality. At the same time, the high maneuverability of HGVs poses a new challenge to the real-time performance of the task decision-making algorithm of the remote sensing constellation.
[0068] Traditional task decision-making algorithms for remote sensing satellites usually adopt a task decision-making method in which a central node satellite makes decisions and multiple satellites cooperate to execute. Such methods impose a relatively large computational pressure on the central node. In complex tracking scenarios, it is difficult to handle the tracking problem of HGVs with uncertainties.
[0069] The present invention designs a staring tracking model for a remote sensing constellation to multiple targets in the near space. The tracking process of the remote sensing constellation for near space vehicles is described from three aspects: constellation configuration method, inter-satellite communication constraint, and infrared sensor constraint.
[0070] The present invention designs a two-satellite geometric positioning model for near space flight. The position of the target is calculated through the angle measurement information of two satellite targets, and considering the influence of angle measurement error on target positioning, a geometric positioning accuracy factor is designed to measure the effect of two-satellite positioning.
[0071] The present invention designs a task decision-making method for a remote sensing constellation based on reinforcement learning. An environmental state space, a behavior space are established and a reward function is designed, etc., to achieve independent real-time intelligent decision-making of the remote sensing constellation. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0073] Figure 1 It is a schematic diagram of the detection range of an infrared sensor;
[0074] Figure 2 Schematic diagram of the staring tracking model of the remote sensing constellation
[0075] Figure 3 Overall process of intelligent decision-making of the remote sensing constellation
[0076] Figure 4 Schematic diagram of double-star positioning
[0077] Figure 5 Parameter update process Specific implementation manners
[0078] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0079] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus should not be construed as a limitation of the present invention. In addition, the terms "first", "second", "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0080] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0081] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0082] The present invention is described in detail below with specific embodiments. Specific Embodiment 1:
[0084] According to Figures 1 to 5 shown, the specific optimized technical solution adopted by the present invention to solve the above technical problems is: The present invention relates to an intelligent cooperative tracking method for an air moving target by a remote sensing constellation.
[0085] An intelligent cooperative tracking method for an aerial moving target by a remote sensing constellation, the method comprising the following steps:
[0086] Step 1: Establish a staring tracking model for multiple targets in the near space by the remote sensing constellation, including constellation configuration and constraint models;
[0087] The specific content of Step 1 is as follows:
[0088] Step 1.1: Conduct inter-satellite communication constraints. Two adjacent satellites adopt laser communication technology, considering the lowest link height H c , and the maximum inter-satellite communication distance is expressed by the following formula:
[0089]
[0090]
[0091] where H s is the satellite orbital height, R e is the radius length of the earth, is the maximum geocentric angle;
[0092] Step 1.2: Establish infrared sensor constraints. The covered airspace of the infrared sensor is determined jointly by the field of view angle α field , the maximum detection distance the satellite orbital height H s and the target flight height H tar . It is related to the payload capacity and the target infrared radiation intensity, and the maximum detection distance is expressed by the following formula:
[0093]
[0094] where D0 is the sensor aperture, D * is the sensor detectivity (m·Hz 1 / 2 ·W -1 ), τ a is the environmental transmittance, τ o is the transmittance of the optical system, Δ is the sensor signal process factor, A d is the detection unit area (m 2 ), Δf is the noise equivalent bandwidth (Hz), SNR min is the minimum signal-to-noise ratio required for the sensor to detect the target, and I is the infrared intensity of the target captured by the sensor;
[0095] The infrared radiation intensity is expressed by the following formula:
[0096]
[0097] Among them, A is the infrared radiation area of the target (m 2 ), λ1 and λ2 are the lower and upper limits of the infrared band, ε is the spectral emissivity of the target surface, C1 is the first radiation constant (W·m 2 ), and C2 is the second radiation constant (m·K);
[0098] Step 1.3: Determine the stagnation temperature, which is described by the following formula:
[0099]
[0100] Among them, T0 is the environmental temperature at the location of the target, ν is the atmospheric adiabatic index, β is the heat transfer recovery coefficient, and M is the Mach number of the target.
[0101] Step 2: Establish a dual-star geometric positioning model for near-space vehicles;
[0102] The specific steps of Step 2 are as follows:
[0103] Step 2.1: Set the positions of the two satellites for cooperative observation as [X a Y a Z a and [X c Y c Z c , calculate the farthest detectable points [X b Y b Z b and [X d Y d Z d within the satellite line of sight through the satellite positions and angle measurement information. First, calculate the common perpendicular of the skew lines, divide the common perpendicular proportionally, estimate the target position, and the positioning is expressed by the following formula:
[0104]
[0105]
[0106] Among them, F and t are intermediate variables, and [X tar Y tar Z tar is the estimated target position;
[0107] Step 2.2: Equivalent the image plane measurement error and Euler angle measurement error of the constellation infrared sensor to the angle measurement error in the two-dimensional plane, project the satellite line of sight onto the plane with the common perpendicular as the normal, and the angle measurement is expressed as:
[0108]
[0109] The projected coordinates of the target in the plane are [x, y], and (x1, y1) and (x2, y2) are the projected coordinates of two satellites in the plane respectively:
[0110]
[0111] The observation matrix H for converting the angle measurement error to the target positioning error θ is:
[0112]
[0113] where L is the distance between the two satellites;
[0114] The geometric positioning accuracy is calculated as follows:
[0115]
[0116] In the formula, σ θ is the angle measurement error.
[0117] Step 3: Based on the constellation configuration in the multi-target staring tracking model, considering the constraint limitations, with the geometric positioning accuracy of the double satellites as the optimization goal, realize the tracking and positioning of the remote sensing constellation for the near-space vehicle.
[0118] The specific content of the said Step 3 is:
[0119] Step 3.1: Set the environmental state space. The environmental state space S is defined as S = {T num , Obs angle , Pos tar , Vel tar , L tar , L score}, where T num is the number of targets, is the observation angle in the antenna coordinate system, n is the satellite number; Pos tar is the target position, Vel tar is the target velocity; is the list of satellites tracking the target, is the number of the target pointed by the sensor of satellite sat n , is the satellite return list, is the return obtained when the sensor of satellite sat n points to the target ID tar , and the dimensions of L tar and L score are the same as the number of satellites. When satellite sat n does not observe the target
[0120] Step 3.2: Establish the behavior space. The behavior space A of the satellite = {Tar1, Tar2,..., Tar k} is described as the target pointed by the sensor at time t; A is a k-dimensional vector, and the initial value of Tar k is -1; Since the behavior space includes not only the target direction but also a behavior of sweeping back to the initial position, when the number of visible targets of the satellite is less than k - 1, A includes the directions of all visible targets;
[0121] Step 3.3: Establish the reward function. The reward function is used to evaluate the value of the satellite's behavior; r t is obtained by calculating through a t , s t . The specific calculation is as follows:
[0122]
[0123] where C m and C a are the observation angle weight and the sensor rotation weight respectively, is the maximum detection distance of the sensor, is the distance between the satellite and the target, is the rotation angle of the sensor from the current direction to the target direction, ω max is the maximum rotation angular velocity of the sensor, and the visibility parameter I vis is expressed by the following formula:
[0124]
[0125] Step 3.4: Perform network update. The decision network gradient update equation is expressed by the following formula:
[0126]
[0127]
[0128] where the policy π θ (a t |s t ) represents the probability that the agent selects the behavior a t at time step t after obtaining the environmental state s t ; is the average value sampled at time step t, is the generalized advantage function at time step t, γ is the discount factor, λ is the discount coefficient of the generalized advantage estimation, r t is the current reward, and V is the evaluation value of the critic network;
[0129] Establish the objective function of the policy gradient, and the objective function is expressed by the following formula:
[0130]
[0131] Among them, is the importance weight; the update equation of the policy gradient is:
[0132]
[0133] Among them, α is the parameter update step size;
[0134] The MAPPO algorithm training and parameter update adopt an experience sharing mechanism, and use global information to train the independent decision-making strategies of each satellite.
[0135] When the number of visible targets is greater than k - 1, A contains the directions of the top k - 1 targets in the visible targets arranged in descending order of the return value. The satellite comprehensively considers the current pointing of the sensor and the environmental state, and decides the pointing of the sensor at the next moment.
[0136] The importance weight is restricted to the range of (1 - ε, 1 + ε) by the hyperparameter ε and the truncation operation CLIP.
[0137] The present invention provides an intelligent cooperative tracking system for an airborne moving target by a remote sensing constellation. The system includes:
[0138] A staring tracking module, which establishes a staring tracking model for multiple targets in the near space by the remote sensing constellation, including a constellation configuration and a constraint model;
[0139] A double-star geometric positioning module, which establishes a double-star geometric positioning model for near-space aircraft;
[0140] A tracking and positioning module, which is based on the constellation configuration in the multi-target staring tracking model, considers constraint limitations, and takes the double-star geometric positioning accuracy as the optimization goal to realize the tracking and positioning of the near-space aircraft by the remote sensing constellation.
[0141] During the tracking process of HGVs, multiple satellites cooperate in observation to ensure continuous coverage of the target. More than two satellites track the same target at the same time to meet the positioning requirements; since the satellites in the constellation early warning system are all low-orbit satellites, both the satellites and HGVs are in a high-speed motion state. The spatio-temporal relationship between the satellites and the target has high dynamics. Multiple satellites need to relay tracking during the flight of the target to achieve continuous multiple coverage of the target. For the constellation early warning system, HGVs are non-cooperative targets, and the position and speed information are obtained by satellite observation and message sharing.
[0142] The present invention provides a computer-readable storage medium, on which a computer program is stored, and the program is executed by a processor to implement an intelligent cooperative tracking method for a remote sensing constellation to an airborne moving target.
[0143] The present invention provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, an intelligent cooperative tracking method for a remote sensing constellation to an airborne moving target is implemented. Specific Embodiment 2:
[0145] The difference between the second embodiment and the first embodiment of this application is only that:
[0146] The steps of a remote sensing constellation mission decision method based on reinforcement learning are as follows:
[0147] Establish a staring tracking model for a remote sensing constellation to multiple near-space targets, including a constellation configuration and a constraint model;
[0148] Establish a two-star geometric positioning model for near-space aircraft;
[0149] Design a remote sensing constellation mission decision method based on reinforcement learning. Based on the constellation configuration in the multi-target staring tracking model, considering constraint limitations, and taking the two-star geometric positioning accuracy as the optimization goal, the tracking and positioning of near-space aircraft by the remote sensing constellation are realized;
[0150] Specific implementation measures
[0151] The following further describes the specific implementation manners of the present invention in detail.
[0152] The tracking method of the remote sensing constellation of the present invention for hypersonic aircraft is carried out in the following steps:
[0153] In step 1), establish a staring tracking model for a remote sensing constellation to multiple near-space targets:
[0154] Constellation Configuration: The remote sensing constellation involved in the present invention takes hypersonic vehicles as observation targets and needs to achieve global all-weather coverage characteristics. Therefore, the Walker-δ constellation configuration is suitable for selection. This constellation configuration consists of multiple satellites with the same orbital inclination and the same orbital altitude. Moreover, the phases of the satellites within each orbital plane of the constellation are evenly distributed, and the ascending nodes between the orbital planes are evenly distributed. The Walker-δ constellation configuration can be expressed as N / P / F, where N is the number of satellites, P is the number of orbital planes, and F is the phase difference between adjacent orbital planes. Near space refers to the airspace at an altitude of 20 - 100 kilometers from the ground. The spaceborne infrared sensor adopts the limb observation method, so its visible time window for near-space vehicles is relatively short. When two or more infrared sensors simultaneously cover a target, its position can be determined. Therefore, the constellation of the star cluster early warning system needs to meet the requirement that the coverage multiplicity of near-space targets in different regions of the world can reach more than 2 folds.
[0155] Inter-satellite Communication Constraint:
[0156] Adjacent satellites adopt laser communication technology. To reduce link loss, it is necessary to avoid passing through the atmosphere. Therefore, the minimum link height H needs to be considered c . The maximum inter-satellite communication distance can be expressed as:
[0157]
[0158] In the formula, H s is the satellite orbital altitude, R e is the length of the Earth's radius, is the maximum geocentric angle.
[0159]
[0160] Infrared Sensor Constraint:
[0161] Affected by the infrared radiation of the Earth and the atmosphere, the sensor needs to meet geometric visibility constraints, limb observation constraints, maximum detection distance constraints, etc. The covered airspace of the infrared sensor is determined by the field of view angle α field , the maximum detection distance the satellite orbital altitude H s and the target flight altitude H tar jointly. The geometric visibility constraint means that the connection line between the satellite and the target cannot be blocked by the Earth or other obstacles. The limb observation constraint means that the deep space background should be maintained within the observation field of view of the satellite, and there should be no infrared radiation objects such as the Earth or the Sun in the field of view. is related to the payload capacity and the target infrared radiation intensity, and is described as follows:
[0162]
[0163] where D0 is the sensor aperture, D * is the sensor detectivity (m·Hz 12 ·W -1 , τ a is the environmental transmittance, τ o is the transmittance of the optical system, Δ is the sensor signal process factor, A d is the area of the detection unit (m 2 ), Δf is the noise equivalent bandwidth (Hz), SNR min is the minimum signal-to-noise ratio required for the sensor to detect the target, and I is the infrared intensity of the target captured by the sensor. The HGVs studied in this paper enter the near space in an unpowered gliding manner, and infrared radiation is generated by the friction between the skin and the air. The infrared radiation intensity can be expressed as:
[0164]
[0165] where A is the infrared radiation area of the target (m 2 ), λ1 and λ2 are the lower and upper limits of the infrared band, ε is the spectral emissivity of the target surface, C1 is the first radiation constant (W·m 2 ), and C2 is the second radiation constant (m·K).
[0166] The stagnation point temperature T is described as follows:
[0167]
[0168] where T0 is the environmental temperature at the target location, ν is the atmospheric adiabatic index, β is the heat transfer recovery coefficient, and M is the target Mach number.
[0169] During the tracking process of the HGVs by the constellation early warning system, multiple satellites cooperate in observation to ensure continuous coverage of the target. More than two satellites need to track the same target simultaneously at the same moment to meet the positioning requirements. Since the satellites in the constellation early warning system are all low-earth orbit satellites, both the satellites and the HGVs are in a high-speed motion state. The spatio-temporal relationship between the satellites and the target has high dynamics. Multiple satellites need to relay tracking during the flight of the target to achieve continuous multiple coverage of the target. For the constellation early warning system, the HGVs are non-cooperative targets. Their position, velocity and other information are obtained by satellite observation and message sharing.
[0170] Step 2) Considering the staring tracking model of the remote sensing constellation for multiple targets in the near space established in Step 1), design a two-satellite geometric positioning method:
[0171] In order for the satellites in the constellation early warning system to obtain the position and motion information of the target, at least two satellites are required to track the same target simultaneously. Therefore, multi-satellite multi-angle tracking and observation of the target is the core task of the mission decision-making of the constellation early warning system and also the core purpose of collaborative decision-making among satellites.
[0172] The satellite positioning method is as follows: Assume that the positions of two satellites for collaborative observation are [X a Y a Z a and [X c Y c Z c . Calculate the farthest detectable points [X b Y b Z b and [X d Y d Z d within the satellite line of sight through the satellite positions and angle measurement information. Since the observation directions of the two satellites caused by angle measurement errors are usually skew lines. First, calculate the common perpendicular of the skew lines, and then divide the common perpendicular proportionally to estimate the target position. The positioning algorithm is as follows:
[0173]
[0174]
[0175] In the formula, F and t are intermediate variables, and [X tar Y tar Z tar is the estimated target position.
[0176] The image plane measurement error and Euler angle measurement error of the constellation infrared sensor can be equivalent to the angle measurement error in the two-dimensional plane. Project the satellite line of sight onto the plane with the common perpendicular as the normal, and the angle measurement positioning method in the two-dimensional plane is as Figure 4 shown.
[0177] The angle measurement is expressed as:
[0178]
[0179] The projected coordinates of the target in the plane are [x, y]. (x1, y1) and (x2, y2) are the projected coordinates of the two satellites in the plane respectively. From Figure 4 the geometric relationship can be seen:
[0180]
[0181] The observation matrix H θ from the angle measurement error to the target positioning error is:
[0182]
[0183] Wherein, L is the distance between two satellites.
[0184] The geometric positioning accuracy is calculated as follows:
[0185]
[0186] Wherein, σ θ is the angular measurement error.
[0187] Step 3): Based on the constellation configuration in Step 1), considering the constraint limitations, and taking the geometric positioning accuracy of the double satellites in Step 2) as the optimization objective, design a remote sensing constellation task decision-making method based on reinforcement learning:
[0188] Environmental state space: The environmental state space S is defined as S = {T num , Obs angle , Pos tar , Vel tar , L tar , L score}, T num is the number of targets, is the observation angle in the antenna coordinate system, and n is the satellite number. Pos tar is the target position, and Vel tar is the target velocity. is the list of satellites tracking the target, is the number of the target pointed by the sensor of satellite sat n . is the satellite reward list, is the reward obtained when the sensor of satellite sat n points to the target ID tar . The dimensions of L tar and L score are the same as the number of satellites. When satellite sat n does not observe the target
[0189] Action space: The action space A of the satellite = {Tar1, Tar2,..., Tar k} can be described as the target pointed by the sensor at time t. A is a k-dimensional vector, and the initial value of Tar k is -1. Since the action space contains a behavior of sweeping back to the initial position in addition to the target direction, when the number of visible targets of the satellite is less than k - 1, A contains the directions of all visible targets. When the number of visible targets is greater than k - 1, A contains the directions of the first k - 1 targets sorted in descending order of the reward value among the visible targets. The satellite comprehensively considers the current pointing of the sensor and the environmental state and decides the pointing of the sensor at the next moment.
[0190] Reward function: The reward function is used to evaluate the value of satellite behavior. r t Through a t , s t Calculate and obtain, the specific calculation is as follows:
[0191]
[0192] In the formula, C m and C a are the observation angle weight and the sensor rotation weight respectively. is the maximum detection distance of the sensor, is the distance between the satellite and the target. is the rotation angle of the sensor from the current direction to the target direction, ω max is the maximum rotation angular velocity of the sensor. The visibility parameter I vis is expressed as:
[0193]
[0194] Network update:
[0195] The gradient update equation of the decision network is:
[0196]
[0197] In the formula, the policy π θ (a t |s t ) represents the probability that the agent selects the action a t after obtaining the environmental state s t at time step t. is the average value of the sample at time step t. is the generalized advantage function at time step t.
[0198]
[0199] In the formula, γ is the discount factor, and λ is the discount coefficient of the generalized advantage estimation. r t is the current reward, and V is the evaluation value of the critic network.
[0200] The objective function of the policy gradient is:
[0201]
[0202] In the formula, This is the Importance Weight. The role of the Importance Weight is to adjust the return value through the ratio of the new and old policy gradients. The weight of behaviors that are more likely to be taken is increased, while the weight of behaviors that are less likely to be taken is decreased. The Importance Weight is restricted to the range (1 - ε, 1 + ε) by the hyperparameter ε and the clipping operation CLIP.
[0203] The update equation for the policy gradient is:
[0204]
[0205] In the formula, α is the parameter update step size.
[0206] The MAPPO algorithm uses an experience sharing mechanism for training and parameter update. Each satellite's independent decision-making policy is trained using global information. The parameter update process is as Figure 5 shown.
[0207] The algorithm includes N local policy networks and one global policy network. At each time step t, the algorithm calculates the behavioral policies of the N Actor networks and obtains the return value of the behavior after interacting with the environment Each local policy network calculates for T e steps. Then the advantages A t , t = 1, 2, 3,... N·T e are calculated. The N local networks share experiences Then the global network is updated using the Adam optimizer, and finally the new global network is copied to each local network.
[0208] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of these features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined. Any process or method description represented in a flowchart or otherwise described herein can be understood to represent a module, segment, or portion of code including one or more N executable instructions for implementing a customized logical function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present invention belong. The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in connection with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or N wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM).In addition, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing when necessary, and then storing it in a computer memory. It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0209] The above is only a preferred embodiment of an intelligent cooperative tracking method for an airborne moving target by a remote sensing constellation. The protection scope of an intelligent cooperative tracking method for an airborne moving target by a remote sensing constellation is not limited to the above embodiments. Any technical solutions falling within this concept belong to the protection scope of the present invention. It should be noted that for those skilled in the art, several improvements and changes made without departing from the principle of the present invention should also be regarded as within the protection scope of the present invention.
Claims
1. An intelligent collaborative tracking method for aerial moving targets by a remote sensing constellation, characterized in that: The method includes the following steps: Step 1: Establish a staring tracking model for multiple near-space targets by a remote sensing constellation, including constellation configuration and constraint models; Specifically, Step 1 is as follows: Step 1.1: Perform inter-satellite communication constraints. For two adjacent satellites, laser communication technology is adopted. Consider the lowest link altitude H c , and the maximum inter-satellite communication distance is expressed by the following formula: Among them, H s is the satellite orbit altitude, R e is the length of the Earth's radius, is the maximum geocentric angle; Step 1.2: Establish the constraints of the infrared sensor. The covered airspace of the infrared sensor is determined by the field of view angle α field , the maximum detection distance the satellite orbital altitude H s and the target flight altitude H tar jointly. It is related to the payload capacity and the target infrared radiation intensity. The maximum detection distance is expressed by the following formula: Among them, D0 is the sensor aperture, D * is the sensor detectivity (m·Hz 12 ·W -1 ), τ a is the environmental transmittance, τ o is the transmittance of the optical system, Δ is the sensor signal process factor, A d is the detection unit area (m 2 ), Δf is the noise equivalent bandwidth (Hz), SNR min is the minimum signal-to-noise ratio required for the sensor to detect the target, and I is the infrared intensity of the target captured by the sensor; The infrared radiation intensity is expressed by the following formula: Among them, A is the infrared radiation area of the target (m 2 ), λ1 and λ2 are the lower and upper limits of the infrared band, ε is the spectral emissivity of the target surface, C1 is the first radiation constant (W·m 2 ), and C2 is the second radiation constant (m·K); Step 1.3: Determine the stagnation temperature, which is described by the following formula: where, T0 is the environmental temperature at the target location, ν is the atmospheric adiabatic index, β is the heat transfer recovery coefficient, and M is the target Mach number; Step 2: Establish a two-satellite geometric positioning model for near-space aircraft; Specifically, Step 2 is as follows: Step 2.1: Set the positions of two satellites for collaborative observation as [X a Y a Z a and [X c Y c Z c . Calculate the farthest detectable points [X b Y b Z b and [X d Y d Z d within the satellite line of sight based on the satellite positions and angle measurement information. First, calculate the common perpendicular of skew lines, divide the common perpendicular proportionally, estimate the target position, and the positioning is expressed by the following formula: where F and t are intermediate variables, and [X tar Y tar Z tar is the estimated target position; Step 2.2: Equivalent the image plane measurement error and Euler angle measurement error of the constellation infrared sensor to the angle measurement error in the two-dimensional plane, project the satellite line of sight into the plane with the common perpendicular as the normal, and the angle measurement is expressed as: The projection coordinates of the target in the plane are [x, y], and (x1, y1) and (x2, y2) are the projection coordinates of the two satellites in the plane respectively: Observation matrix H for converting angular measurement error to target positioning error θ is as follows: where, L is the distance between the two satellites; The geometric positioning accuracy is calculated as follows: where σ θ is the angular measurement error; Step 3: Based on the constellation configuration in the multi-target staring tracking model, considering the constraint limitations, and taking the two-satellite geometric positioning accuracy as the optimization goal, realize the tracking and positioning of the remote sensing constellation for near-space aircraft; Specifically, Step 3 is as follows: Step 3.1: Set the environmental state space. The environmental state space S is defined as S = {T num , Obs angle , Pos tar , Vel tar , L tar , L score}, where T num is the number of targets, is the observation angle in the antenna coordinate system, and n is the satellite number; Pos tar is the target position, and Vel tar is the target velocity; is the list of targets tracked by the satellite, is the number of the target pointed by the sensor of satellite sat n , is the satellite return list, is the ID of the target pointed by the sensor of satellite sat n when the return is obtained, and the dimensions of L tar and L tar are the same as the number of satellites. When satellite sat score does not observe the target, n Step 3.2: Establish the behavior space. The behavior space A of the satellite = {Tar1, Tar2,..., Tar k} is described as the target pointed to by the sensor at time t; A is a k-dimensional vector, and the initial value of Tar k is -1; Since the behavior space includes not only the target direction but also a behavior of scanning back to the initial position, when the number of visible targets of the satellite is less than k - 1, A includes the directions of all visible targets; Step 3.3: Establish a reward function, which is used to evaluate the value of satellite behavior; r t Through a t , s t It is calculated as follows: Among them, C m and C a are the observation angle weight and the sensor rotation weight respectively, is the maximum detection distance of the sensor, is the distance between the satellite and the target, is the rotation angle of the sensor from the current direction to the target direction, ω max is the maximum rotation angular velocity of the sensor, and the visibility parameter I vis is expressed by the following formula: Step 3.4: Perform network update, and the decision network gradient update equation is expressed by the following formula: Among them, the policy π θ (a t |s t ) represents the probability of the agent selecting action a t after obtaining the environmental state s t at time step t; is the average value sampled at time step t, is the generalized advantage function at time step t, γ is the discount factor, λ is the discount coefficient for generalized advantage estimation, r t is the current return, and V is the evaluation value of the critic network; Establish the objective function of the policy gradient, and the objective function is expressed by the following formula: Among them, is the importance weight; the update equation of the policy gradient is: where, α is the parameter update step size; The MAPPO algorithm training and parameter update adopt the experience sharing mechanism, and use the global information to train the independent decision-making strategies of each satellite.
2. The method according to claim 1, characterized in that: When the number of visible targets is greater than k - 1, A contains the directions of the first k - 1 targets arranged in descending order of return values among the visible targets. The satellite comprehensively considers the current pointing of the sensor and the environmental state, and decides the sensor pointing at the next moment.
3. The method according to claim 2, characterized in that: The importance weight is limited to the range of (1 - ε, 1 + ε) by the hyperparameter ε and the truncation operation CLIP.
4. An intelligent cooperative tracking system for airborne moving targets by a remote sensing constellation, characterized in that: The system includes: A staring tracking module, which establishes a staring tracking model for multiple near-space targets by a remote sensing constellation, including constellation configuration and constraint models; Inter-satellite communication constraints are imposed. For two adjacent satellites, laser communication technology is adopted, and the lowest link altitude H is considered. c , and the maximum inter-satellite communication distance is expressed by the following formula: Among them, H s is the satellite orbital altitude, R e is the length of the Earth's radius, is the maximum geocentric angle; Establish infrared sensor constraints. The covered airspace of the infrared sensor is determined by the field of view angle α field , the maximum detection distance satellite orbital altitude H s and the target flight altitude H tar jointly determine, related to the payload capacity and the target infrared radiation intensity. The maximum detection distance is expressed by the following formula: Among them, D0 is the sensor aperture, D * is the sensor detectivity (m·Hz 12 ·W -1 ), τ a is the environmental transmittance, τ o is the transmittance of the optical system, Δ is the sensor signal process factor, A d is the detection unit area (m 2 ), Δf is the noise equivalent bandwidth (Hz), SNR min is the minimum signal-to-noise ratio required for the sensor to detect the target, I is the infrared intensity of the target captured by the sensor; The infrared radiation intensity is expressed by the following formula: Among them, A is the infrared radiation area of the target (m 2 ), λ1 and λ2 are the lower and upper limits of the infrared band, ε is the spectral emissivity of the target surface, C1 is the first radiation constant (W·m 2 ), and C2 is the second radiation constant (m·K); Determine the stagnation temperature, which is described by the following formula: where, T0 is the environmental temperature at the target location, ν is the atmospheric adiabatic index, β is the heat transfer recovery coefficient, and M is the target Mach number; A two-satellite geometric positioning module, which establishes a two-satellite geometric positioning model for near-space aircraft; Set the positions of two satellites for collaborative observation as [X a Y a Z a and [X c Y c Z c , and calculate the farthest detectable points [X b Y b Z b and [X d Y d Z d within the satellite line of sight based on the satellite positions and angular measurement information. First, calculate the common perpendicular of skew lines, divide the common perpendicular proportionally, estimate the target position, and the positioning is expressed by the following formula: where F and t are intermediate variables, and [X tar Y tar Z tar is the estimated target position; Equivalent the image plane measurement error and Euler angle measurement error of the constellation infrared sensor to the angle measurement error in the two-dimensional plane, project the satellite line of sight into the plane with the common perpendicular as the normal, and the angle measurement is expressed as: The projection coordinates of the target in the plane are [x, y], and (x1, y1) and (x2, y2) are the projection coordinates of the two satellites in the plane respectively: Observation matrix H for converting angular measurement error to target positioning error θ is as follows: where, L is the distance between the two satellites; The geometric positioning accuracy is calculated as follows: where σ θ is the angular measurement error; A tracking and positioning module, which is based on the constellation configuration in the multi-target staring tracking model, considers the constraint limitations, and takes the two-satellite geometric positioning accuracy as the optimization goal, and realizes the tracking and positioning of the remote sensing constellation for near-space aircraft; Set up the environmental state space. The environmental state space S is defined as S = {T num , Obs angle , Pos tar , Vel tar , L tar , L score}, where T num is the number of targets, is the observation angle in the antenna coordinate system, and n is the satellite number; Pos tar is the target position, and Vel tar is the target velocity; is the list of targets tracked by the satellite, is the number of the target pointed by the sensor of satellite sat n , and is the satellite return list, is the return obtained when the sensor of satellite sat n points to the target ID tar . The dimensions of L tar and L score are the same as the number of satellites. When satellite sat n does not observe the target, Establish a behavior space. The behavior space A of the satellite = {Tar1, Tar2,..., Tar k} is described as the target pointed to by the sensor at time t; A is a k-dimensional vector, and the initial value of Tar k is -1; Since the behavior space includes not only the target direction but also a behavior of scanning back to the initial position, when the number of visible targets of the satellite is less than k - 1, A includes the directions of all visible targets; Step 3.3: Establish a reward function, which is used to evaluate the value of satellite behavior; r t through a t , s t is calculated as follows: Among them, C m and C a are the observation angle weight and the sensor rotation weight respectively, is the maximum detection distance of the sensor, is the distance between the satellite and the target, is the rotation angle of the sensor from the current direction to the target direction, ω max is the maximum rotation angular velocity of the sensor, and the visibility parameter I vis is expressed by the following formula: Perform network update, and the decision network gradient update equation is represented by the following formula: Among them, strategy π θ (a t |s t ) means that at time step t, the agent obtains the environment state s t Post-selection behavior t probability; is the average value of the samples at time step t, is the generalized advantage function at time step t, γ is the discount factor, λ is the discount coefficient of the generalized advantage estimate, r t is the current return, V is the evaluation value of the critic network; Establish the objective function of the policy gradient, and the objective function is represented by the following formula: Among them, is the importance weight; the update equation of the policy gradient is as follows: where α is the parameter update step size; The MAPPO algorithm uses an experience sharing mechanism for training and parameter update, and uses global information to train the independent decision-making strategies of each satellite.
5. The system according to claim 4, characterized in that: During the tracking process of HGVs, multiple satellites cooperate in observation to ensure continuous coverage of the target. More than two satellites need to track the same target at the same time to meet the positioning requirements; since the satellites in the constellation early warning system are all low-earth orbit satellites, both the satellites and HGVs are in a high-speed motion state, and the spatio-temporal relationship between the satellites and the target is highly dynamic. Multiple satellites need to relay the tracking of the target during its flight to achieve continuous multiple coverage of the target. For the constellation early warning system, HGVs are non-cooperative targets, and their position and velocity information are obtained through satellite observation and message sharing.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by a processor to implement the method according to any one of claims 1-3.
7. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that: When the processor executes the computer program, the method according to any one of claims 1-3 is implemented.
Citation Information
Patent Citations
Multi-satellite coordinated real-time tracking method for spatial dynamic target
CN110412869A
Satellite constellation in-orbit distributed cooperative scheduling method
CN115535297A