Time-sensitive Target Observation Method and System Based on Cooperative Sensing of Multiple Satellite Platforms

Through the method of collaborative perception of multi-satellite platforms and Bayesian inference theory combined with greedy algorithms, the problem of time-sensitive target positioning and tracking is solved, and the time-sensitive target is detected in the shortest time, improving the collaborative perception and autonomous planning capabilities of multi-satellite platforms.

CN115730712BActive Publication Date: 2025-06-27HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211424866.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2025-06-27
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

The prior art is difficult to quickly locate and track time-sensitive targets, especially when their potential area has large geometric dimensions, complex geographical structures and agnostic movement laws.

Method used

The method based on multi-satellite platform collaborative perception is adopted to predict the time-sensitive target state through Bayesian inference theory, and a greedy algorithm is used to solve the satellite platform observation strategy to achieve time-sensitive target detection in the shortest time.

Benefits of technology

The time-sensitive targets are detected in the shortest time, the collaborative perception and autonomous planning capabilities of multi-star platforms are improved, and the need to maximize mission efficiency under limited satellite resources is met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115730712B_ABST
    Figure CN115730712B_ABST
Patent Text Reader

Abstract

The present invention provides a method, system, storage medium and electronic device for observing time-sensitive targets based on multi-satellite platform collaborative perception, which relates to the technical field of satellite observation. The present invention regards the detection process of time-sensitive targets as a partially observable Markov decision process, predicts the state position of time-sensitive targets in a deterministic manner through Bayesian inference theory, and solves the satellite platform observation strategy with a step size of 1 through a greedy algorithm, and selects an action strategy to ensure the maximum probability of discovering time-sensitive targets; through the search of a rolling horizon periodic observation strategy with a step size of N, the optimal action strategy of the satellite platform is selected within the rolling horizon period N, so as to detect time-sensitive targets in the shortest time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of satellite observation, and particularly relates to a time-sensitive target observation method, system, storage medium and electronic device based on multi-satellite platform collaborative perception. Background Art

[0002] A time-sensitive target is short for a time-sensitive object, referring to a target that appears within a certain time window. Time-sensitive targets have characteristics such as difficult positioning, difficult tracking, and strong timeliness. The real-time perception ability of time-sensitive targets is an important measure of the development level of military technology. Multi-satellite platform collaborative perception is the most effective means to achieve the positioning and surveillance of time-sensitive targets. At the same time, however, the detection of time-sensitive targets also poses extremely high requirements for the intelligent task planning of multi-satellite platform collaborative perception, urgently demanding to improve the multi-satellite platform collaborative perception ability and enhance the autonomous planning ability to maximize the task efficiency with limited satellite resources.

[0003] Due to the large geometric size and complex geographical structure of the potential area of time-sensitive targets; they are themselves mobile, and their movement patterns are unknown and uncertain; different from the observation of stationary point targets and regional targets, the search for moving targets requires a certain time resolution requirement.

[0004] Therefore, the key to time-sensitive target detection is to use sensor network technology to achieve the tracking, monitoring and rapid positioning of mobile targets. Most of the existing earth observation technologies complete the observation tasks when the observation requirements are known and the geographical location of the observation target is clear. Due to the characteristics of difficult positioning, difficult tracking and strong timeliness of time-sensitive targets, the existing earth observation technologies are difficult to meet the requirements. Summary of the Invention

[0005] (1) Technical Problems to be Solved

[0006] In view of the deficiencies of the prior art, the present invention provides a time-sensitive target observation method, system, storage medium and electronic device based on multi-satellite platform collaborative perception, and solves the technical problem of being unable to quickly locate time-sensitive targets.

[0007] (2) Technical Solutions

[0008] To achieve the above object, the present invention is realized through the following technical solutions:

[0009] A time-sensitive target observation method based on multi-satellite platform collaborative perception, comprising:

[0010] S1. It is known that the state of the multi-satellite platform at time k is Store the possible states of the multi-satellite platform at time k + 1 into the queue Initialize reward1 = 0; where, given a multi-satellite platform set Q = {S1, S2,... S q}; num represents the possible state set of the multi-satellite platform at time k + 1, represents the satellite S i in the num-th possible state at time k + 1, k = 1, 2,..., N; reward1 represents the threshold for the action strategy to detect time-sensitive targets;

[0011] S2. Let j = 1. If the queue queue is empty, the iteration ends; otherwise, select the first state variable in the queue queue as Obtain the state transition of the multi-satellite platform set from time k to time k + 1 as Then remove this element from the queue queue and transfer to S3;

[0012] S3. Let j = j + 1, initialize an empty list adjnodes, generate the set of all possible state transitions of the multi-satellite platform from time k to time k + j, and put all elements into the list adjnodes, Transfer to S4; where numl is the number of its elements, represents the num1-th possible state transition from time k to time k + j;

[0013] S4. If j ≥ N or adjnodes is empty, calculate R for all possible cases within the time interval j according to the short-term effectiveness evaluation function Set N and use the one-step greedy algorithm to obtain the instantaneous effectiveness evaluation function Let the confidence cumulative return expectation function value r of the collaborative action strategy of the multi-satellite platform set be r = R + R I + R N , and obtain its set r set , transfer to S5; otherwise, transfer to S3; where, R Nl represents the short-term effectiveness evaluation function value corresponding to the state transition process; respectively represent the confidence levels of the state estimations of the multi-satellite platform collaboration for time-sensitive targets at times k and k + 1;

[0014] S5. For the elements r in r set sort them in ascending order according to the number of time intervals experienced, and successively select r from r If r ≤ reward1, then reward1 = r, set ​Determine that the multi-satellite platform collaboratively detects a time-sensitive target at time k + j, and end the iterative process; otherwise, go to S2; where, Indicates that the state of the multi-satellite platform set at time k + 1 enables the shortest time for collaborative detection of the time-sensitive target at time k + j; * represents the optimal state.

[0015] Preferably, in S3, a one-step greedy algorithm is used to obtain

[0016]

[0017] where, the specific solution process of v is as follows:

[0018] S10. Initialize the confidence of the state prediction of the time-sensitive target at time k + 1

[0019] S20. Initialize reward2 = 0, and store all possible action strategies of the satellite platform set Q at time k into the action set list; where, reward2 is a variable used to record the maximum value of the probability that the multi-satellite platform set Q collaboratively detects the time-sensitive target;

[0020] S30. Obtain the confidence in the prediction stage according to the state transition probability P(γ k+1 |γ k ) where, γ k and γ k+1 represent the position states of the time-sensitive target at times k and k + 1 respectively;

[0021] S40. If i ≥ L list , end the iteration; otherwise, go to S50; where, L list represents the length of the action set list;

[0022] S50. When the multi-satellite platform set Q is in the state at time k and takes the action then the state of the satellite platform at time k + 1 is The probability that the multi-satellite platform set Q collaboratively detects the time-sensitive target is where, represents the probability that the sensor does not detect the time-sensitive target at time k + 1; d is the differential symbol;

[0023] S60. If v > reward2, let go to S70; otherwise, directly go to S70; where, represents the action strategy that maximizes the probability of discovering the time-sensitive target in the current iterative operation, Denotes the state of the multi-satellite platform set Q that maximizes the probability of discovering time-sensitive targets in the current iteration operation at time k + 1;

[0024] S70, if i < L list , let i = i + 1, go to S50; otherwise, go to S80;

[0025] S80, the confidence level of the time-sensitive target state in the update phase at time k + 1 Go to S40.

[0026] Preferably, the confidence level in the prediction phase The specific solution process is as follows:

[0027] Using the posterior probability of the time-sensitive target state obtained at time k Predict the state of the time-sensitive target at time k + 1 to obtain the prior probability Calculated according to the Chapman-Kolmogorov formula:

[0028]

[0029]

[0030] Among them, γ k+1|k Represents the predicted state of the time-sensitive target at time k + 1; Respectively represent the continuous observation of the time-sensitive target by the satellite set cooperation from the initial time to time k, and the state set from the initial time to time k; Define the probability of the time-sensitive target position at the initial time

[0031] The confidence level in the update phase The specific solution process is as follows:

[0032] Continuously observe the time-sensitive target from the initial time until time k + 1, based on the prior probability of the time-sensitive target state at the previous time Calculate the posterior probability of observing the time-sensitive target state at the current time Use the recursive Bayesian rule to calculate the posterior probability:

[0033]

[0034] Among them, γ k+1|k+1 Represents the updated state of the time-sensitive target at time k + 1; p(z k+1 |γ k+1 , z 1:k , s 1:k+1 ) is the observation likelihood function of the satellite platform at time k + 1, (zk+1 |z 1:k ,s 1:k+1 ) is the normalized likelihood function, which is calculated by using the marginalization method.

[0035] Preferably, the short-term effectiveness evaluation function R in S3 N The specific solution process is as follows:

[0036] Use the local expected detection time LET to replace the global expected detection time:

[0037]

[0038] The probability of not detecting a time-sensitive target event within the time interval (k + 1, k + j) is:

[0039]

[0040] Among them, is the confidence level of the multi-satellite platform in switching the cooperative perception of the time-sensitive target position from the (k + j - 1)th moment to the (k + j)th moment, and the local expected detection time is

[0041] Since the short-term effectiveness evaluation function R N is only related to the time interval N and the behavior strategies formulated at each time node, therefore:

[0042]

[0043] A time-sensitive target observation system based on multi-satellite platform cooperative perception, comprising:

[0044] An initialization module, used to execute S1. It is known that the state of the multi-satellite platform at the kth moment is Store the possible states of the multi-satellite platform at the (k + 1)th moment into the queue Initialize reward1 = 0; among them, given the multi-satellite platform set Q = {S1, S2,... S q}; num represents the set of possible states of the multi-satellite platform at the (k + 1)th moment, represents the numth possible state of satellite S i at the (k + 1)th moment, k = 1, 2,..., N; reward1 represents the threshold for the action strategy to detect a time-sensitive target;

[0045] A selection module, used to execute S2. Let j = 1. If the queue queue is empty, the iteration ends; otherwise, select the first state variable in the queue queue as Obtain the state transition of the multi-satellite platform set from the kth moment to the (k + 1)th moment as Then remove the element from the queue queue and transfer it to the generation module to execute S3;

[0046] The generation module is used to execute S3, set j = j + 1, initialize an empty list adjnodes, generate all possible state transition sets of the multi-satellite platform from time k to time k + j, and put all elements into the list adjnodes. Transfer to the evaluation module to execute S4; where num1 is the number of its elements. It represents the num1-th possibility of state transition from time k to time k + j.

[0047] The evaluation module is used to execute S4. If j ≥ N or adjnodes is empty, calculate the R of all possible cases within the time interval j according to the short-term effectiveness evaluation function Calculate the R of all possible cases within the time interval j according to the short-term effectiveness evaluation function N Set And adopt a one-step greedy algorithm to obtain the instantaneous effectiveness evaluation function Set the confidence cumulative return expectation function value r of the collaborative action strategy of the multi-satellite platform set to r = R I +R N , and obtain its set r set , transfer to the determination module to execute S5; otherwise, transfer to the generation module to execute S3; where R NI Represents the short-term effectiveness evaluation function value corresponding to The state transition process. Respectively represent the confidence levels of the state estimations of the multi-satellite platform in collaborating on time-sensitive targets at times k and k + 1.

[0048] The determination module is used to execute S5. For the element r in r set Arrange the elements in r in ascending order according to the number of time intervals experienced, and sequentially select r from r , if r ≤ reward1, then set Determine that the multi-satellite platform collaboratively detects the time-sensitive target at time k + j and ends the iterative process; otherwise, transfer to the selection module to execute S2; where Represents that the state of the multi-satellite platform set at time k + 1 makes the time for collaboratively detecting the time-sensitive target at time k + j the shortest; * represents the optimal state. Represents that the state of the multi-satellite platform set at time k + 1 makes the time for collaboratively detecting the time-sensitive target at time k + j the shortest; * represents the optimal state.

[0049] A storage medium stores a computer program for observing time-sensitive targets based on the collaborative perception of multi-satellite platforms. Among them, the computer program enables a computer to execute the method for observing time-sensitive targets based on the collaborative perception of multi-satellite platforms as described above.

[0050] An electronic device includes:

[0051] One or more processors;

[0052] A memory; and

[0053] One or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the programs include a method for observing time-sensitive targets based on multi-satellite platform collaborative perception as described above.

[0054] (III) Beneficial effects

[0055] The present invention provides a method, system, storage medium, and electronic device for observing time-sensitive targets based on multi-satellite platform collaborative perception. Compared with the prior art, the following beneficial effects are achieved:

[0056] The present invention regards the detection process of time-sensitive targets as a partially observable Markov decision process, predicts the state position of time-sensitive targets in a deterministic manner through Bayesian inference theory, and solves the satellite platform observation strategy with a step size of 1 through a greedy algorithm, selecting an action strategy to ensure the maximum probability of discovering time-sensitive targets; through the search of a rolling horizon period observation strategy with a step size of N, the optimal action strategy of the satellite platform is selected within the rolling horizon period N, and the time-sensitive targets are detected in the shortest time. Description of the drawings

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0058] Figure 1 It is a schematic diagram of the confidence cumulative return expectation function of a multi-satellite platform collaborative action strategy provided by an embodiment of the present invention;

[0059] Figure 2 It is a schematic flowchart of a method for observing time-sensitive targets based on multi-satellite platform collaborative perception provided by an embodiment of the present invention. Detailed implementation manners

[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present invention belong to the scope of protection of the present invention.

[0061] Embodiments of the present application provide a time-sensitive target observation method, system, storage medium, and electronic device based on multi-satellite platform collaborative perception, which solve the technical problem of being unable to quickly locate time-sensitive targets.

[0062] The technical solutions in the embodiments of the present application to solve the above technical problems have the following general idea:

[0063] Define T as the time variable when a time-sensitive target is detected, and it is a random variable. Let k be the discrete time series, k = 1, 2,..., N. The objective function for multi-satellite platform collaborative observation of time-sensitive targets is: Since only partial position information of the time-sensitive target can be known during the search process of the time-sensitive target, the time-sensitive target search problem is also a partially observable Markov decision process POMDP (Partially Observable Markov Decision Process). The essence of the shortest time problem for time-sensitive target detection is to give the optimal action strategy for the satellite platform to complete the positioning of the time-sensitive target in the shortest time.

[0064] Define the action strategy The joint detection probability of multi-satellite platform collaborative perception of time-sensitive targets is:

[0065] Since the time T to detect the time-sensitive target is a random variable, and the confidence of the time-sensitive target state only depends on the choice of action decision. Based on the influence of the current action decision on the future confidence of the time-sensitive target The effectiveness evaluation of the influence can be composed of 3 different evaluation functions: the instantaneous effectiveness evaluation function R I , the short-term effectiveness evaluation function R N , the future or long-term effectiveness evaluation function R H , as Figure 1 shown.

[0066] Therefore, the rolling horizon control strategy is adopted to solve the optimal action strategy. Let N be the rolling horizon control period, then the cumulative return expectation function of the confidence of the multi-satellite platform collaborative action strategy is expressed as: The action strategy can be expressed as The above formula shows that the multi-satellite platform collaborative behavior strategy only relates to the change of the system state and the prior probability of the time-sensitive target.

[0067] Among them, represents the confidence of the multi-satellite platform in switching from the k-th moment to the (k + N)-th moment for collaborative perception of the position of the time-sensitive target. R I represents the expected time to detect the time-sensitive target at the k-th moment. R Ndenotes the expected time to detect a time-sensitive target from time k+1 to time k+N, R H denotes the expected time to detect a time-sensitive target after time k+N+1. It can be seen that the instantaneous effectiveness evaluation function R I is independent of time and only related to the current satellite platform state and the confidence level of the time-sensitive target state; the short-term effectiveness evaluation function R N is related to the time interval N and the behavior strategy formulated for N time nodes; the long-term effectiveness evaluation function R H denotes the expected impact on the future. As the confidence level value of the satellite platform for the time-sensitive target state improves and increases, in the embodiments of the present invention, R H can be set to zero. Therefore, R I The function can be solved using the greedy algorithm, and R N The function is suitable for solving using the rolling time domain strategy.

[0068] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific implementation manners.

[0069] Embodiment:

[0070] As Figure 2 shown, the embodiments of the present invention provide a method for observing time-sensitive targets based on multi-satellite platform collaborative perception, including:

[0071] S1. Given that the state of the multi-satellite platform at time k is Store the possible states of the multi-satellite platform at time k+1 into the queue Initialize reward1 = 0; where, given a multi-satellite platform set Q = {S1, S2,... S q}; num represents the set of possible states of the multi-satellite platform at time k+1, represents the satellite S i The num-th possible state at time k+1, k = 1, 2,..., N; reward1 represents the threshold for the action strategy to detect a time-sensitive target;

[0072] S2. Let j = 1. If the queue queue is empty, the iteration ends; otherwise, select the first state variable in the queue queue as Obtain the state transition of the multi-satellite platform set from time k to time k+1 as Then remove this element from the queue queue and transfer to S3;

[0073] S3. Let \(j = j + 1\), initialize an empty list \(adjnodes\), generate all possible state transition sets of the multi - satellite platform from time \(k\) to time \(k + j\), and put all elements into the list \(adjnodes\), then go to S4; where \(num1\) is the number of its elements, representing the \(num1\) - th possibility of the state transition from time \(k\) to time \(k + j\);

[0074] S4. If \(j\geq N\) or \(adjnodes\) is empty, calculate \(R\) for all possible cases within the time interval \(j\) according to the short - term effectiveness evaluation function and use a one - step greedy algorithm to obtain the instantaneous effectiveness evaluation function N set Let the confidence cumulative return expectation function value \(r\) of the collaborative action strategy of the multi - satellite platform set be \(r = R\) + \(R\) I + \(R\) N , to obtain its set \(r\) set , then go to S5; otherwise, go to S3; where \(R\) Nl represents the short - term effectiveness evaluation function value corresponding to the state transition process; respectively represent the confidence levels of the state estimations of the multi - satellite platform in collaborating to time - sensitive targets at times \(k\) and \(k + 1\);

[0075] S5. For the element \(r\) in \(r\) set sort the elements in \(r\) in ascending order according to the number of time intervals experienced, and sequentially select \(r\) from \(r\) . If \(r\leq reward1\), then \(reward1 = r\), set determine that the multi - satellite platform collaboratively detects the time - sensitive target at time \(k + j\), and end the iterative process; otherwise, go to S2; where, represents that the state of the multi - satellite platform set at time \(k + 1\) enables the shortest time to collaboratively detect the time - sensitive target at time \(k + j\); * represents the optimal state.

[0076] In the embodiment of the present invention, the detection process of the time - sensitive target is regarded as a partially observable Markov decision process. The state position of the time - sensitive target is predicted in a deterministic manner through the Bayesian inference theory, and the satellite platform observation strategy with a step size of 1 is solved by a greedy algorithm. The action strategy is selected to ensure the maximum probability of discovering the time - sensitive target; through the search of a rolling - horizon periodic observation strategy with a step size of \(N\), the optimal action strategy of the satellite platform is selected within the rolling - horizon period \(N\) to achieve detecting the time - sensitive target in the shortest time.

[0077] Next, each step of the above - mentioned scheme will be introduced in detail:

[0078] First, it should be noted that generally, the partially observable Markov decision process is used to describe how an Agent moves in an uncertain environment. Formally, it can be represented by a six-tuple {S, A, T, R, Ω, O}. S is the Agent's model of the environment it is in, usually referring to a finite set of states; A is the set of actions that the Agent can take under certain conditions; T is a set of conditional transition probabilities between states; R is the reward function; Ω: the set of observable information of the Agent; O: the observation function of the Agent, which can be used to calculate the possible observed values when entering the next state after taking action a.

[0079] And the set of satellite platforms Q = {S1, S2,... S q} given in the embodiments of the present invention, denotes the state of S i at time k (k = 1, 2,..., N). The state of the satellite set Q at time k is being S i in the state the action taken under, the action taken by the satellite platform set Q in the state under. denotes the set of possible action decisions of the i-th satellite platform S i in the state under, is the set of possible action decisions of the satellite platform set Q in the state under.

[0080] The state of multiple satellite platforms at time k + 1 is jointly determined by the state at time k and the action , that is

[0081] denotes the observation result of S i on the time-sensitive target in the state under, is the state of S i at time k (k = 1, 2,..., N) to the state at time k + N the decision sequence made. The position state of the time-sensitive target at time k is γ k , and the state transition probability of the time-sensitive target is P(γ k |γ k-1 ). Given the probability of the position of the time-sensitive target at the initial time The sensor detecting the target event is denoted as D, and the event of not detecting is denoted as The probability that the sensor detects the target at time k is expressed as The confidence of the collaborative state estimation of time-sensitive targets by multiple satellite platforms at time k is where are respectively the continuous observation of the time-sensitive target by the satellite set collaboration from the initial time to time k, and the state set from the initial time to time k, γ k|k-1 and γ k|k represent the predicted state and the updated state of the time-sensitive target at time k respectively.

[0082] The initial state confidence of the time-sensitive target at time k is The prior probability of the time-sensitive target Also known as the confidence in the prediction stage, at time k the Agent is in the state to make an observation to obtain the posterior probability of the state of the time-sensitive target The posterior probability is also known as the confidence in the update stage.

[0083] The prediction stage means using the posterior probability of the state of the time-sensitive target obtained at time k - 1 to predict the state of the time-sensitive target at time k, that is, to obtain the prior probability Calculated according to the Chapman-Kolmogorov formula:[[]]

[0084]

[0085] The update stage means continuously observing the time-sensitive target from the initial time until time k, and based on the prior probability of the state of the time-sensitive target at the previous time to calculate the posterior probability of observing the state of the time-sensitive target at the current time The posterior probability can be calculated using the recursive Bayesian rule:[[]]

[0086]

[0087] where p(z k |γ k , z 1:k-1 , s 1:k ) is the observation likelihood function of the satellite platform at time k, and p(z k |z 1:k-1 , s 1:k ) is the normalized likelihood function depending on the known information, usually calculated using the marginalization method. Let

[0088] Consider the satellite observation strategy from time k to k+N. As shown in steps S1 to S4, perform a search for the satellite platform observation strategy within a complete rolling time domain period N, and design a search termination judgment function. This algorithm is also known as the Limited First Depth Search (LDFS) algorithm, which realizes the selection of the optimal behavior strategy of the satellite within the rolling time domain period N and detects time-sensitive targets in the shortest time.

[0089] In step S1, the state of multiple satellite platforms at time k is known as Store the possible states of the multiple satellite platforms at time k+1 into the queue Initialize reward l = 0; where, given a set of multiple satellite platforms Q = {S1, S2,... S q}; num represents the set of possible states of the multiple satellite platforms at time k+1, represents the num-th possible state of satellite S i at time k+1, k = 1, 2,..., N; reward1 represents the threshold for the action strategy to detect a time-sensitive target, which is specifically used to determine whether the cumulative return expectation function value of the collaborative action strategy of the multiple satellite platform set reaches the threshold. If so, the time-sensitive target can be detected according to this action strategy.

[0090] In step S2, let j = 1. If the queue queue is empty, the iteration ends; otherwise, select the first state variable in the queue queue as Obtain the state transition of the multiple satellite platform set from time k to time k+1 as Then remove this element from the queue queue and transfer to S3.

[0091] In step S3, let j = j+1, initialize an empty list adjnodes, generate the set of all possible state transitions of the multiple satellite platforms from time k to time k+j, and put all elements into the list adjnodes, Transfer to S4; where num1 is the number of its elements, represents the num1-th possible state transition from time k to time k+j.

[0092] In step S4, if j≥N or adjnodes is empty, according to the short-term effectiveness evaluation function Calculate R for all possible cases within the time interval j N Set And use the one-step greedy algorithm to obtain the instantaneous effectiveness evaluation function Let the cumulative return expectation function value r of the collaborative action strategy of the multiple satellite platform set be r = RI +R N , obtain its set r set , transfer to S5; otherwise, transfer to S3; where R Nl represents the short-term effectiveness evaluation function value corresponding to the state transition process; respectively represent the confidence levels of the state estimations of the multi-satellite platform collaborative targeting of time-sensitive targets at times k and k + 1.

[0093] In this step, a one-step greedy algorithm OSSGA is designed to solve the problem of multi-satellite platform collaborative search for time-sensitive targets on the sea surface. The i-th satellite platform S i selects an action strategy based on the current state and the confidence level of the time-sensitive target state to ensure that the probability of discovering a time-sensitive target by this action strategy is the largest. Specifically, it is obtained by using a one-step greedy algorithm

[0094]

[0095] where the specific solution process of v is as follows:

[0096] S10. Initialize the confidence level of the state prediction of the time-sensitive target at time k + 1

[0097] S20. Initialize reward2 = 0, and store all possible action strategies of the satellite platform set Q at time k into the action set list; where reward2 is a variable used to record the maximum value of the probability that the multi-satellite platform set Q collaboratively detects a time-sensitive target; thus, the action strategy with the largest probability of discovering a time-sensitive target can be found during the algorithm iteration process.

[0098] S30. Obtain the confidence level in the prediction stage according to the state transition probability P(γ k+1 |γ k ) where γ k and γ k+1 respectively represent the position states of the time-sensitive target at times k and k + 1;

[0099] where the confidence level in the prediction stage The specific solution process is as follows:

[0100] Use the posterior probability of the time-sensitive target state obtained at time k to predict the state of the time-sensitive target at time k + 1, that is, obtain the prior probability Calculated according to the Chapman-Kolmogorov formula:​

[0101]

[0102] Among them, γ k+1|k represents the predicted state of the time-sensitive target at the (k + 1)-th moment; respectively represent that the satellite set collaboratively observes the time-sensitive target continuously from the initial moment to the k-th moment, and the state set from the initial moment to the k-th moment; define the probability of the position of the time-sensitive target at the initial moment

[0103] S40. If i ≥ L list , the iteration ends; otherwise, go to S50; where L list represents the length of the action set list.

[0104] S50. When the multi-satellite platform set Q is in the state at the k-th moment and takes the action then the state of the satellite platform at the (k + 1)-th moment is The probability that the multi-satellite platform set Q collaboratively detects the time-sensitive target is Among them, represents the probability that the sensor does not detect the time-sensitive target at the (k + 1)-th moment; d is the differential symbol.

[0105] S60. If v > reward2, let reward2 = v, go to S70; otherwise, directly go to S70; where represents the action strategy that maximizes the probability of detecting the time-sensitive target in the current iteration operation, represents the state of the multi-satellite platform set Q at the (k + 1)-th moment that maximizes the probability of detecting the time-sensitive target in the current iteration operation.

[0106] S70. If i < L list , let i = i + 1, go to S50; otherwise, go to S80.

[0107] S80. The confidence level of the state of the time-sensitive target at the (k + 1)-th moment in the update stage Go to S40;

[0108] Among them, the confidence level in the update stage The specific solution process is as follows:

[0109] Continuously observe the time-sensitive target from the initial moment until the (k + 1)-th moment, and based on the prior probability of the state of the time-sensitive target at the previous moment calculate the posterior probability of observing the state of the time-sensitive target at the current moment Calculate the posterior probability using the recursive Bayesian rule:

[0110]

[0111] Among them, γ k+1|k+1 represents the updated state of the time-sensitive target at time k + 1; p(z k+1| γ k+1 , z 1:k , s 1:k+1 ) is the observation likelihood function of the satellite platform at time k + 1, and (z k+1 |z 1:k , s 1:k+1 ) is the normalized likelihood function, which is calculated by the marginalization method.

[0112] In addition, in step S4, the specific solution process of the short-term effectiveness evaluation function R N is as follows:

[0113] Use the local expected detection time LET to replace the global expected detection time:

[0114]

[0115] The probability of not detecting the time-sensitive target event within the time interval (k + 1, k + j) is:

[0116]

[0117] Among them, is the confidence level of the multi-satellite platform to switch from time k + j - 1 to time k + j for collaborative sensing of the time-sensitive target position, and the local expected detection time is

[0118] Since the short-term effectiveness evaluation function R N is only related to the time interval N and the behavior strategies formulated at each time node, therefore:

[0119]

[0120] In step S5, for the elements r in r set sort them in ascending order according to the number of time intervals experienced, and select r from r in turn. If r ≤ reward1, then reward1 = r, set determine that the multi-satellite platform collaboratively detects the time-sensitive target at time k + j, and end the iteration process; otherwise, transfer to S2; among them, represents that the state of the multi-satellite platform set at time k + 1 makes the collaborative detection of the time-sensitive target at time k + j the shortest; * represents the optimal state.

[0121] ​An embodiment of the present invention provides a time-sensitive target observation system based on multi-satellite platform collaborative perception, including:

[0122] An initialization module, configured to execute S1. It is known that the state of the multi-satellite platform at time k is Store the possible states of the multi-satellite platform at time k + 1 into a queue Initialize reward1 = 0; where, given a multi-satellite platform set Q = {S1, S2,... S q}; num represents the set of possible states of the multi-satellite platform at time k + 1, represents the num-th possible state of satellite S i at time k + 1, k = 1, 2,..., N; reward1 represents the threshold for the action strategy to detect a time-sensitive target;

[0123] A selection module, configured to execute S2. Let j = 1. If the queue queue is empty, the iteration ends; otherwise, select the first state variable in the queue queue as Obtain the state transition of the multi-satellite platform set from time k to time k + 1 as Then remove this element from the queue queue and transfer to the generation module to execute S3;

[0124] A generation module, configured to execute S3. Let j = j + 1, initialize an empty list adjnodes, generate all possible state transition sets of the multi-satellite platform from time k to time k + j, and put all elements into the list adjnodes, Transfer to the evaluation module to execute S4; where num1 is the number of its elements, represents the numl-th possible state transition from time k to time k + j;

[0125] An evaluation module, configured to execute S4. If j ≥ N or adjnodes is empty, calculate R for all possible cases within the time interval j according to the short-term effectiveness evaluation function N set And use a one-step greedy algorithm to obtain the instantaneous effectiveness evaluation function Let the confidence cumulative return expectation function value r of the collaborative action strategy of the multi-satellite platform set be r = R I +R N , and obtain its set r set , transfer to the determination module to execute S5; otherwise, transfer to the generation module to execute S3; where, R Nl represents the short-term effectiveness evaluation function value corresponding to the state transition process; respectively represent the confidence levels of the state estimations of the multi-satellite platform's collaborative targeting of time-sensitive targets at times k and k + 1;

[0126] a determination module, configured to execute S5. For r set in the elements r of sort them in ascending order according to the number of time intervals experienced, and sequentially select r from r set If r ≤ reward1, then determine that the multi-satellite platform collaboratively detects the time-sensitive target at time k + j, and end the iterative process; otherwise, transfer to the selection module to execute S2; where represents that the state of the multi-satellite platform set at time k + 1 enables the shortest time for collaborative detection of the time-sensitive target at time k + j; * represents the optimal state.

[0127] An embodiment of the present invention provides a storage medium storing a computer program for time-sensitive target observation based on multi-satellite platform collaborative sensing, wherein the computer program causes a computer to execute the time-sensitive target observation method based on multi-satellite platform collaborative sensing as described above.

[0128] An embodiment of the present invention provides an electronic device, including:

[0129] one or more processors;

[0130] a memory; and

[0131] one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include those for executing the time-sensitive target observation method based on multi-satellite platform collaborative sensing as described above.

[0132] In summary, compared with the prior art, the following beneficial effects are achieved:

[0133] In the embodiment of the present invention, the detection process of time-sensitive targets is regarded as a partially observable Markov decision process. The state position of time-sensitive targets is predicted in a deterministic manner through Bayesian inference theory, and a satellite platform observation strategy with a step size of 1 is solved through a greedy algorithm. An action strategy is selected to ensure the maximum probability of discovering time-sensitive targets; through the search of a rolling horizon period observation strategy with a step size of N, an optimal action strategy of the satellite platform is selected within the rolling horizon period N, so as to detect time-sensitive targets in the shortest time.

[0134] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0135] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A time-sensitive target observation method based on collaborative perception of multiple satellite platforms, characterized in that, including: S1. The state of the multi-satellite platform at time k is known as Store the possible states of the multi-satellite platform at time k+1 into the queue Initialize reward1 = 0; where, given the multi-satellite platform set Q = {S1, S2,... S q}; num represents the set of possible states of the multi-satellite platform at time k+1, represents the satellite S i the num-th possible state at time k+1, k = 1, 2,..., N; reward1 represents the threshold for the action strategy to detect time-sensitive targets; S2. Let \(j = 1\). If the queue `queue` is empty, the iteration ends; otherwise, select the first state variable in the queue `queue` as Obtain the state transition of the multi-satellite platform set from time \(k\) to time \(k + 1\) as Then remove this element from the queue `queue` and transfer to S3; S3. Let \(j = j + 1\), initialize an empty list \(adjnodes\), generate all possible state transition sets of the multi - satellite platform from time \(k\) to time \(k + j\), and put all elements into the list \(adjnodes\). Go to S4; where \(num1\) is the number of its elements. It represents the \(num1\) - th possibility of state transition from time \(k\) to time \(k + j\). S4. If j ≥ N or adjnodes is empty, calculate the R for all possible cases within the time interval j according to the short-term performance evaluation function and obtain the instantaneous performance evaluation function by using the one-step greedy algorithm N Set Let the cumulative return expectation function value r of the cooperative action strategy of the multi-satellite platform set be r = R +R I +R N to obtain its set r set , and then go to S5; otherwise, go to S3; where R Nl represents the short-term performance evaluation function value corresponding to the state transition process; respectively represent the confidence levels of the state estimations of the multi-satellite platform cooperation for time-sensitive targets at times k and k + 1; S5. For r set Arrange the elements r in ascending order according to the number of time intervals experienced, and select r from r set in turn. If r ≤ reward1, then reward1 = r. Determine that the multi-satellite platform collaboratively detects the time-sensitive target at time k + j, and end the iteration process; otherwise, go to S2; where represents that the state of the multi-satellite platform set at time k + 1 enables the shortest time for collaborative detection of the time-sensitive target at time k + j; * represents the optimal state.

2. The time-sensitive target observation method based on multi-satellite platform collaborative perception according to claim 1, wherein in S3, the one-step greedy algorithm is adopted to obtain wherein, the specific solution process of v is as follows: S10. Initialize the confidence of the state prediction of the time-sensitive target at the (k + 1)-th moment i = 1; S20. Initialize reward2 = 0, and store all possible action strategies of the satellite platform set Q at the k-th moment into the action set list; where reward2 is a variable used to record the maximum value of the probability that the multi-satellite platform set Q collaboratively detects time-sensitive targets. ​ S30. Obtain the confidence level in the prediction stage according to the state transition probability P(γ k+1 |γ k ) where γ k and γ k+1 respectively represent the position states of the time-sensitive target at times k and k + 1; S40. If i ≥ L list , the iteration ends; otherwise, go to S50; where L list represents the length of the action set list; The state of the multi-satellite platform set Q at time k Take action at Then the state of the satellite platform at time k+1 is The probability that the multi-satellite platform set Q collaboratively detects a time-sensitive target is Among them, Indicates the probability that the sensor does not detect a time-sensitive target at time k+1; d is the differential symbol; S60. If v > reward2, then let reward2 = v, Go to S70; otherwise, directly go to S70; where represents the action strategy that maximizes the probability of discovering time-sensitive targets in the current iterative operation, represents the state of the multi-satellite platform set Q that maximizes the probability of discovering time-sensitive targets at time k + 1 in the current iterative operation; S70. If i < L list , set i = i + 1 and go to S50; otherwise, go to S80; Confidence of the state of the time-sensitive target at the (k + 1)-th moment in the update phase Transfer to S40.

3. The time-sensitive target observation method based on multi-satellite platform collaborative perception according to claim 2, characterized in that Confidence in the prediction phase And the body solution process is as follows: The posterior probability of the time-sensitive target state obtained at time k Predict the state of the time-sensitive target at time k+1 to obtain the prior probability Calculated according to the Chapman-Kolmogorov formula: where γ k+1|k represents the predicted state of the time-sensitive target at time k + 1; respectively represent the continuous observation of the time-sensitive target by the satellite set collaboration from the initial time to time k, and the state set from the initial time to time k; define the probability of the position of the time-sensitive target at the initial time Confidence in the update stage The specific solution process is as follows: Continuously observe the time-sensitive target from the initial moment until the (k + 1)-th moment, based on the prior probability of the state of the time-sensitive target at the previous moment Calculate the posterior probability of observing the state of the time-sensitive target at the current moment Calculate the posterior probability using the recursive Bayesian rule: Among them, γ k+1|k+1 represents the updated state of the time-sensitive target at time k + 1; p(z k+1 |Y k+1 , z 1:k , s 1:k+1 ) is the observation likelihood function of the satellite platform at time k + 1, and (z k+1 |z 1:k , s 1:k+1 ) is the normalized likelihood function, which is calculated by the marginalization method.

4. The time-sensitive target observation method based on multi-satellite platform collaborative perception according to any one of claims 1 to 3, characterized in that, The short-term performance evaluation function R in S3 N The specific solution process is as follows: the local expected detection time LET is used to replace the global expected detection time: the probability of not detecting the time-sensitive target event in the time interval (k + 1, k + j) is: Among them, is the confidence level of the multi-satellite platform for collaborative perception of the time-sensitive target position when transitioning from time k + j - 1 to time k + j, and the local expected detection time is Since the short-term effectiveness evaluation function R N is only related to the time interval N and the behavior strategies formulated at each time node, therefore:

5. A time-sensitive target observation system based on multi-satellite platform collaborative perception, characterized in that, including: Initialization module, used to execute S1. The state of the known multi-satellite platform at time k is Store the possible states of the multi-satellite platform at time k+1 into the queue Initialize reward1 = 0; where the given multi-satellite platform set Q = {S1, S2,... S q}; num represents the set of possible states of the multi-satellite platform at time k+1, represents the satellite S i the num-th possible state at time k+1, k = 1, 2,..., N; reward1 represents the threshold for the action strategy to detect time-sensitive targets; A selection module, which is used to execute S2: set j = 1. If the queue "queue" is empty, the iteration ends; otherwise, select the first state variable in the queue "queue" as obtain the state transition of the multi-satellite platform set from time k to time k + 1 as then remove this element from the queue "queue" and transfer to the generation module to execute S3; A generation module, which is used to execute S3, set j = j + 1, initialize an empty list adjnodes, generate a set of all possible state transitions of the multi-satellite platform from time k to time k + j, and put all elements into the list adjnodes, transfer to the evaluation module to execute S4; where num1 is the number of its elements, indicating the num1-th possibility of state transition from time k to time k + j; An evaluation module, configured to execute S4. If j≥N or the adjacency nodes are empty, according to the short-term performance evaluation function calculate the R of all possible cases within the time interval j N set and use a one-step greedy algorithm to obtain the instantaneous performance evaluation function Let the confidence cumulative return expectation function value r of the multi-satellite platform set collaborative action strategy be r = R I +R N , and obtain its set r set , transfer to the determination module to execute S5; otherwise, transfer to the generation module to execute S3; where R Nl represents the short-term performance evaluation function value corresponding to the state transition process; respectively represent the confidence levels of the state estimations of the multi-satellite platform collaboration for time-sensitive targets at times k and k+1; Determination module, for performing S5, for r set among the elements r in sort the number of time intervals experienced in ascending order, and select r from r in turn. If r ≤ reward1, then reward1 = r, set If r ≤ reward1, then reward1 = r, Determine that the multi-satellite platform collaboratively detects a time-sensitive target at the k + j moment, and end the iterative process; otherwise, transfer to the selection module to execute S2; where represents that the state of the multi-satellite platform set at the k + 1 moment enables the shortest time for collaborative detection of the time-sensitive target at the k + j moment; * represents the optimal state.

6. A storage medium, characterized in that, it stores a computer program for time-sensitive target observation based on multi-satellite platform collaborative perception, wherein the computer program enables the computer to execute the time-sensitive target observation method based on multi-satellite platform collaborative perception according to any one of claims 1 to 4.

7. An electronic device, characterized in that, including: one or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include those for executing the time-sensitive target observation method based on multi-satellite platform collaborative perception according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-satellite online cooperative task planning method for time-sensitive moving target tracking

    CN112862306A

  • Multi-satellite emergency task scheduling method and system considering minimum disturbance and maximum benefit

    CN114781784A