A distributed satellite dynamic grid load allocation method based on potential energy game

By combining grid modeling and potential energy game theory with the SeTVBRP algorithm that selects time-varying parameters, the problem of lack of mathematical models for dynamic grid load allocation in distributed satellite systems is solved, and efficient task allocation and collaboration under decentralized decision-making are achieved.

CN120124462BActive Publication Date: 2026-01-02NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510197251.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2026-01-02
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

In distributed satellite systems, the lack of a standardized mathematical model for dynamic grid load allocation makes it difficult to achieve scientific and intuitive task allocation. Furthermore, centralized methods are not suitable for distributed environments, and there is a lack of high-quality solutions for decentralized decision-making.

Method used

A grid model is used for unified description, and the objective function is decomposed by introducing potential energy game theory. The better response process (SeTVBRP) algorithm with time-varying parameters is designed and used for distributed satellite dynamic grid load allocation. Reasonable decision-making is achieved in the absence of a central decision node through the potential energy game model and the SeTVBRP algorithm.

Benefits of technology

It effectively solves the challenges of dynamic grid load allocation in distributed satellite systems, improves the scientific nature and efficiency of task allocation, and achieves high-quality decision-making and collaborative efficiency in the absence of a central controller.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124462B_ABST
    Figure CN120124462B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of multi-star dynamic grid load distribution methods based on potential energy game, by satellite data and the grid data of the management end based on satellite position division are obtained in ground station, and carry out DGA problem modeling, ground station carries out sDGA problem modeling, establishes the potential energy game model of distributed satellite dynamic grid load distribution, to obtain the distributed satellite grid load distribution scheme based on the selection time-varying optimal response process, the scheme of the present application is more efficient and there is superiority relative to prior art in solving DGA problem process.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of task planning and scheduling, and particularly relates to a multi-satellite dynamic grid load distribution method based on potential game. BACKGROUND

[0002] A distributed satellite system refers to a plurality of satellite platforms operating on different space orbits, which are equipped with sensors to perform observation tasks. The so-called observation task refers to the imaging activity of the sensor carried by the satellite within a specific time period. With the rapid increase in observation task demand and the continuous progress of space technology, satellite technology has evolved from a single large spacecraft to a large-scale distributed satellite system. For example, Planet Company has successfully launched more than 150 Dove satellites, building a large-scale distributed satellite network. Although the distributed satellite system has obvious advantages, it also brings new challenges in control and collaborative operation. In a dynamic environment, real-time allocation of a large number of observation tasks to satellites is a complex task, which requires huge computing resources. Therefore, as shown in the figure, the current trend is to divide observation tasks into different grids according to geographic information for target allocation. Figure 1

[0003] The advantages of introducing a grid model to represent online imaging tasks are as follows: (1) In the grid model, the number of tasks within each grid is quantitatively represented by the observation load value of each grid. The arrival and disappearance of online tasks (whether completed or failed due to exceeding the deadline) can achieve dynamic balance on the grid observation load value, thereby reducing the dynamic nature of the star cluster collaboration problem to some extent and reducing the energy consumption caused by inter-satellite communication. (2) Due to the fixedness and regularity of the grid boundary, combined with the prediction information of the satellite orbit, the visibility calculation of the geographic grid can be completed in advance. Compared with real-time visibility calculation based on tasks, this method not only shortens the preprocessing time for handling a large number of online tasks, but also unifies the time points for visibility calculation by each satellite in the absence of a cluster head satellite, effectively avoiding confusion in the collaboration process. (3) In addition, using the geographic grid model to represent satellite imaging tasks can achieve unified identity coding of remote sensing targets, facilitating the effective transfer of remote sensing targets, collaboration schemes, and execution results between different autonomous satellites in the collaboration process, reducing the complexity of inter-satellite communication protocols, and saving inter-satellite communication link resources as much as possible; (4) At the same time, by monitoring the changes in grid load values, each satellite can macroscopically control the completion of remote sensing tasks at a minimum storage cost to cope with the loss of the global macroscopic perspective caused by the absence of a cluster head satellite. (5) Considering the possible imbalance in task allocation, the grid model takes the minimization of the maximum grid observation load value as the objective function, striving to achieve the most balanced allocation of tasks and improve resource utilization efficiency. ​

[0004] However, at present, in the actual management process, the distributed satellite dynamic grid load allocation is faced with the following problems: 1) Lack of standard and intuitive distributed satellite dynamic grid load mathematical model, which presents time-dependent characteristics of constraints and benefits, thereby guiding the design of the algorithm. Due to the limited observation time window, time-dependent observation quality and unpredictable changes in grid observation load, how to scientifically and intuitively describe the dynamic network and its time-dependent characteristics, standardize the mathematical expressions of decision variables, constraints and benefits, help understand the combinatorial optimization connotation of the dynamic grid load allocation problem, and then better guide the design of the algorithm and operator, is an important topic. 2) Due to the consideration of robustness, scalability and management overhead, the centralized method designed for a single large satellite is not suitable for solving problems involving distributed satellites in uncertain environments. The concept of distributed self-organization provides a comprehensive solution to the resource allocation problem between tasks in a satellite group in a distributed environment. Unlike centralized allocation mechanisms, distributed allocation mechanisms do not rely on a single master controller; instead, each agent achieves the global goal through mutual interaction and communication. 3) However, without a centralized controller, it is difficult to guarantee a high-quality solution theoretically. Therefore, designing an individual decision algorithm that is theoretically conducive to the global goal has become a core challenge of the distributed satellite dynamic grid load allocation method. SUMMARY

[0005] In view of the problems of the prior art, the present application proposes a distributed satellite dynamic grid load allocation method based on potential game for a dynamic grid composed of online tasks: ① A grid model is used to uniformly describe the online arriving satellite tasks, and a distributed satellite dynamic grid load allocation (DGA) mathematical modeling is established; ② In order to realize reasonable decision of each satellite without a central decision node, the potential game theory is introduced to decompose the DGA problem objective function to each satellite's local performance function, and an equivalent global performance function and potential function are designed; ③ In order to improve the online decision and collaboration efficiency of the satellite, a Selective Time Variant-parameter Better Reply Process (SeTVBRP) algorithm is proposed; ④ Algorithm analysis experiments, algorithm comparison experiments are carried out under different task scale and satellite group scale simulation scenes, and comparison experiments are carried out with five advanced potential game learning algorithms, which illustrates the efficiency and superiority of the algorithm in solving the DGA problem, and objectively verifies the performance of the algorithm.

[0006] The technical solutions adopted by the present application are as follows:

[0007] A multi-satellite dynamic grid load distribution method based on potential game, executes the following steps:

[0008] S1. The ground station obtains satellite data and grid data divided by the satellite position based management end, and models the DGA problem;

[0009] N={1, 2,..., n} represents a satellite index set, wherein the satellite s i has an index i, i∈N;

[0010] M={1, 2,..., m} represents a grid index set, wherein the grid r j has an index j, j∈M;

[0011] S2. The ground station divides the DGA problem into a single-stage DGA problem according to the time slicing method, and models the sDGA problem; wherein the single-stage DGA problem is the sDGA problem;

[0012] Each grid r j needs a satellite to complete the observation task in the grid to reduce its observation load β i ;

[0013] In the sDGA problem modeling, the observation ability of the satellite s i to the grid r j is represented by α ij , and the amount of time that the satellite s i allocates to the grid r j is represented by x ij ;

[0014] The remaining observation load y j of the grid r j is defined as follows:

[0015]

[0016] S3. Establish a potential game model of distributed satellite dynamic grid load distribution;

[0017] Let G P ={N, {A i}, U i} represent a game, wherein N={1, 2,..., n} is a satellite set; A=A1×A2...×A n is a joint action set, and U i : A i →R is a local utility function of participant i;

[0018] For the action set a=(a1, a2,..., a n )∈A, wherein; a i ∈Ai is an action of participant i,

[0019] S4. Obtain a distributed satellite grid load allocation scheme based on the selection time-varying optimal response process;

[0020] Generate an initial allocation scheme for all satellites according to the greedy strategy;

[0021] Randomly select a satellite update action at each iteration time to generate a new allocation scheme.

[0022] Further, in step S1,

[0023] At any time, each satellite can only perform an imaging task in a specific grid;

[0024] The observation load of the grid changes dynamically over time;

[0025] The change of the observation load is directly affected by the increase or decrease of the number of imaging tasks within its grid.

[0026] Further, in the modeling of the DGA problem, for each satellite s i , the corresponding grid r j The time window is w ij =[c ij ,d ij ],

[0027] Where c ij represents the start time of the time window, and d ij represents the end time of the time window.

[0028] Further, in step S1, in the modeling of the DGA problem, the state change point t k is the specific time when the time window changes, and the time interval [t k , t k +Δt] is called stage k.

[0029] Further, in the modeling of the DGA problem, is the observation capability of satellite s i to grid r j ; the values of observation capability and observation load are non-negative constants;

[0030] represents the visible grid set of satellite s i in decision stage k, that is:

[0031]

[0032] Further, in the modeling of the DGA problem, For the assignment grid set of satellite s i in decision stage k, and

[0033] If there is no intersection between the assignment grid sets of two adjacent decision stages, i.e. Satellite s i needs to adjust its imaging payload angle to steer from the grids in set to the grids in set .

[0034] Further, in the modeling of the DGA problem, the switching time η i of satellite s ik in decision stage k is:

[0035]

[0036] where H represents a non-negative constant to represent the simplified imaging payload switching time;

[0037] When the intersection of the assignment grid sets of satellite s i in k-1 decision stage and k decision stage is an empty set, the switching time of the two stages is H, otherwise it is 0.

[0038] Further, in step S2, if the satellite is assigned to two or more grids, the time for the satellite to perform imaging payload attitude switching between different grids is considered.

[0039] Further, in the modeling of the sDGA problem, represents the assignment grid set of satellite s i , satisfying For satellite s i , the switching time ρ i is calculated as follows:

[0040]

[0041] Further, in the modeling of the sDGA problem, in decision stage k, the problem model P0 is constructed as follows

[0042]

[0043] where x ij is the decision variable, representing the amount of time satellite s i is assigned to grid r j ;

[0044] Since the sDGA problem is a single-stage DGA problem divided from the DGA problem according to the time slicing method, satellite s i and grid rj The same as the physical meaning of model DGA expression;

[0045] Δt is the total allocatable time of this stage; the objective function U(x) is an exponential function of minimizing the maximum value of the residual observation load, where represents the observation load of the grid r j in all satellite total allocation schemes x.

[0046] represents the conversion time ρ i and the sum of the satellite total allocation time s i cannot exceed the allocatable time Δt. BRIEF DESCRIPTION OF DRAWINGS

[0047] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings required to be used in the specific embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0048] Figure 1 Grid model schematic diagram.

[0049] Figure 2 is a satellite and grid distribution diagram.

[0050] Figure 3 is an observation ability diagram of different satellites to different grids.

[0051] Figure 4 is a conversion time schematic diagram.

[0052] Figure 5 is a single-stage DGA problem objective function schematic diagram.

[0053] Figure 6 is a multi-stage dynamic allocation schematic diagram.

[0054] Figure 7 is an improved strategy ablation experiment result analysis under a small-scale scenario.

[0055] Figure 8 is the convergence performance of 500 iterations of 6 algorithms under a small-scale scenario.

[0056] Figure 9 is the convergence performance of 2000 iterations under a large-scale scenario. DETAILED DESCRIPTION

[0057] Exemplary embodiments of the present application are described herein with reference to the accompanying drawings, which are cited by way of example only. The various details of the embodiments of the application can be combined in other embodiments and can be used with other applications without departing from the scope of the application. Therefore, the application should not be construed as being limited to the embodiments described in the following description. Rather, the description is provided as an exemplification of the principles of the application.

[0058] In order to make personnel in the technical field better understand the application scheme, the technical solutions in the embodiments of the application will be described clearly and completely below in conjunction with the drawings in the embodiments of the application. Obviously, the described embodiments are only a part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work should belong to the protection scope of the application.

[0059] A distributed satellite dynamic grid load allocation method based on potential game is proposed for dynamic grid composed of online tasks: ①The grid model is used to uniformly describe the online arriving satellite tasks, and the mathematical modeling of distributed satellite dynamic grid load allocation (DGA) is established; ②In order to realize the reasonable decision of each satellite under the condition of no central decision node, the potential game theory is introduced to decompose the DGA problem objective function to the local performance function of each satellite, and the equivalent global performance function and potential function are designed; ③In order to improve the online decision and collaborative efficiency of satellites, a Selective Time Variant-parameter Better Reply Process (SeTVBRP) algorithm is proposed; ④The algorithm analysis experiment, algorithm comparison experiment are carried out under different task scale and constellation scale simulation scenes, and the comparison experiment is carried out with five advanced potential game learning algorithms, which illustrates the efficiency and superiority of the algorithm in solving the DGA problem, and the objective experimental verification is carried out on the algorithm performance.

[0060] The application relates to a dynamic grid with a time-window allocation (DGA) problem. Firstly, according to the characteristics of the satellite and the grid, the application defines the research subjects of the DGA problem, and the properties of the grid and the satellite are defined. Further, in order to solve the dynamic nature of the DGA problem, the DGA problem is decomposed into a single-stage dynamic grid with a time-window allocation (sDGA) problem. After analyzing the relationship between each stage of the DGA problem, the application constructs a mathematical model of the sDGA problem.

[0061] S1. The ground station acquires satellite data and grid data divided by the satellite position based on the management end, and performs DGA problem modeling

[0062] The DGA problem involves a group of satellites S={s i |i∈N} and a group of grids R={r j |j∈M}, wherein N={1, 2,..., n} and M={1, 2,..., m} represent the index set of the satellite and the grid respectively. At any time, each satellite can only perform an imaging task in a specific grid. The observation load of the grid changes dynamically over time. The change of the observation load is directly affected by the increase and decrease of the number of imaging tasks in the grid. As shown in the figure, the dots in the figure represent specific imaging tasks, and the color depth of the grid intuitively maps the observation load of the grid. Figure 2

[0063] For each satellite s i , the time window w j of the corresponding grid r ij is defined as [c ij , d ij ], wherein c ij represents the start time of the time window, and d ij represents the end time of the time window. In view of the certainty and regularity of the grid boundary, combined with the prediction information of the satellite orbit, the calculation of the visibility of the geographical grid can be completed in advance. The specific time when the time window changes is defined as the state change point t k , and the time interval [t k , t k +ΔT k ] is called stage k. According to the visible time window of the satellite and the grid, the entire time axis of the DGA can be divided into several decision stages, as shown in the figure. Figure 3

[0064] ​​Therefore, in order to reflect the dynamic characteristics of DGA at each stage, the DGA problem is divided into a series of consecutive single-stage DGAs (sDGAs) problems according to different decision stages. In each stage, different model parameters reflect the specific properties of DGA at this stage. For each stage k, the observation load of grid r j is denoted as Figure 4 As shown in the figure, a time window can be divided into multiple decision stages, and the observation ability of the satellite in each stage is different. For example, the decision stage close to the center of the time window has higher observation ability. At the same time, the observation ability of satellite s i to grid r j is defined as According to the definition of stage division of the DGA problem in the previous section, the observation ability and the observation load are both non-negative constants. Let denote the visible grid set of satellite s i in decision stage k, that is:

[0065]

[0066] Let denote the assigned grid set of satellite s i in decision stage k, and If there is no intersection between the assigned grid sets of two adjacent decision stages, that is satellite s i needs to spend a certain time to adjust its imaging payload angle in order to switch from the grid set to the grid set . Therefore, the transition time η i of satellite s ik in decision stage k is defined as:

[0067]

[0068] where H represents a non-negative constant, which is used to represent the simplified imaging payload transition time. As shown in the figure, when the intersection of the assigned grid sets of satellite s i in k-1 decision stage and k decision stage is an empty set, the transition time of the two stages is H, otherwise it is 0.

[0069] ​Based on the above analysis, this section divides the timeline of the DGA problem into several sub-problems sDGA to cope with the dynamic changes of the grid observation load and satellite observation capacity within the time window. The relationship between adjacent decision stages is mainly reflected in the imaging load conversion time between the allocation results, which makes the DGA problem into a sequence decision problem, that is, the sDGA allocation result of the next stage depends on the sDGA allocation result of the previous stage. Therefore, only by solving the single-stage DGA problem (sDGA problem) one by one in order, the whole DGA problem can be effectively solved.

[0070] S2. Ground station models sDGA problem

[0071] The present application will analyze and model the sDGA problem (also known as the single-stage DGA problem). In sDGA, R i represents the set of visible grids of satellite s i , S j represents the set of visible satellites of grid r j . Each grid r j needs a satellite to complete the observation task within the grid to reduce its observation load β i . The observation capacity of satellite s i to grid r j is represented by α ij , and the amount of time that satellite s i is allocated to grid r j is represented by x ij . The remaining observation load y j of grid r j is defined as follows:

[0072]

[0073] In order to achieve effective and balanced allocation of grid load, the optimization objective of the sDGA problem is set to minimize the maximum load remaining amount.

[0074] In order to further intuitively understand the meaning of the objective function, Figure 5 an example case containing 10 grids in the sDGA problem is shown, where Figure 5 (a) The height of the rectangle represents the observation load size of each grid, and Figure 5 (b) The height of the rectangle represents the remaining observation load of each grid after the satellite performs the imaging task, where the vertical coordinate value corresponding to the dashed line is the maximum load remaining amount maxy j .

[0075] If a satellite is allocated to two or more grids, the time for the satellite to perform imaging load attitude conversion between different grids needs to be considered. Define Indicates satellite s i The allocated grid set satisfies For satellites i Its conversion time ρ i The calculation formula is as follows:

[0076]

[0077] Based on the above analysis, in decision-making stage k, the problem model P0 is constructed as follows:

[0078]

[0079] st

[0080]

[0081] Where, x ij Let be the decision variable, representing the satellite s i Assigned to grid r j The time required. Since the sDGA problem is a single-stage DGA problem derived from the DGA problem using the time-slicing method, the satellite s i and grid r j This has the same physical meaning as expressed by the DGA model. Δt represents the total allocatable time for this stage. The objective function U(x) is an exponential function that minimizes the maximum value of the remaining observation load, where... This represents the total satellite allocation scheme x for grid r. j The observed load. Formula (6) represents the conversion time ρ. i Total satellite allocation time s i The sum must not exceed the allocatable time Δt.

[0082] S3. Establish a potential energy game model for distributed satellite dynamic grid load allocation.

[0083] Let G P ={N,{A i},{U i} represents a game, where N = {1, 2, ..., n} is the set of participants, which in this chapter is the satellite set; A = A1 × A2 ... × A n For joint action sets, U i :A i →R is the local utility function of participant i. For the action set a = (a1, a2, ..., a...), ... n )∈A, where a i ∈A i It is an action of participant i.

[0084] U(x) is the objective function of the mathematical model P0 of the problem, and can also be used as the game theory tool G.P = {N, {A i}, {U i}} To ensure that the set of actions of participants satisfy the constraints of problem model P0, the initial set of actions of participants will be screened preliminarily. Specifically, when constructing the action space in the potential game model, the actions of participants that do not meet the constraints of P0 will be deleted. Under the above premise, the set of actions a i of participant i is expressed as:

[0085] a i = (x i1 ,x i2 ,...,x ij ,...,x im ) (7)

[0086] where x ij is the amount of time that satellite s i allocates to grid r j . Let a -i = (a1,a2...,a i-1 ,a i+1 ,...,a n ) be the set of actions of satellites other than satellite s i , denote the specific empty action of participant i. Therefore, each x ij in a i is zero.

[0087] According to the global utility function U(x), based on the Wonderful Life Utility rule, the local utility function of individuals is designed as

[0088]

[0089] where is the set of grids that satellite s i allocates in action a i , is the set of satellites that allocates to grid r j in action a space.

[0090] Then, the local utility function proposed in the above formula is normalized as

[0091]

[0092] In formula (9), a max is the maximum observation ability of satellites, b max is the maximum observation load value of grids, N max represents the maximum number of grids that satellites can be allocated, i.e.

[0093]

[0094] Let The local utility function U i (a i ,a i ) of satellite s -i is given as follows

[0095]

[0096] In potential game, Nash equilibrium is a key concept, which describes a state of game participants: when the strategies of other participants remain unchanged, all participants have no motivation to change their strategies unilaterally, then the state is called Nash equilibrium. The formal definition of Nash equilibrium is as follows.

[0097] Definition 6.1 Nash Equilibrium: For a game G P = {N, {A i} i∈N , {U i} i∈N}, action a N is called a Nash equilibrium if and only if for each player i, U

[0098]

[0099] Definition 6.2 Exact Potential Game: A game G P = {N, {A i} i∈N , {U i} i∈N , φ} is called an exact potential game if and only if its potential function φ: A i → R satisfies

[0100] Lemma 6.1 If the potential function φ is defined as formula 14, then the task allocation game G A = {N, {A i}, {U i}} is a potential game.

[0101]

[0102] Proof: According to the definition of local utility function, we have:

[0103]

[0104] Therefore, we can get:

[0105]

[0106] According to Definition 6.2, φ = 1 / P·U is the game G. P The potential energy function, and the game G with function φ as the potential energy function. P This is a strict potential energy game. Q.E.D.

[0107] Theorem 6.2 applies to the potential energy game G defined in Lemma 6.1. P ={N,{A i} i∈N ,{U i} i∈N ,φ}, the Nash equilibrium of its game a N That is, the optimal solution a to problem P0. * .

[0108] Proof: Assume the optimal solution Not a potential energy game G P The Nash equilibrium, that is, the existence of

[0109]

[0110] According to definition 6.2, we can obtain

[0111]

[0112] Therefore, we can conclude that...

[0113]

[0114] The above formula and a * This contradicts the condition that this is the optimal solution to problem P0. Therefore, the proof is complete.

[0115] S4. Obtain a distributed satellite grid load allocation scheme based on the selection-time change optimization response process.

[0116] For the potential game model of the above distributed satellite dynamic grid load allocation, a distributed satellite grid load allocation method based on a selective time-varying better reply process is provided below. A selective time-varying parameter better reply process (SeTVBRP) algorithm is adopted, which adopts a decentralized online collaborative mode and can dynamically adjust parameters to adapt to the changing actions and states of each game participant, and then is extended to the dynamic multi-stage DGA problem. In addition, unlike the recursive average action update mode commonly used in potential games, SeTVBRP adopts a single-step update method, which will only select one participant to update the action in one negotiation step, and has higher adaptability to satellite systems with limited communication conditions and delay tolerance. The specific process of the SeTVBRP algorithm is as follows:

[0117] First, an initial allocation scheme for all satellites is generated according to the greedy strategy, and then a satellite update action is randomly selected at each iteration time to generate a new allocation scheme. The satellite update action rules include: after updating the time-varying parameters ε(t) and ω(t) at the current time, the satellite s i The selective action set at time t is obtained by applying the selective action strategy Exchange allocation scheme information with other satellites in the information exchange satellite set Γ i After exchanging allocation scheme information with other satellites, the local performance function value is calculated according to formula (11) Where a i (t) is the action of satellite s i at time t, a -i (t) is the action set of satellites other than satellite s i at time t, and then an action is selected from the selective action set to form a better response set

[0118]

[0119] If the better response set of satellite s i at the current time t is not empty, a to-be-updated action (i.e. a to-be-updated allocation scheme) is selected from it with equal probability, and the next action a i (t+1) is updated to the selected to-be-updated action In the above process, the SeTVBRP algorithm is based on the traditional BRP algorithm, and the invention introduces two improved mechanisms of the selective action strategy and the time-varying parameter strategy.

[0120] (1) Selective action strategy

[0121] In the BRP algorithm, the performance function needs to be calculated for all potential actions in each iteration, but only one action is selected for the next round of update, which leads to low computational efficiency. To solve this problem, the SeTVBRP algorithm designs a selective action strategy. This strategy guides the participants to calculate the performance function for only part of the action set in each iteration, and then obtains a better action set B i (a). Specifically, the selective action strategy randomly selects ω(t) · |A i |actions from the action set A i |at each iteration t to form the selective action set B where ω(t) ∈ [0, 1] represents the selection ratio, which is a monotone non-decreasing value related to the iteration time t.

[0122] (2) Time-varying parameter strategy

[0123] The design idea of the time-varying parameter strategy comes from the simulated annealing algorithm, where the temperature ζ determines the amplitude of the noise, indicating the probability of the participants taking the best response action. When ζ tends to 0, the participants are more likely to choose the best response action. In the potential game model constructed in this chapter, the utility function U i has a similar effect on the parameter ε, where a larger ε provides a global perspective of the utility function, and the participants (i.e., satellites in the DGA problem) tend to take actions that reduce the sum of all grid observation loads. However, when ε tends to zero, the utility function U i will guide the participants to prefer actions that directly minimize the global objective (i.e., the maximum grid observation load). The time-varying parameter method designs a time-varying parameter ε that monotonically decreases as the iteration time t increases. This method ensures that the SeTVBRP algorithm focuses on exploring the possibility of better actions in the early iterations and utilizes and improves the current action in the later iterations.

[0124] The specific pseudo code of the SeTVBRP algorithm is as follows:

[0125]

[0126]

[0127] However, the original SeTVBRP algorithm can only solve single-stage DGA problems, and now the SeTVBRP algorithm is extended to the multi-stage case. Specifically, when the system state (including the grid load value β or the satellite observation capability α) changes, a single-stage decentralized coordination process will be triggered. Within a single stage, the SeTVBRP algorithm proposed in the previous section is used to generate an action a kUnlike SeTVBRP algorithm in single stage, the satellites s i The influence of the action in the last stage must be considered in the current stage. This influence mainly reflects in the transition time η ik which is closely related to the action in the last stage. Therefore, after receiving the new dynamic parameters, all satellites combine the action a k-1 of the last stage to get the initial action of the current stage. Then, the satellites in the constellation adopt a "relay" communication mechanism. Under the pre-set communication order, the satellites execute the SeTVBRP algorithm from the 4th line to the 16th line to update their allocation schemes in turn, and pass them to the next satellite in a relay manner until the termination condition is reached.

[0128] The cooperative termination condition is: ① the communication times reach the pre-set maximum iteration number T Max ; ② the satellites reach a consensus that the current action a k reaches the Nash equilibrium state, that is, in this state, all satellites cannot improve their local performance values by updating the action. Once the cooperative termination condition is reached, the satellite receiving the final action will broadcast the scheme to all satellites, thus ending the coordination of the decision stage. After that, the satellite receiving the final action will execute the scheme until another change in the system state triggers the coordination of the next decision stage.

[0129] The application also sets up an experimental process to verify the effect of the scheme.

[0130] The application designs two simulation experiment scenarios, named "small-scale scenario" and "large-scale scenario". In the small-scale scenario, the simulation involves 25 heterogeneous remote sensing satellites, aiming to simulate the decentralized coordination effect of the constellation when the number of satellites is small; the large-scale scenario involves 100 heterogeneous remote sensing satellites, to explore the complexity and efficiency of the decentralized coordination of the constellation when the number of satellites is large.

[0131] The basic parameters of the remaining experimental scenarios are shown in Table 1. The grid is divided into rectangular units with a longitude and latitude length of 10° to ensure the universality of the coverage. The observation load values of each grid are randomly generated between 30 and 80, which reflects the observation demand and load difference of different grid areas. The observation capacity of the satellites is initialized to an integer between 2 and 3, which reflects the difference in observation capacity of different satellites in different decision stages.

[0132] Table 1 Basic parameters of experimental scenarios

[0133]

[0134] The present application introduces a series of ablation experiments to explore the specific impact of time-varying parameter strategy and selective action strategy on the performance of SeTVBRP algorithm. By comparing the performance of the complete algorithm (i.e. SeTVBRP algorithm) with the original BRP algorithm, the BRP algorithm after introducing the time-varying parameter strategy (i.e. TVBRP algorithm), and the BRP algorithm after introducing the selective action strategy (i.e. SeBRP algorithm), the contribution of different strategies to the algorithm effect is quantitatively discussed. Table 2 provides the statistical results of each algorithm running independently for 50 times in a small-scale scenario, including the best value, the worst value, the average value of the target value after 50 algorithm runs average running time T, variance σ and the number of optimal solutions N best And the convergence curves of different algorithms in the regional scenario are as shown in Figure 7

[0135] From the statistical data in Table 2, it can be observed that the proposed SeTVBRP algorithm has the best performance in each performance indicator, showing the significant impact of time-varying parameter strategy and selective action strategy on improving the solution quality and efficiency of the algorithm:

[0136] The time-varying parameter strategy significantly improves the optimization ability of the algorithm by dynamically adjusting the balance between exploration and development. Specifically, the average target value of BRP algorithm is about 30% higher than that of TVBRP algorithm after introducing the time-varying parameter, and the average target value of SeBRP algorithm is about 48% higher than that of SeTVBRP algorithm. This result shows that the time-varying parameter strategy plays a key role in improving the optimization of target value.

[0137] The selective action strategy effectively reduces the computational cost by reducing the size of the solution space considered in each iteration. The calculation time of BRP algorithm is about 35% less than that of SeBRP algorithm, and the calculation time of TVBRP algorithm is about 38% less than that of SeTVBRP algorithm, verifying that the selective action strategy can effectively reduce the calculation time of the algorithm.

[0138] Table 2 Allocation results of different algorithms in a small-scale scenario

[0139]

[0140] In the phase of iteration number 100 to 300, SeBRP and SeTVBRP algorithms using selective action method show faster convergence speed. This is because the selective action strategy allows the algorithm to more concentratedly search the promising area in the solution space, thereby accelerating the early convergence process. In the phase of iteration number 300 to 500, TVBRP and SeTVBRP algorithms using time-varying method show better performance, indicating that the time-varying parameter strategy provides more fine adjustment in the later stage, which helps the algorithm more accurately approach the optimal solution. ​

[0141] In summary, the time-varying parameter strategy and the selective action strategy can effectively improve the original BRP algorithm, and significantly improve the quality and efficiency of the solution. Among them, the time-varying parameter strategy enhances the search ability of the algorithm for the optimal solution, and the selective action strategy improves the calculation efficiency of the algorithm. The synergistic effect of the two strategies enables the SeTVBRP algorithm to quickly and accurately find high-quality solutions when solving the potential game problem in this chapter.

[0142] The following is the comparison result of the scheme of the application and other algorithms.

[0143] The following is a set of comparative experiments to illustrate the effectiveness of the SeTVBRP algorithm proposed in the application. The comparison algorithms include: BRA

[150] , DT2A

[146] , SeSAP

[149] , TVLLA

[161] and TVCBLLA

[160] . The solving efficiency is reflected by two indicators: global target value and CPU solving time.

[0144] SeTVBRP algorithm successfully obtains 32 optimal solutions in 50 runs, with an average calculation time of about 0.7 seconds. In contrast, the number of optimal solutions obtained by BRA, DT2A, SeSAP and TVLLA algorithms is less than 15. Secondly, SeTVBRP algorithm has the smallest variance of 50 experimental results among the 6 algorithms, and the value of the worst-case solution is also the smallest, which shows that SeTVBRP algorithm has high robustness and good convergence in various experimental scenarios. In contrast, although BRA introduces a pure greedy strategy to calculate the solution time, it lacks randomness and performs poorly in terms of solution quality. While the DT2A algorithm has higher action selection randomness, the average solution quality is slightly better than BRA, but the calculation time is longer and the robustness is insufficient. The TVCBLLA algorithm cannot converge under 500 iterations, so the results after 20000 iterations are shown in Table 3, denoted as TVCBLLA(20000). Compared with SeTVBRP algorithm, TVCBLLA(20000) not only has more iterations, but also has longer calculation time, about 10 times of SeTVBRP algorithm, and its efficiency and robustness are not as good as SeTVBRP algorithm.

[0145] Table 3 Allocation results of different algorithms in small-scale scenarios

[0146]

[0147]

[0148] Next, we focus on the convergence curves of each potential game algorithm. After 200 iterations, the average objective value of SeTVBRP algorithm is significantly better than other baseline algorithms. Before 200 iterations, the performance of SeTVBRP algorithm is slightly worse than BRA, SeSAP and TVLLA algorithms. This phenomenon can be attributed to the strong greedy feature shared by these comparative algorithms, which tend to quickly capture better solutions in early iterations, but may sacrifice the ability to converge to the optimal solution in the long run. In contrast, the DT2A algorithm introduces greater randomness in the rules for updating participant actions. This design, while potentially leading to poor performance in the early evolution process, helps the algorithm avoid getting stuck in local optima and ultimately achieve better convergence results in the long run. It is worth noting that the initial solution value obtained by the DT2A algorithm is significantly different from other algorithms, as the DT2A algorithm initializes the solution in a random way rather than using the greedy approach of other algorithms. The SeTVBRP algorithm proposed in this invention balances the greediness and randomness. This balanced strategy enables the SeTVBRP algorithm to perform well in early iterations and gradually exhibit its best performance in subsequent iterations. Therefore, after 200 iterations, the average objective value of the SeTVBRP algorithm is better than other algorithms, showing that it can effectively narrow the solution space and quickly converge to high-quality solutions in the solution process.

[0149] In the algorithm comparison experiment under large-scale scenarios, the results are roughly consistent with the experimental results under small-scale scenarios, while revealing the advantages of the SeTVBRP algorithm in handling larger-scale problems. Due to the increase in the number of grids and satellites in large-scale scenarios, the number of iterations is increased from 500 to 2000 to ensure that the algorithm has enough iterations to explore the solution space. At the same time, the time-varying parameter τ value is set to 0.85 to adapt to large-scale scenarios.

[0150] In this case, the SeTVBRP algorithm proposed in this chapter still has an average objective value better than other algorithms. It is worth noting that the SeTVBRP algorithm obtained 17 optimal solutions in 50 runs, while other algorithms did not find the optimal solution, highlighting the significant advantage of the SeTVBRP algorithm in large-scale search capabilities compared to other algorithms.

[0151] Table 4 Allocation results of different algorithms under large-scale scenarios

[0152]

[0153]

[0154] From Figure 9It can be seen that the average objective value of SeTVBRP algorithm has a rapid decline phase when the iteration number reaches about 1700. In the convergence curve of small-scale scenario, there is also a similar trend when the iteration number is between 300 and 350, which indicates that the time-varying parameter strategy of SeTVBRP algorithm effectively enters the development and optimization stage in the later iteration. Therefore, since the value of τ is set to 0.85 at this time, the convergence curve of SeTVBRP starts to accelerate when the iteration reaches τT max = 1700 times. However, this does not mean that SeTVBRP can only achieve better results than other algorithms when τ = 0.85. In fact, as shown in Figure 6 , SeTVBRP algorithm is generally superior to similar algorithms when τ is in the range of 0.2 to 0.95, showing the wide applicability of the algorithm. In addition, although the performance of SeTVBRP algorithm during the early iterations may not be outstanding, its calculation time is the shortest among all algorithms, only 4.628 seconds.

[0155] Based on the above comparative test analysis, the summary of the results of the algorithm comparison experiment is as follows:

[0156] (1) Performance of SeTVBRP algorithm: In two different experimental scenarios of small-scale and large-scale scenarios, SeTVBRP algorithm shows significant performance: the average objective value of the algorithm reaches 0.66 and 1.16 respectively, and the algorithm can gradually approach the optimal solution as time goes to infinity. In addition, in terms of solution efficiency and stability of the problem in this chapter, SeTVBRP algorithm also achieves good results compared with the rest of the comparison algorithms.

[0157] (2) Applicability and limitations of the algorithm: Through the experimental results, it can be concluded that SeTVBRP algorithm has certain advantages in dealing with large-scale and complex task allocation problems. However, the algorithm may encounter performance bottlenecks when facing allocation problems with extremely small time granularity. Specifically, in order to further verify the robustness of SeTVBRP algorithm, the number of satellites is increased to 200 and 300 respectively. The results show that even in the case of a significant increase in the number of satellites, SeTVBRP algorithm can still maintain good performance. However, when the measurement granularity of the allocation time is reduced, for example, from minutes to seconds, the performance of the algorithm decreases. This is because as the time granularity decreases, the space of the set of available allocation actions will grow exponentially, greatly increasing the difficulty of the satellite finding the optimal solution in the huge action space.

[0158] Those skilled in the art should understand that the modules or steps of the present application described above can be realized by a general computer device, or alternatively, they can be realized by program codes executable by a computing device, so that they can be stored in a storage device and executed by a computing device, or they can be respectively manufactured into individual integrated circuit modules, or a plurality of modules or steps among them can be manufactured into a single integrated circuit module. The present application is not limited to any specific combination of hardware and software.

[0159] The above describes the specific embodiments of the present application in conjunction with the accompanying drawings, but is not a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications or changes made by those skilled in the art on the basis of the technical solutions of the present application without creative labor are still within the protection scope of the present application.

[0160] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A multi-satellite dynamic grid load distribution method based on potential energy game, characterized in that, The following steps are performed: S1. The ground station obtains satellite data and grid data divided by the satellite position based management terminal, and performs DGA problem modeling; N = {1, 2,..., n} represents a satellite index set, where satellite s i has an index of i, i ∈ N; M = {1, 2,..., m} represents a set of grid index, wherein the grid r j has an index j, j e M; S2. The ground station divides the DGA problem into a single-stage DGA problem according to the time slicing method, and performs sDGA problem modeling; wherein the single-stage DGA problem is an sDGA problem; Each grid r j requires a satellite to complete the observation task within the grid to reduce its observation load β i ; In sDGA problem modeling, satellite s i For grid r j The observation capability of α ij It indicates that satellite s i Assigned to grid r j The amount of time is x ij express; Definition of the grid r j of the remaining observed load quantity y j as follows: S3. A potential game model of distributed satellite dynamic grid load distribution is established; Let G P ={N,{A i },{U i } represents a game, where N = {1, 2, ..., n} is the set of satellites; A = A1 × A2 ... × A n For joint action sets, U i :A i →R is the local utility function of participant i; for a set of actions a = (ai, a2,..., an) e A, where; a n ∈ A i is an action of participant i; and i is a set of actions of all participants. S4. A distributed satellite grid load distribution scheme based on a selection time-varying optimal response process is obtained; According to the greedy strategy, an initial distribution scheme of all satellites is generated; At each iteration time, a satellite update action is randomly selected to generate a new distribution scheme.

2. The method of claim 1, wherein, In step S1, At any time, each satellite can only perform an imaging task in a specific grid; The observation load of the grid changes dynamically over time; The change of the observation load is directly affected by the increase and decrease of the number of imaging tasks in the grid.

3. The multi-satellite dynamic grid load distribution method based on potential game according to claim 2, wherein In the DGA problem modeling, for each satellite s i , the corresponding grid r j , and the time window w ij = [c ij , d ij ], where c ij represents the start time of the time window, d ij represents the end time of the time window.

4. The multi-satellite dynamic grid load distribution method based on potential game according to claim 1, wherein In step S1, in the modeling of the DGA problem, the state change point t k is a specific moment in time when the time window changes, and the time interval [t k , t k + Δt] is called phase k.

5. The multi-satellite dynamic grid load distribution method based on potential game according to claim 4, wherein In the DGA problem modeling, For satellite s i The observation ability of grid r j The observation ability is a non-negative constant; the observation load is a non-negative constant And the observation load The values of the observation ability and the observation load are non-negative constants; represents the set of visible grids of satellite s in decision phase k, i.e.: i represents the set of visible grids of satellite s in decision phase k, i.e.:

6. The multi-satellite dynamic grid load distribution method based on potential game according to claim 5, wherein In the DGA problem modeling, for the assignment grid set in decision stage k for satellite s i , and if there is no intersection between the assignment grid sets of two adjacent decision stages, i.e. then satellite s i needs to adjust its imaging payload angle to steer from a grid in set to a grid in set .

7. The multi-satellite dynamic grid load distribution method based on potential game according to claim 6, wherein In the DGA problem modeling, the switching time η i of satellite s ik in decision stage k is: H represents a non-negative constant, which represents the simplified imaging load conversion time; When the intersection of the set of allocated grids of satellite s i under k-1 decision phase and k decision phase is empty set, the transition time of two phases is H, otherwise 0.

8. The multi-satellite dynamic grid load distribution method based on potential game according to claim 7, wherein In step S2, if a satellite is assigned to two or more grids, the time for the satellite to perform imaging load attitude conversion between different grids is considered.

9. The multi-satellite dynamic grid load distribution method based on potential game according to claim 8, wherein In the sDGA problem modeling, represents the set of allocation grids of satellite s i , satisfying For satellite s i , its conversion time p i The calculation formula is as follows:

10. The multi-satellite dynamic grid load distribution method based on potential game according to claim 9, wherein In the sDGA problem modeling, in the decision stage k, the problem model P0 is constructed as follows where x ij is a decision variable representing the amount of time to assign satellite s i to grid r j ; Since the sDGA problem is a single-stage DGA problem divided from the DGA problem by time slicing, the satellite s i and the grid r j have the same physical meaning as the model DGA expression. Δt is the total allocable time for this phase; the objective function U(x) is an exponential function that minimizes the maximum of the residual observation load, where represents the observation load for grid r j in all satellites total allocation scheme x; denotes the conversion time p i and the total satellite allocation time s i The sum must not exceed the allocatable time At.

Citation Information

Patent Citations

  • Satellite group task allocation method based on group game

    CN115339651A

  • GNSS optimized aircraft control system and method

    US20110264307A1