Hierarchical intelligent cross-domain resource scheduling method for spatial information networks
By constructing a step-by-step intelligent cross-domain resource scheduling method in the spatial information network, combining task attributes and resource status information, resource scheduling is carried out in stages, the problems of task attribute differences and network scale expansion are solved, and the number of emergency tasks completed and resource utilization efficiency are improved.
Patent Information
- Application Number
- CN202310670688.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-07
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2043-06-07
AI Technical Summary
The existing technology fails to effectively consider the task attribute differences in different domains in spatial information networks, resulting in low number of emergency tasks, unable to adapt to high dynamic and differentiated task requirements, and as the network scale expands, it is difficult to efficiently utilize the resources of the entire network.
A progressive intelligent cross-domain resource scheduling method is constructed through neural networks, and the policy function and state value function of the inter-domain and intra-domain resource scheduling stages are combined with task attribute information and resource state information, and resource scheduling is divided into two phases inter-domain and intra-domain stages. A distributed training mechanism is used to dynamically adjust the resource scheduling strategy.
The number of completed tasks of emergency tasks is improved, the adaptability to differentiated and highly dynamic tasks is enhanced, the impact of network scale growth on solution complexity is reduced, and the resource scheduling performance is achieved.
Smart Images

Figure CN116582171B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of satellite communication technology, and in particular relates to a hierarchical intelligent cross-domain resource scheduling method, which can be used for intelligent scheduling of space information networks to obtain a higher number of task completions for the entire network. Background Art
[0002] Space information networks include communication systems that provide diverse service needs, navigation systems that provide navigation and positioning services, and observation systems that provide various monitoring and observation services. With the advancement of communication technology, space information networks are becoming heterogeneous, multifunctional, and large-scale. Typically, a system with a specific service function is defined as a domain, and each domain is independent of the others. During service provision, satellites within each domain do not share resources or exchange data with satellites in other domains. However, the increasing differentiation of mission requirements and the rapidly growing number of missions make it difficult to provide satisfactory and timely service guarantees using the resources of a single domain. This problem becomes particularly prominent when faced with emergencies requiring high timeliness. Therefore, breaking down the barriers between domain resources and enabling cross-domain resource scheduling, thereby efficiently utilizing resources across the entire network and improving the number of missions completed across the network, has become a key approach to service delivery in future space information networks.
[0003] Currently, there is very little research on cross-domain resource scheduling methods for spatial information networks, and the solutions considered do not consider the attribute differences of tasks in different domains, which is not practical. Therefore, it is necessary to design a cross-domain resource scheduling method that considers the attributes of tasks in each domain and provides efficient services for the task requirements of each domain. In addition, as the scale of the network continues to expand and the differentiation of task requirements in each domain increases significantly, cross-domain resource scheduling solutions need to focus on the differentiation and dynamic nature of task requirements, as well as the impact of network scale. To this end, it is necessary to design an inter-domain and intra-domain hierarchical intelligent cross-domain resource scheduling method to enhance the adaptability of cross-domain resource scheduling strategies and avoid the increase in solution complexity caused by the growth of network scale.
[0004] Qi Hao et al. proposed a two-stage cross-domain resource scheduling method for satellite networks in their paper “A Multi-aspect Expanded Hypergraph Enabled Cross-domain Resource Management in Satellite Networks” (IEEE Transactions on Communications, July 2022). This method focuses on the coordinated scheduling of resources between satellites in different domains, and by sharing resources in each domain, it achieves better resource scheduling performance than non-cross-domain scenarios. However, this method ignores the attributes of tasks in different domains, reduces all domain tasks to data, and assumes that all domain tasks have the same completion priority, which is unrealistic and will affect the number of urgent tasks completed. At the same time, since the essence of the two-stage cross-domain resource scheduling method designed by this method is the traditional solution idea of optimization problems, it is a static scheduling method and needs to be re-solved for different task requirements. Therefore, it cannot be applied to differentiated and highly dynamic task requirements. In addition, since the solution complexity of this method is high and is related to the network scale, it cannot be applied to space information networks with continuously expanding network scale. Summary of the Invention
[0005] The purpose of the present invention is to address the deficiencies of the above-mentioned existing technologies and propose a hierarchical intelligent cross-domain resource scheduling method for spatial information networks, so as to effectively improve the number of completed urgent tasks, enhance the adaptability to the differentiated and highly dynamic task requirements of various domains, avoid the increase in solution complexity caused by the growth of network scale, and obtain higher resource scheduling performance.
[0006] To achieve the above object, the technical solution adopted by the present invention includes the following steps:
[0007] (1) Construct a planned space information network with multiple satellite systems and ground stations, and regard each satellite system in the network as a domain;
[0008] (2) Taking τ as the time slot length, divide the planning time Ts of the spatial information network to be planned into T time slots: T = Ts / τ;
[0009] (3) Determine the auxiliary domain for each satellite to assist its transmission mission in addition to its own domain based on the number of domains K in the network;
[0010] (4) Determine the relay satellite set for each satellite:
[0011] (4a) According to the “one satellite, four links” model, determine the relay satellites for each satellite in its domain;
[0012] (4b) selecting one satellite in each of the two auxiliary domains of each satellite as an inter-domain relay satellite;
[0013] (4c) selecting intra-domain relay satellites and inter-domain relay satellites of each satellite to form a relay satellite set for the satellite;
[0014] (5) Determine the set of connection relationships between each satellite and the relay satellite and ground station:
[0015] (5a) The intersatellite link relationship between each satellite and its relay satellite within the communication range is evaluated as visible, which is represented by "1"; the intersatellite link relationship between each satellite and its relay satellite that is not within the communication range is evaluated as invisible, which is represented by "0";
[0016] (5b) The satellite-to-ground connection relationship of each ground station that can be covered by the satellite is evaluated as visible, represented by "1"; the satellite-to-ground connection relationship that cannot be covered is evaluated as invisible, represented by "0";
[0017] (5c) The inter-satellite connection relationships between each satellite in any time slot and all its intra-domain relay satellites constitute the intra-domain relay satellite connection relationship set of the satellite, the inter-satellite connection relationships between each satellite and all its inter-domain relay satellites constitute the inter-domain relay satellite connection relationship set of the satellite, and the satellite-to-ground connection relationships between each satellite and all its ground stations constitute the ground station connection relationship set of the satellite;
[0018] (6) Applying neural networks to construct the policy function and state value function of each satellite in the intra-domain and inter-domain resource scheduling phases, respectively, and setting the sets Mj = φ and Mn = φ for collecting training experience in the inter-domain and intra-domain resource scheduling phases, where φ represents the empty set;
[0019] (7) Setting the initial resource status and mission attribute information of all satellites in each training cycle;
[0020] (8) Obtaining the intra-domain and inter-domain joint status information of all satellites in the current time slot at each planning time slot;
[0021] (9) Select the inter-domain and intra-domain resource scheduling strategy for each satellite:
[0022] (9a) Determine the inter-domain action set of each satellite based on its intra-domain, inter-domain, and ground station connection relationship set, and apply the policy function of the inter-domain resource scheduling phase to select the inter-domain resource scheduling policy of each satellite from the inter-domain action set;
[0023] (9b) Determine the intra-domain action set of each satellite based on its intra-domain and inter-domain connection relationship set, and apply the policy function of the intra-domain resource scheduling phase to select the intra-domain resource scheduling policy executed by each satellite from the intra-domain action set;
[0024] (10) Execute the inter-domain resource scheduling strategy and the intra-domain resource scheduling strategy, obtain the benefit value and penalty value obtained by each satellite during the mission transmission process, and subtract the benefit value from the penalty value to obtain the reward of the intra-domain and inter-domain resource scheduling stage;
[0025] (11) Obtaining intra-domain joint state information of each satellite for the next time slot and inter-domain joint state information for the next time slot;
[0026] (12) The obtained joint state information of the current time slot, resource scheduling strategy, reward, time slot number and joint state information of the next time slot in the inter-domain and intra-domain resource scheduling phase are used as training experience data and stored in Mj and Mn respectively;
[0027] (13) Training the policy function and state value function of the inter-domain resource scheduling phase:
[0028] (13a) Determine whether the number of experience data |Mj| in the experience set of the inter-domain resource scheduling phase reaches the number of inter-domain training data. If so, execute (13b); otherwise, execute (14);
[0029] (13b) Update the training parameters of the inter-domain state value function and the policy function, and set Mj = φ;
[0030] (14) Training the policy function and state value function of the domain resource scheduling stage:
[0031] (14a) Determine whether the number of experience data |Mn| in the experience set of the intra-domain resource scheduling phase reaches the number of intra-domain training data. If so, execute (14b); otherwise, execute (15);
[0032] (14b) Update the training parameters of the state value function and the policy function in the domain, and set Mn = φ;
[0033] (15) Loop through (8) to (14) until the Tth time slot is completed;
[0034] (16) Loop through (7) to (15) until convergence, and obtain the trained policy function and state value function for the inter-domain and intra-domain resource scheduling phase;
[0035] (17) Using the trained inter-domain and intra-domain resource scheduling strategy functions, we can obtain the inter-domain and intra-domain resource scheduling strategies for each time slot and complete the tasks of each domain.
[0036] Compared with the prior art, the present invention has the following advantages:
[0037] 1. The present invention combines task attribute information with resource status information to construct inter-domain and intra-domain joint status information, and inputs it into the policy function and state value function constructed in the inter-domain and intra-domain resource scheduling stages to select and train scheduling strategies. By learning the characteristics of tasks in different domains and dynamically adjusting resource scheduling strategies, a higher number of task completions for the entire network can be achieved while meeting the requirements of tasks in different domains. This overcomes the problem that the existing technology ignores the attributes of tasks in different domains when performing resource scheduling, resulting in a low number of emergency task completions.
[0038] 2. Since the present invention divides cross-domain resource scheduling into two stages, inter-domain and intra-domain resource scheduling, a policy function and a state value function are constructed in each stage, and by training under various resource and task attribute states, not only a policy function with high adaptability to differentiated and dynamic task requirements can be obtained, but also the resource scheduling policy can be intelligently adjusted according to task requirements during resource scheduling, overcoming the problem that the static scheduling method used in the prior art cannot be applied to differentiated and highly dynamic task requirements.
[0039] 3. By treating each satellite as an intelligent agent and constructing a constant set of inter-domain and intra-domain actionable actions for it, distributed training is performed during the inter-domain and intra-domain resource scheduling phases. This ensures that the set of actionable actions used to select resource scheduling strategies at each phase does not scale with the growth of the network scale, effectively mitigating the impact of network scale on the complexity of resource scheduling solutions. This overcomes the high solution complexity of existing technologies, which makes them unsuitable for the ever-expanding space information network. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 It is an implementation flow chart of the present invention;
[0041] Figure 2 It is a schematic diagram of the spatial information network structure in the present invention;
[0042] Figure 3 This is a diagram showing simulation results of the number of urgent tasks in different first domains using the present invention and the existing method. DETAILED DESCRIPTION
[0043] The embodiments and effects of the present invention are described in further detail below with reference to the accompanying drawings.
[0044] Reference Figure 1 , the implementation steps of this example are as follows:
[0045] Step 1: Build the spatial information network to be planned.
[0046] Construct a system consisting of K satellites and N esThe space information network to be planned is constructed by taking each satellite system as a domain, and the domain set is D = {D1,…,D k ,…,D K}, where D k represents the kth domain, K≥2, N es ≥1;
[0047] Assume that each domain includes I orbits, each orbit includes J satellites, and the satellite set is: Among them, k∈{1,2,…,K}, I≥1, J≥2, represents the jth satellite of the i-th orbit of the k-th domain.
[0048] In this embodiment, a space information network is constructed including three satellite systems, namely a satellite communication system, a satellite observation system and a satellite navigation system, and four ground stations. Figure 2 As shown in the figure, the satellite communication system is called the first domain, denoted by D1, which includes 3 orbits, each orbit includes 3 communication satellites; the satellite observation system is called the second domain, denoted by D2, which includes 3 orbits, each orbit includes 3 observation satellites; the satellite navigation system is called the third domain, denoted by D3, which includes 3 orbits, each orbit includes 3 navigation satellites; the overall domain set D = {D1, D2, D3}, K = 3, N es =4.
[0049] Step 2: Divide the planning time of the spatial information network to be planned.
[0050] Taking τ as the time slot length, the planning time Ts of the spatial information network to be planned is divided into T time slots: T = Ts / τ.
[0051] In this embodiment, the time slot length τ=100s, the planning time Ts=6 hours, and the total number of time slots T=216.
[0052] Step 3: Determine the auxiliary domain for each satellite to perform mission transmission.
[0053] This step is to determine the auxiliary domain of each satellite to assist its transmission task in addition to its own domain based on the number of domains K in the network:
[0054] When K≤3, remove the kth domain D from the network. k The other domains outside the kth domain are the jth satellite of the i-th orbit of the k-th domain Auxiliary domain of
[0055] When K>3, remove the kth domain D from the network k Randomly select two domains from other domains as Auxiliary domain of , and let the kth domain D kAll satellites in select the same auxiliary domain.
[0056] In this embodiment, the auxiliary domains of each satellite in the first domain D1 are the second domain D2 and the third domain D3, the auxiliary domains of each satellite in the second domain D2 are the first domain D1 and the third domain D3, and the auxiliary domains of each satellite in the third domain D3 are the first domain D1 and the second domain D2.
[0057] Step 4: Determine the relay satellite set for each satellite.
[0058] 4.1) Using the "one satellite, four chains" model, establish the jth satellite of the ith orbit in the kth domain and the j-1th satellite of the ith orbit of the kth domain respectively The j+1th satellite of the ith orbit of the kth domain The jth satellite of the i-1th orbit of the kth domain and the jth satellite of the i+1th orbit of the kth domain The links between these four satellites establish a total of four inter-satellite links;
[0059] In each domain, the satellite orbit number i∈{1,2,…,I}, the satellite number of each orbit j∈{1,2,…,J}, the satellite orbit number and the satellite number of each orbit cannot exceed I and J respectively, and cannot be less than 1. According to this condition:
[0060] If i=I, i+1>I, then With the jth satellite of the 1st orbit of the kth domain Establish a link;
[0061] If i=1, i-1<1, then With the jth satellite of the Ith orbit of the kth domain Establish a link;
[0062] If j=J, j+1>J, then and the Jth satellite of the ith orbit of the kth domain Establish a link;
[0063] If j=1, j-1<1, then and the first satellite of the i-th orbit of the k-th domain Establish a link;
[0064] 4.2) Combine the above with Four satellites that establish intersatellite links and As Satellite's intra-domain relay satellite.
[0065] 4.3) Select one satellite from each of the two auxiliary domains of each satellite as an inter-domain relay satellite;
[0066] 4.4) The intra-domain relay satellites and inter-domain relay satellites of each satellite are selected to form the relay satellite set of the satellite.
[0067] This example determines the relay satellite set for each satellite as follows Figure 2 As shown, it is based on the satellite of the third domain For example, The relay satellites in the domain are and and Select the first field D1 Satellite and D2 in the second domain The satellite is its inter-domain relay satellite, The relay satellite set is Intra-domain and inter-domain relay satellites Figure 2 Marked with a rectangular frame.
[0068] Step 5: Determine the set of connection relationships between each satellite and the relay satellite and ground station.
[0069] 5.1) The inter-satellite link relationship between each satellite and its relay satellite within the communication range is evaluated as visible, represented by "1"; the inter-satellite link relationship between each satellite and its relay satellite that is not within the communication range is evaluated as invisible, represented by "0";
[0070] 5.2) The satellite-to-ground connection relationship where each satellite can cover the ground station is evaluated as visible, represented by "1"; the satellite-to-ground connection relationship where it cannot cover the ground station is evaluated as invisible, represented by "0";
[0071] 5.3) The inter-satellite connection relationships of each satellite in any time slot and all its intra-domain relay satellites form the intra-domain relay satellite connection relationship set of the satellite, which is expressed as follows:
[0072]
[0073] in, Indicates the tth time slot The set of intra-domain relay satellite connection relationships, Indicates the tth time slot The inter-satellite connection relationship with the relay satellites in its domain, express A collection of relay satellites within the domain, represents the jth satellite of the ith orbit of the kth domain, t∈{1,2,…,T};
[0074] 5.4) The inter-satellite connection relationships of each satellite in any time slot and all its inter-domain relay satellites form the inter-domain relay satellite connection relationship set of the satellite, which is expressed as follows:
[0075]
[0076] in, Indicates the tth time slot The set of inter-domain relay satellite connection relationships, Indicates the tth time slot The inter-satellite connection relationship with its inter-domain relay satellites, express A collection of inter-domain relay satellites;
[0077] 5.5) The satellite-to-ground connection relationships of each satellite in any time slot and all its ground stations form the ground station connection relationship set of the satellite, which is expressed as follows:
[0078]
[0079] in, Indicates the tth time slot The ground station connection relationship set, Indicates the tth time slot The satellite-to-ground connection relationship with the ground station.
[0080] Step 6: Construct intra-domain and inter-domain resource scheduling functions and initialize parameters.
[0081] 6.1) Set the sets Mj = φ and Mn = φ for collecting training experience in the inter-domain and intra-domain resource scheduling phases, where φ represents the empty set;
[0082] 6.2) Apply neural networks to construct the policy functions for each satellite's intra-domain and inter-domain resource scheduling phases, which are expressed as follows:
[0083]
[0084]
[0085] in, and Respectively Intra-domain and inter-domain policy functions, and Respectively The training parameters of the intra-domain and inter-domain policy functions, and Represents the tth time slot Intra-domain and inter-domain joint state information, and Represents the tth time slot The intra-domain and inter-domain scheduling strategy, P(·|·) represents the conditional probability, s and a represent the state and action of the satellite respectively, represents the jth satellite of the i-th orbit of the k-th domain;
[0086] In this embodiment, the neural network for constructing intra-domain and inter-domain policy functions is composed of a cascade of three parts: an input layer, a hidden layer, and an output layer. The input layer includes four parallel fully connected layers, the hidden layer includes one fully connected layer, and the output layer of the policy function includes a Softmax layer. The number of neurons in each fully connected layer is set to 32.
[0087] 6.3) Apply neural networks to construct the state value functions of each satellite's intra-domain and inter-domain resource scheduling phases, which are expressed as follows:
[0088]
[0089]
[0090] in, and Respectively The state value function within and between domains of and Respectively The training parameters of the state value function within and between domains, and Represents the tth time slot The rewards of the intra-domain and inter-domain resource scheduling stages are calculated, and E(·) represents the expected reward.
[0091] In this embodiment, the neural network for constructing intra-domain and inter-domain state value functions is composed of a cascade of three parts: an input layer, a hidden layer, and an output layer. The input layer includes four parallel fully connected layers, the hidden layer includes one fully connected layer, and the output layer includes one Linear layer. The number of neurons in each fully connected layer is set to 32.
[0092] Step 7: Set the initial resource status and mission attribute information of all satellites.
[0093] Set the initial resource status and mission attribute information of all satellites in each training cycle, including: initial resource status information, including the first time slot The amount of data for the storage task Remaining battery energy Initial mission attribute information, including the mission survival time slot Rs generated by the satellite in the kth domain k ,in, represents the jth satellite of the i-th orbit of the k-th domain.
[0094] In this embodiment, the first time slot is set The amount of data for the storage task Remaining battery energy The mission survival time slot Rs1 generated by the satellite of the first domain is 18, the mission survival time slot Rs2 generated by the satellite of the second domain is 72, and the mission survival time slot Rs3 generated by the satellite of the third domain is 6, where E max =100KJ represents the battery capacity of the satellite.
[0095] Step 8: Obtain the intra-domain and inter-domain joint status information of all satellites in the current time slot.
[0096] 8.1) Obtain the intra-domain joint status information of all satellites in the current time slot at each planning time slot, as shown below:
[0097]
[0098] in, Indicates the tth time slot Intra-domain federation state information, Indicates the tth time slot Local state information, S (t) (m) represents the tth time slot Local status information of relay satellites in the domain, express A collection of relay satellites within the domain, represents the jth satellite of the ith orbit of the kth domain, ∪ represents the union of two sets;
[0099] 8.2) Calculate the average value of the joint state information in the domain:
[0100]
[0101] in, Indicates the tth time slot The average value of the joint state information in the domain, |·| represents the number of elements in the acquired set;
[0102] 8.3) In each planning time slot, the inter-domain joint state information is formed by the average value of the intra-domain joint state information and the local state information of the inter-domain relay satellites, which is expressed as follows:
[0103]
[0104] in, Indicates the tth time slot Inter-domain joint state information, S (t) (n) represents the tth time slot The local status information of the inter-domain relay satellite, express A collection of inter-domain relay satellites.
[0105] In this embodiment, the tth time slot Local status information By the tth time slot Available communication resources Relative value of the amount of storage task data Relative value of remaining battery power and the average remaining lifetime slots of all stored tasks Composition, among which B max =60Gbits represents the capacity of satellite onboard memory, Indicates the tth time slot The amount of data stored in the task, Indicates the tth time slot The remaining energy of the battery, E min =(1-η)·E max Indicates the minimum remaining capacity of the battery, η=75% indicates the maximum discharge depth of the battery, E max =100KJ represents the satellite's battery capacity, E o =0.5KJ represents the energy consumed by the satellite to maintain normal operation in each time slot.
[0106] Step 9: Select the inter-domain and intra-domain resource scheduling strategies for each satellite.
[0107] 9.1) According to the jth satellite of the ith orbit of the kth domain in the tth time slot The set of intra-domain relay satellite connection relationships Time slot t The set of inter-domain relay satellite connection relations and the tth time slot Ground station connection relationship set Sure Inter-domain action set
[0108] if or Then D k Set as an action and add it to middle;
[0109] if Then it will be The domain where the satellite's connection relationship is visible and the relay satellite is located is set as an active action and added to middle;
[0110] 9.2) Application The policy function of the inter-domain resource scheduling phase from Inter-domain action set Select Inter-domain resource scheduling strategy
[0111] 9.3) According to the jth satellite of the ith orbit of the kth domain in the tth time slot The set of intra-domain relay satellite connection relationships and the tth time slot Ground station connection relationship set Sure The set of actions that can be taken within the domain
[0112] if Then Set as an actionable action and add it to In, it means Can directly transmit tasks to the ground station;
[0113] if Then it will be The satellite connection relationship is visible in the domain relay satellite is set as an actionable action and added to middle;
[0114] 9.4) Application The policy function of the domain resource scheduling phase from The set of actions that can be taken within the domain Select Intra-domain resource scheduling strategy
[0115] Step 10: Execute the resource scheduling strategy of each satellite and obtain its reward.
[0116] Execute the tth time slot Inter-domain resource scheduling strategy and domain resource scheduling strategies Obtain the amount of data successfully transmitted to the ground station during the mission transmission process, and use this amount of data as the revenue value;
[0117] Get The amount of data that failed to be received by the relay satellite is used as a penalty value if the satellite does not select the relay satellite as the intra-domain scheduling strategy. Then set the penalty value of the satellite to 0;
[0118] Will The difference between the benefit value and the penalty value is used to obtain the reward for the resource scheduling stage within the domain. and rewards for inter-domain resource scheduling
[0119] Step 11: Obtain the joint state information of all satellites in the next time slot and collect the training experience of each satellite.
[0120] 11.1) Acquisition Intra-domain joint state information for the next time slot and the inter-domain joint state information of the next time slot
[0121] 11.2) The joint state information of the inter-domain and intra-domain resource scheduling phases of the current time slot, resource scheduling strategy, reward, time slot sequence number, and the joint state information of the next time slot are used as training experience data and stored in Mj and Mn respectively.
[0122] Step 12: training the policy function and state value function of the inter-domain resource scheduling phase.
[0123] 12.1) Set the number of data points Qj required for inter-domain training. In this embodiment, Qj = 64.
[0124] 12.2) Determine whether the number of experience data |Mj| in the experience set during the inter-domain resource scheduling phase reaches the required number Qj of inter-domain training data:
[0125] If yes, go to step 12.3).
[0126] Otherwise, go to step 13;
[0127] 12.3) Update the training parameters of the inter-domain state value function:
[0128]
[0129] in, is the parameter of the inter-domain state value function that is cyclically updated during the training process, represents the parameter value of the inter-domain state value function obtained at the p+1th update, represents the parameter value of the inter-domain state value function obtained at the pth update, p represents the number of updates of the parameters of the inter-domain state value function, represents the learning rate of the inter-domain state value function, Express Find the gradient, express exist The evaluation value under the state, Indicates the tth time slot The reward for the inter-domain resource scheduling phase, express exist The state value in the state, and Represent the tth time slot and the t+1th time slot respectively The inter-domain joint state information, γ∈[0,1) represents the discount factor, represents the jth satellite of the i-th orbit of the k-th domain;
[0130] 12.4) Update the training parameters of the inter-domain policy function:
[0131]
[0132] in, are the parameters of the inter-domain policy function that are cyclically updated during the training process, represents the parameter value of the inter-domain strategy function obtained at the p+1th update, represents the parameter value of the inter-domain policy function obtained at the p-th update, Express Find the gradient, express The inter-domain timing difference error, α θj represents the learning rate of the inter-domain policy function, represents the multiplication operation;
[0133] 12.5) Set Mj = φ.
[0134] Step 13: training the policy function and state value function of the domain resource scheduling phase.
[0135] 13.1) Set the number of data points Qn required for in-domain training. In this embodiment, Qn=64.
[0136] 13.2) Determine whether the number of experience data |Mn| in the experience set during the domain resource scheduling phase reaches the required number of domain training data Qn:
[0137] If yes, go to step 13.3).
[0138] Otherwise, go to step 14;
[0139] 13.3) Update the training parameters of the state value function in the domain:
[0140]
[0141] in, is the parameter of the domain state value function that is cyclically updated during the training process, represents the parameter value of the domain state value function obtained at the q+1th update, Indicates the parameter value of the domain state value function obtained at the qth update, q indicates the number of updates of the parameters of the domain state value function, represents the learning rate of the state value function in the domain, Express Find the gradient, express exist The evaluation value under the state, Indicates the tth time slot The reward for the intra-domain resource scheduling phase, express exist The state value in the state, and Represent the tth time slot and the t+1th time slot respectively The joint state information between domains, γ∈[0,1) represents the discount factor, represents the jth satellite of the i-th orbit of the k-th domain;
[0142] 13.4) Update the training parameters of the policy function in the domain:
[0143]
[0144] in, are the parameters of the domain policy function that are cyclically updated during the training process, represents the parameter value of the domain policy function obtained at the q+1th update, represents the parameter value of the domain policy function obtained at the qth update, Express Find the gradient, express The inter-domain timing difference error, α θn represents the learning rate of the policy function in the domain, represents the multiplication operation;
[0145] 13.5) Set Mn = φ.
[0146] Step 14: cyclically execute steps 8 to 13 until the Tth time slot is completed.
[0147] In step 15, steps 7 to 14 are executed repeatedly until the sum of the rewards of all satellites converges, and the trained policy function and state value function of the inter-domain and intra-domain resource scheduling phase are obtained.
[0148] Step 16: Use the trained strategy function to complete the task.
[0149] 16.1) Using the trained policy functions for the inter-domain and intra-domain resource scheduling phase, obtain the inter-domain and intra-domain resource scheduling policies for each time slot:
[0150]
[0151]
[0152] in, Indicates all time slots The set of inter-domain scheduling strategies, Indicates all time slots A set of intra-domain scheduling strategies, Indicates the tth time slot Inter-domain scheduling strategy, Indicates the tth time slot Intra-domain scheduling strategy, represents the jth satellite of the ith orbit of the kth domain, t∈{1,2,…,T};
[0153] 16.2) Execute the inter-domain and intra-domain resource scheduling strategies for each time slot to complete the tasks of each domain.
[0154] In this embodiment, each satellite first obtains the inter-domain resource scheduling policy using the policy function of the inter-domain resource scheduling phase, and then obtains the intra-domain resource scheduling policy using the policy function of the intra-domain resource scheduling phase. Then, each satellite executes the inter-domain and intra-domain resource scheduling policies respectively, transmits the tasks stored on the satellite in each time slot, and completes the tasks of each domain.
[0155] The effects of the present invention are further described below in conjunction with simulation experiments:
[0156] 1. Simulation conditions
[0157] Construct a space information network consisting of three domains and 10 ground stations. The first domain is represented by D1, which includes six orbits, each of which includes 11 satellites; the second domain is represented by D2, which includes eight orbits, each of which includes six satellites; the third domain is represented by D3, which includes three orbits, each of which includes eight satellites;
[0158] Set the time slot length τ = 100s, the planning time Ts = 6 hours, and the total number of time slots T = 216;
[0159] Set each domain to generate both regular tasks and emergency tasks, and the regular task survival slot Rs1 generated by the satellite in the first domain is 18, the regular task survival slot Rs2 generated by the satellite in the second domain is 72, and the regular task survival slot Rs3 generated by the satellite in the third domain is 6;
[0160] Set the satellite's battery capacity E max=100KJ, the minimum remaining capacity of the battery is 25KJ, the maximum discharge depth of the battery is η = 75%, and the energy consumed by the satellite to maintain normal operation in each time slot is E o =0.5KJ, the capacity of satellite onboard memory B max =60Gbits, the data volume of the first time slot storage task of each satellite is 0 and the remaining energy of the battery in the first time slot is 60KJ;
[0161] Set the number of data required for inter-domain training Qj = 64, the number of data required for intra-domain training Qn = 64, and the learning rate of the inter-domain and intra-domain state value functions Learning rate α of inter-domain and intra-domain policy functions θj =α θn =0.00025, discount factor γ=0.99.
[0162] 2. Simulation content
[0163] Simulation experiment 1: Based on the above simulation conditions, the number of regular tasks generated by the first domain D1 is further set to 10560, the data volume of each task is 1Gbits, the number of regular tasks generated by the second domain D2 is 1824, the data volume of each task is 3Gbits, the number of regular tasks generated by the third domain D3 is 4800, the data volume of each task is 0.5Gbits, and the survival time slot of the emergency task generated by the first domain is set to 3. The present invention and the existing method are used to simulate different numbers of emergency tasks generated by the first domain. The simulation results are as follows: Figure 3 shown.
[0164] from Figure 3 As can be seen from the figure, the present invention achieves better performance for both routine and emergency tasks, and the total number of completed tasks increases with the number of emergency tasks. It can also be seen that the greater the number of emergency tasks, the more significant the performance improvement of the present invention compared to existing methods. This is because the present invention considers task attribute information when selecting resource scheduling strategies and divides the resource scheduling phase into two parts: inter-domain and intra-domain. This allows for better learning of inter-domain resource collaboration strategies. When intra-domain resources cannot be supplied in a timely manner, effective inter-domain collaboration can be used to find feasible task transmission solutions for emergency tasks.
[0165] Simulation experiment 2:
[0166] Based on the above simulation conditions, three new domains are added, namely the fourth domain D4, the fifth domain D5, and the sixth domain D6; D4 includes 6 orbits, each orbit includes 4 satellites; D5 includes 6 orbits, each orbit includes 10 satellites; D6 includes 8 orbits, each orbit includes 6 satellites.
[0167] Set the parameters for normal and emergency tasks:
[0168] The first domain D1 generates 6600 regular tasks and 3960 emergency tasks, with each task having a data volume of 1 Gbits.
[0169] The second domain D2 generates 960 regular tasks and 864 emergency tasks, with each task having a data volume of 3 Gbits.
[0170] The third domain D3 generates 2928 regular tasks and 1872 emergency tasks, with each task having a data volume of 0.5 Gbits.
[0171] The fourth domain D4 generates 2928 regular tasks and 1872 emergency tasks, with a data volume of 0.5 Gbits per task.
[0172] The fifth domain D5 generates 3600 regular tasks and 3600 emergency tasks, with each task having a data volume of 1 Gbits;
[0173] The sixth domain D6 generates 960 regular tasks and 864 emergency tasks, with each task having a data volume of 3 Gbits.
[0174] The fourth domain generates a regular task survival slot Rs4 = 6, the fifth domain generates a regular task survival slot Rs5 = 18, and the sixth domain generates a regular task survival slot Rs6 = 72;
[0175] The lifetime of all domain-generated emergency tasks is set to 3.
[0176] Under the above parameters, the present invention and the existing method are used to simulate the task completion performance of conventional tasks and emergency task modes under different network scales. The simulation results are shown in Table 1.
[0177] Table 1 Comparison of task completion performance of routine and emergency task modes under different network scales
[0178]
[0179] As can be seen from Table 1, the total number of tasks completed by the present invention is better than that of the existing methods under different numbers of domains. It can also be seen that as the network scale increases, the improvement in the total number of tasks completed by the present invention compared with the existing methods becomes more and more obvious. For example, when the network has 6 domains, the total number of tasks completed by the present invention is 2740 more than that of the existing methods. This is because the present invention does not need to increase the dimension of the feasible action set when the number of domains increases, avoiding the increase in solution complexity caused by the increase in network scale, and reduces the difficulty of training and learning the entire network through hierarchical scheduling between and within domains, thereby achieving better training effects.
[0180] The above simulation results show that the present invention can achieve a higher number of emergency task completions in the face of dynamically changing emergency task demands, and has good performance for both routine tasks and emergency tasks with differentiated demands, and has good adaptability to differentiated and dynamic task demands; at the same time, the present invention can achieve better completion performance for routine tasks and emergency task modes under different network scales, and can be applied to future large-scale spatial information networks.
[0181] The above description is only a specific example of the present invention. It is obvious that for professionals in this field, after understanding the content and principles of the present invention, it is possible to make various modifications and changes in form and details without departing from the principles and structure of the present invention. However, these modifications and changes based on the ideas of the present invention are still within the scope of protection of the claims of the present invention.
Claims
1. A hierarchical intelligent cross-domain resource scheduling method for spatial information networks, characterized in that: The steps include: (1) Construct a planned space information network with multiple satellite systems and ground stations, and regard each satellite system in the network as a domain; (2) Taking τ as the time slot length, divide the planning time Ts of the spatial information network to be planned into T time slots: T = Ts / τ; (3) Determine the auxiliary domain for each satellite to assist its transmission mission in addition to its own domain based on the number of domains K in the network; (4) Determine the relay satellite set for each satellite: (4a) According to the "one satellite, four links" model, determine the relay satellites of each satellite in its domain; (4b) selecting one satellite in each of the two auxiliary domains of each satellite as an inter-domain relay satellite; (4c) selecting intra-domain relay satellites and inter-domain relay satellites of each satellite to form a relay satellite set for the satellite; (5) Determine the set of connection relationships between each satellite and the relay satellite and ground station: (5a) The intersatellite link relationship between each satellite and its relay satellite within the communication range is evaluated as visible, which is represented by "1"; the intersatellite link relationship between each satellite and its relay satellite that is not within the communication range is evaluated as invisible, which is represented by "0"; (5b) The satellite-to-ground connection relationship of each ground station that can be covered by the satellite is evaluated as visible, represented by "1"; the satellite-to-ground connection relationship that cannot be covered is evaluated as invisible, represented by "0"; (5c) The inter-satellite connection relationships between each satellite in any time slot and all its intra-domain relay satellites constitute the intra-domain relay satellite connection relationship set of the satellite, the inter-satellite connection relationships between each satellite and all its inter-domain relay satellites constitute the inter-domain relay satellite connection relationship set of the satellite, and the satellite-to-ground connection relationships between each satellite and all its ground stations constitute the ground station connection relationship set of the satellite; (6) Applying neural networks to construct the policy function and state value function of each satellite in the intra-domain and inter-domain resource scheduling phases, respectively, and setting the sets Mj = φ and Mn = φ for collecting training experience in the inter-domain and intra-domain resource scheduling phases, where φ represents the empty set; (7) Setting the initial resource status and mission attribute information of all satellites in each training cycle; (8) Obtaining the intra-domain and inter-domain joint status information of all satellites in the current time slot at each planning time slot; (9) Select the inter-domain and intra-domain resource scheduling strategy for each satellite: (9a) Determine the inter-domain action set of each satellite based on its intra-domain, inter-domain, and ground station connection relationship set, and apply the policy function of the inter-domain resource scheduling phase to select the inter-domain resource scheduling policy of each satellite from the inter-domain action set; (9b) Determine the intra-domain action set of each satellite based on its intra-domain and inter-domain connection relationship set, and apply the policy function of the intra-domain resource scheduling phase to select the intra-domain resource scheduling policy executed by each satellite from the intra-domain action set; (10) Execute the inter-domain resource scheduling strategy and the intra-domain resource scheduling strategy, obtain the benefit value and penalty value obtained by each satellite during the mission transmission process, and subtract the benefit value from the penalty value to obtain the reward of the intra-domain and inter-domain resource scheduling stage; (11) Obtaining intra-domain joint state information of each satellite for the next time slot and inter-domain joint state information for the next time slot; (12) The obtained joint state information of the current time slot, resource scheduling strategy, reward, time slot number and joint state information of the next time slot in the inter-domain and intra-domain resource scheduling phase are used as training experience data and stored in Mj and Mn respectively; (13) Training the policy function and state value function of the inter-domain resource scheduling phase: (13a) Determine whether the number of experience data |Mj| in the experience set of the inter-domain resource scheduling phase reaches the number of inter-domain training data. If so, execute (13b); otherwise, execute (14); (13b) Update the training parameters of the inter-domain state value function and the policy function, and set Mj = φ; (14) Training the policy function and state value function of the domain resource scheduling stage: (14a) Determine whether the number of experience data |Mn| in the experience set of the intra-domain resource scheduling phase reaches the number of intra-domain training data. If so, execute (14b); otherwise, execute (15); (14b) Update the training parameters of the state value function and the policy function in the domain, and set Mn = φ; (15) Loop through (8) to (14) until the Tth time slot is completed; (16) Loop through (7) to (15) until convergence, and obtain the trained policy function and state value function for the inter-domain and intra-domain resource scheduling phase; (17) Using the trained inter-domain and intra-domain resource scheduling strategy functions, we can obtain the inter-domain and intra-domain resource scheduling strategies for each time slot and complete the tasks of each domain.
2. The method according to claim 1, characterized in that In step (1), a space information network to be planned is constructed, which includes K satellite systems and N es ground stations, and each satellite system is regarded as a domain, and the domain set is D = {D1,…,D k ,…,D K }, where D k represents the kth domain, K≥2, N es ≥1; Assume that each domain includes I orbits, each orbit includes J satellites, and the satellite set is: Among them, k∈{1,2,…,K}, I≥1, J≥2, represents the jth satellite of the i-th orbit of the k-th domain.
3. The method according to claim 1, characterized in that In step (3), the auxiliary domain of each satellite is determined according to the number K of domains in the network, which is implemented as follows: When K≤3, remove the kth domain D from the network. k The other domains outside the kth domain are the jth satellite of the i-th orbit of the k-th domain Auxiliary domain of When K>3, remove the kth domain D from the network k Randomly select two domains from other domains as Auxiliary domain of , and let the kth domain D k All satellites in select the same auxiliary domain.
4. The method according to claim 1, wherein The "one satellite, four links" model described in step (4a) and the selected intra-domain relay satellites are implemented as follows: Establish the jth satellite of the i-th orbit of the k-th domain respectively and the j-1th satellite of the ith orbit of the kth domain The j+1th satellite of the ith orbit of the kth domain The jth satellite of the i-1th orbit of the kth domain and the jth satellite of the i+1th orbit of the kth domain The links between these four satellites establish a total of four inter-satellite links; Combine the above with Four satellites that establish intersatellite links and As Satellite's intra-domain relay satellite.
5. The method according to claim 1, wherein The intra-domain relay satellite connection relationship set, the inter-domain relay satellite connection relationship set, and the ground station connection relationship set in step (5c) are respectively expressed as follows: in, Indicates the tth time slot The set of intra-domain relay satellite connection relationships, Indicates the tth time slot The inter-satellite connection relationship with the relay satellites in its domain, express A collection of relay satellites within the domain, Indicates the tth time slot The set of inter-domain relay satellite connection relationships, Indicates the tth time slot The inter-satellite connection relationship with its inter-domain relay satellites, express A collection of inter-domain relay satellites, Indicates the tth time slot The ground station connection relationship set, Indicates the tth time slot The satellite-to-ground connection relationship with the ground station, represents the jth satellite of the ith orbit of the kth domain, t∈{1,2,…,T}.
6. The method according to claim 1, wherein The policy functions of the intra-domain and inter-domain resource scheduling phases and the state value functions of the intra-domain and inter-domain resource scheduling phases in step (6) are expressed as follows: in, and Respectively Intra-domain and inter-domain policy functions, and Represent the training parameters of the intra-domain and inter-domain policy functions, and Represents the tth time slot Intra-domain and inter-domain joint state information, and Represents the tth time slot The intra-domain and inter-domain scheduling strategies, P(·|·) represents the conditional probability; and denote the state value functions within and between domains, respectively. and Represent the training parameters of the state value function within and between domains, and Represents the tth time slot The rewards of the intra-domain and inter-domain resource scheduling phase, E(·) represents the expected value, s and a represent the state and action of the satellite respectively, represents the jth satellite of the i-th orbit of the k-th domain.
7. The method according to claim 1, characterized in that The intra-domain and inter-domain joint state information in step (8) are represented as follows: in, and Represents the tth time slot Intra-domain and inter-domain joint state information, Indicates the tth time slot Local state information, S (t) (m) represents the tth time slot Local status information of relay satellites in the domain, Indicates the tth time slot The average value of the joint state information in the domain, Indicates the tth time slot The local status information of the inter-domain relay satellite, express A collection of relay satellites within the domain, express A collection of inter-domain relay satellites represents the jth satellite of the ith orbit of the kth domain, and ∪ represents the union of two sets.
8. The method according to claim 1, characterized in that The inter-domain action set in step (9a) is based on the jth satellite of the ith orbit of the kth domain in the tth time slot. The set of intra-domain relay satellite connection relationships Time slot t The set of inter-domain relay satellite connection relations and the tth time slot Ground station connection relationship set Sure Inter-domain action set if or Then D k Set as an action and add it to middle; if Then it will be The domain where the satellite's connection relationship is visible and the relay satellite is located is set as an active action and added to middle; The domain action set in step (9b) is based on and Sure The set of actions that can be taken within the domain if Then Set as an actionable action and add it to In, it means Can directly transmit tasks to the ground station; if Then it will be The satellite connection relationship is visible in the domain relay satellite is set as an actionable action and added to middle.
9. The method according to claim 1, characterized in that The training parameters of the updated inter-domain state value function and the policy function in step (13b) are expressed as follows: in, and are the parameters of the inter-domain state value function and the policy function that are updated cyclically during the training process, and They represent the parameter values of the inter-domain state value function and the policy function obtained at the p+1th update, and They represent the parameter values of the inter-domain state value function and the policy function obtained at the pth update, and p represents the number of updates of the parameters of the inter-domain state value function and the policy function. and α θj Represent the learning rates of the inter-domain state value function and the policy function, Express Find the gradient, express exist The evaluation value under the state, Indicates the tth time slot The reward for the inter-domain resource scheduling phase, express exist The state value in the state, and Represent the tth time slot and the t+1th time slot respectively Inter-domain federated state information, express The inter-domain temporal difference error, γ∈[0,1) represents the discount factor, represents the jth satellite of the ith orbit of the kth domain, and · represents the multiplication operation.
10. The method according to claim 1, characterized in that The training parameters of the state value function and the policy function in the update domain in step (14b) are expressed as follows: in, and are the parameters of the domain state value function and the policy function that are updated cyclically during the training process, and They represent the parameter values of the domain state value function and the policy function obtained at the q+1th update, and They represent the parameter values of the domain state value function and the policy function obtained at the qth update, q represents the number of updates of the parameters of the domain state value function and the policy function, and α θn Represent the learning rates of the state value function and the policy function in the domain, Express Find the gradient, express exist The evaluation value under the state, Indicates the tth time slot The reward for the intra-domain resource scheduling phase, express exist The state value in the state, and Represent the tth time slot and the t+1th time slot respectively The joint state information between domains, γ∈[0,1) represents the discount factor, represents the jth satellite of the ith orbit of the kth domain, and · represents the multiplication operation.
11. The method according to claim 1, wherein The inter-domain and intra-domain resource scheduling strategies obtained in step (17) are expressed as follows: in, Indicates all time slots The set of inter-domain scheduling strategies, Indicates all time slots A set of intra-domain scheduling strategies, Indicates the tth time slot Inter-domain scheduling strategy, Indicates the tth time slot Intra-domain scheduling strategy, represents the jth satellite of the ith orbit of the kth domain, t∈{1,2,…,T}.
Citation Information
Patent Citations
Intelligent resource joint scheduling method under environment uncertainty remote sensing satellite network
CN112422171A
Location-specific or range-based licensing system
US20060277312A1