Relay satellite system access and switching planning method and system based on reinforcement learning

By using a reinforcement learning-based approach to plan access and handover for relay satellite systems, the system dynamically optimizes mission access and handover decisions, thereby solving the problem of low mission planning efficiency in relay satellite systems and improving mission completion rate and resource utilization efficiency.

CN121036833AActive Publication Date: 2025-11-28CHINA ACADEMY OF SPACE TECHNOLOGY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511303581.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-12
Publication Date
2025-11-28
Estimated Expiration
2045-09-12

AI Technical Summary

Technical Problem

Existing relay satellite systems fail to effectively handle the differences in mission requests during mission planning, making it difficult to improve the planning efficiency of some types of missions, thus affecting mission completion rate and resource utilization efficiency.

Method used

A reinforcement learning-based approach is adopted to construct a satellite network model and model mission planning as a Markov decision process. The access and switching decisions are optimized by an agent, and the adaptive capability of deep reinforcement learning is used to dynamically optimize the access and switching decisions of the mission, and differentiated processing is performed for continuous and non-continuous missions.

Benefits of technology

It significantly improved the mission completion rate, reduced the number of handovers, optimized the utilization efficiency of satellite communication resources, and met the transmission needs of different types of missions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121036833A_ABST
    Figure CN121036833A_ABST
Patent Text Reader

Abstract

The invention provides a relay satellite system access and switching planning method based on reinforcement learning, and the method comprises the steps: determining a satellite parameter, a beam resource, and a time window, type and data size of a low-orbit satellite task through constructing a network model containing a low-orbit satellite and a medium-and-high-orbit relay satellite; constructing a target function for maximizing the sum of completion quantities of continuous and discontinuous transmission tasks so as to determine optimal access and switching time, relay satellites and wave beams; constructing an intelligent agent, and modeling a problem into a Markov decision process; and initializing Actor and Critic neural network parameters of the intelligent agent, and obtaining a target strategy for outputting an access and switching decision by interacting and iteratively optimizing the parameters with the environment. The invention further provides a relay satellite system access and switching planning system based on reinforcement learning. Therefore, the task completion rate and the resource utilization rate can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of satellite communication, in particular to a relay satellite system access and switching planning method and system based on reinforcement learning. BACKGROUND

[0002] With the continuous progress of space technology, the number of satellites in orbit is increasing, and the amount of data that needs to be transmitted to the ground is also growing exponentially. Among them, the visible time of low-orbit satellites and ground stations is limited and the revisit period is longer, which seriously restricts the real-time transmission efficiency of in-orbit data.

[0003] The tracking and data relay satellite (hereinafter referred to as relay satellite) in geosynchronous orbit can effectively extend the communication window length of low-orbit satellites and ground stations through data relay service due to its high orbital height, and plays a key role in the entire space information network. However, due to the extremely limited resources of relay satellites, in the face of rising task demand, how to scientifically allocate relay satellite resources to ensure task completion rate has become a research hotspot in the field of space information network.

[0004] Existing research on relay satellite system task planning models the task planning problem as a parallel machine scheduling problem with time windows, considers the relay satellite antenna as a "machine", and the task request of the user satellite as a "to be served work", and assumes that each work can only be provided by a machine without interruption. Single service, thereby converting task planning into a task scheduling problem on multiple unrelated machines. However, in fact, the tasks in the space information network have diversity characteristics, and the time window tolerance, service time and other parameters of different task requests differ significantly. The above research does not classify and process the differences in task requests during task planning, resulting in difficulty in improving the planning efficiency of some types of tasks.

[0005] Therefore, it is urgent to propose a relay satellite system access and switching planning method that can effectively improve the task completion rate, reduce the delay, and optimize the utilization efficiency of satellite communication resources. SUMMARY

[0006] In view of the deficiencies of the prior art, the present application provides a relay satellite system access and switching planning method and system based on reinforcement learning, which optimizes the time decision, relay satellite decision and satellite beam decision in low-orbit satellite access and switching, thereby effectively improving the task completion rate, reducing the switching times and improving the task completion rate.

[0007] In order to achieve the above technical effects, the present application provides a relay satellite system access and switching planning method based on reinforcement learning, comprising the steps of:

[0008] A satellite network model comprising low-orbit satellites and medium-high-orbit relay satellites is constructed, the positions and quantities of the satellites in the satellite network model are determined, the number of beam resources of the relay satellites is determined, and the time window, the type of task, and the amount of data of the low-orbit satellite task are determined;

[0009] A target function is constructed, the target function is to maximize the sum of the number of completed discontinuous transmission tasks and continuous transmission tasks, and the optimal time of access and switching, the optimal relay satellite, and the optimal relay satellite beam of the task are determined according to the target function;

[0010] An agent is constructed to model the access and switching planning problem as a Markov decision process;

[0011] The Actor neural network parameters and Critic neural network parameters of the agent are randomly initialized, the target function is optimized, and the Actor neural network parameters and Critic neural network parameters are iteratively optimized and updated by the agent in the process of interacting with the environment to obtain a target policy for outputting access and switching decisions.

[0012] Optionally, the target function is represented as:

[0013]

[0014] Wherein, t i represents the access and switching time decision, represents the satellite decision of access and switching, represents the beam decision of access and switching, i represents the i-th task, y i represents whether the i-th discontinuous transmission task is completed, represents whether the continuous transmission task is completed.

[0015] Optionally, whether the discontinuous transmission task is completed is represented as:

[0016]

[0017] Wherein, represents the last access time of the i-th task, represents the expected transmission time of the last access of the i-th task, end i represents the task end window of the i-th task; f(t) is defined as an indication function of whether the task is completed, when t≥0, that is, At this time, the task cannot be completed, f(t)=0; when t<0, that is, At this time, the task can be completed, f(t)=1.

[0018] Optionally, whether the continuity transmission task is completed transmission is represented as:

[0019]

[0020] wherein, represents the i th continuity task at the h th access selection of the relay satellite at the time point; represents that for the continuity task, the end time of the last access should be equal to the start time of this access.

[0021] Optionally, define the decision time step in the Markov decision process as v, and the i th decision task as v v , the state S v = [O v , U v ] includes the relay satellite beam state matrix O v and the task state vector U v ; the agent executes the access action v according to the state S The access action includes selecting a relay satellite and a beam; the environment updates the state according to the executed action, and gives the agent a reward value R;

[0022] wherein, the relay satellite beam state matrix O v is an N×K matrix, wherein the matrix element represents the remaining occupation time of the n th relay satellite k th beam at the v th decision time; the state vector U v contains the state information of the current to-be-decided task and the subsequent L to-be-decided tasks.

[0023] Optionally, the reward value R is specifically valued according to the following strategy:

[0024] For non-continuous tasks, if the task is completed, a first positive reward value R1 is given, if switching occurs, a first negative reward value R2 is given, and if the task fails, a second negative reward value R3 is given.

[0025] For continuous tasks, if the task is completed, a second positive reward value R4 is given, if switching occurs, a third negative reward value R5 is given, if the task fails, a fourth negative reward value R6 is given, and if delayed access occurs, a fifth negative reward value R7 is given; wherein, R1, R4>0; R2, R3, R5, R6, R7<0.

[0026] Optionally, the target strategy for outputting access and switching decisions is obtained by iteratively optimizing and updating the Actor neural network parameters and Critic neural network parameters of the agent during interaction with the environment.

[0027] ​Through the interaction between the agent and the environment, experience data is collected and stored in a replay pool, and the experience data is used to iteratively optimize and update parameters of the Actor neural network and the Critic neural network by calculating a discount reward and a value estimate, constructing a clipped importance sampling target and an entropy reward item, so as to obtain a target strategy for outputting access and switching decisions.

[0028] Optionally, the interaction between the agent and the environment includes:

[0029] State information of the relay satellite is acquired and feature encoded to obtain relay satellite feature encoding;

[0030] State information of a to-be-decided task and a subsequent task is acquired and feature encoded to obtain task feature encoding;

[0031] The relay satellite feature encoding and the task feature encoding are fused through a multi-head attention mechanism to obtain a fused feature vector;

[0032] The fused feature vector is input into an Actor neural network to obtain an action probability distribution, and input into a Critic neural network to obtain a state value estimate.

[0033] Optionally, the calculating a discount reward and a value estimate, constructing a clipped importance sampling target and an entropy reward item specifically includes:

[0034] The discount reward is calculated based on the following formula:

[0035] G t =R (z) (t+1)+γG t+1 ;

[0036] Wherein, γ is a discount factor, R (z) (t+1) represents a reward value obtained after the agent performs an action at the t+1th time step of the zth round;

[0037] The advantage estimate is calculated based on the following formula:

[0038] A t =R (z) (t+1)+γ×value (z) (t+1)-value (z) (t);

[0039] Wherein, value (z) (t) represents a state value of the tth time step of the zth round estimated by the Critic network;

[0040] The new-old strategy probability ratio is calculated based on the following formula:

[0041] r t = exp(logit (z) (t) - logπ θ (a t |S t ));

[0042] wherein, logit (z) (t) represents the probability of selecting an action by the Actor network policy composed of old parameters, π θ (a t |S t ) represents the probability of selecting an action a t by the Actor network policy composed of current parameters in state S t ;

[0043] The clipping target function is constructed as follows:

[0044]

[0045] wherein, ε is a clipping parameter

[0046] The entropy reward term is calculated as follows:

[0047]

[0048] wherein, β is an entropy coefficient.

[0049] On the other hand, the application also provides a reinforcement learning-based relay satellite system access and switching planning system for implementing the method as described above.

[0050] Compared with the prior art, the application can dynamically optimize the access and switching decisions of tasks, and significantly improve the task completion rate, by modeling the satellite communication access and switching problem as a Markov decision process, and utilizing the adaptive learning ability of deep reinforcement learning. The algorithm can comprehensively consider subsequent task information, and realize collaborative optimization planning of multiple tasks, to improve the utilization rate of relay satellite beam resources, based on the state definition of task time window and resource occupation. Meanwhile, the differences between continuous tasks and non-continuous tasks are modeled and processed, which can better meet the transmission requirements of different types of tasks. BRIEF DESCRIPTION OF DRAWINGS

[0051] Figure 1 The step flowchart of the reinforcement learning-based relay satellite system access and switching planning method provided by an embodiment of the application is shown in the figure;

[0052] Figure 2 The relay satellite resource transfer schematic diagram of the reinforcement learning-based relay satellite system access and switching planning method provided by an embodiment of the application is shown in the figure;

[0053] Figure 3 The reinforcement learning training flowchart of the relay satellite system access and switching planning method based on reinforcement learning provided by an embodiment of the present application. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0055] It should be noted that the use of the terms "one embodiment", "an embodiment", "example embodiment", etc. in the specification refers to the described embodiment including a particular feature, structure or characteristic, but not every embodiment necessarily includes the particular feature, structure or characteristic. In addition, such expressions do not refer to the same embodiment. Further, when a particular feature, structure or characteristic is described in connection with an embodiment, it is indicated that such a feature, structure or characteristic is incorporated into other embodiments within the knowledge of those skilled in the art, whether or not it is explicitly described.

[0056] In addition, some terms are used in the specification and subsequent claims to refer to specific components or parts, and those skilled in the art should understand that manufacturers can use different names or terms to refer to the same component or part. The present specification and subsequent claims do not distinguish components or parts by name, but by functional differences. Throughout the specification and subsequent claims, "including" and "containing" are open terms, which should be interpreted as "including but not limited to". In addition, the term "connected" includes any direct and indirect electrical connection means. Indirect electrical connection means includes connection through other devices.

[0057] The following will further describe in detail a method for planning access and switching of a relay satellite system based on reinforcement learning provided by an embodiment of the present application in combination with the drawings of the specification, as shown in the drawings, the specific implementation of the method can include the following steps: Figure 1

[0058] S101: Construct a satellite network model including low-orbit satellites and medium-high-orbit relay satellites, determine the positions and quantities of the satellites in the satellite network model, the number of beam resources of the relay satellites, and determine the time window, task type and task data volume of the low-orbit satellite task; wherein the time window includes a start time window and an end time window.

[0059] ​S102: Construct a target function, which is to maximize the sum of the number of completed non-continuous transmission tasks and continuous transmission tasks; and determine the optimal time of access and switching, the optimal relay satellite, and the optimal relay satellite beam of the task according to the target function.

[0060] Specifically, the target function of the embodiment is expressed as:

[0061]

[0062] Where, t i represents the access and switching time decision, represents the satellite decision of access and switching, represents the beam decision of access and switching, i represents the ith task, y i represents whether the ith non-continuous transmission task is completed transmission, represents whether the continuous transmission task is completed transmission.

[0063] S103: Construct an agent to model the access and switching planning problem as a Markov decision process; that is, the embodiment models the optimization problem as a Markov decision process by constructing an agent.

[0064] S104: Randomly initialize the Actor neural network parameters and Critic neural network parameters of the agent, solve the optimization target function, and update the Actor neural network parameters and Critic neural network parameters through the agent in the process of interacting with the environment to obtain a target strategy for outputting access and switching decisions. The target strategy of the embodiment refers to the low-orbit satellite with a task accessing or switching to a specified relay satellite and the corresponding beam at a specified time.

[0065] In the process of access and switching of low-orbit satellite tasks, the beam resources of relay satellites (GEO and MEO) need to be dynamically adjusted according to the task transmission demand, Figure 2 which intuitively shows this process.

[0066] The following will be described in detail in combination with a specific application scenario.

[0067] Define the ith task, the low-orbit satellite m, make the hth access decision, and the access time selected by the decision is The selected relay satellite is n, and the selected beam size is The expected transmission time is denoted as

[0068] Further, the release time of each access is calculated according to the access selection:

[0069]

[0070] where, denotes the end time of the zth time window for the low earth orbit satellite n and the decision selected middle-high earth orbit relay satellite m.

[0071] If the expected service time is greater than the remaining service time, switching is performed at the time when the service is not visible, and the access decision is re-performed, assuming that the access decision is re-performed at the time when the service is not visible If the expected service time is still greater than the remaining service time at this time, the third access is performed, assuming that a total of H i accesses are performed.

[0072] Each access will occupy beam resources, and the occupancy relationship between the low earth orbit satellite m and the relay satellite n at any time is calculated:

[0073]

[0074] Further, whether the non-continuous transmission task is completed transmission is represented as:

[0075]

[0076] where, denotes the time of the last access of the ith task, denotes the expected transmission time of the last access of the ith task, end i denotes the task end window of the ith task; f(t) is defined as an indication function of whether the task is completed, when t≥0, that is, at this time the task cannot be completed, f(t)=0; when t<0, that is, at this time the task can be completed, f(t)=1.

[0077] Whether the continuous transmission task is completed transmission is represented as:

[0078]

[0079] where, denotes the relay satellite selected by the ith continuous task at the time when the hth access is performed; denotes that for the continuous task, the end time of the last access should be equal to the start time of this access.

[0080] In the step S2, the optimization objective mathematical expression is

[0081]

[0082] s.t.

[0083]

[0084]

[0085] wherein, indicates access constraint, only the last access is successful, this time can make access decision; indicates that at the same time, the low-orbit satellite only accesses one relay satellite; indicates that each access needs to meet the visibility requirement, indicates that the number of accessed user satellites cannot exceed the relay satellite beam capacity, indicates that the continuous task constraint the start time of the hth access should be equal to the end time of the (h-1) th access.

[0086] The embodiment defines the decision time step in the Markov decision process as v, and the vth decision task as i v , the state S v =[O v ,U v ] includes the relay satellite beam state matrix O v and the task state vector U v ; the intelligent agent performs access action according to the state S v ; the environment updates the state according to the executed action, and gives the intelligent agent a reward value R;

[0087] wherein, the relay satellite beam state matrix O v is an N×K matrix, wherein the matrix element indicates the remaining occupation time of the nth relay satellite kth beam at the vth decision time; the state vector U v contains the state information of the current to-be-decided task and the subsequent L to-be-decided tasks.

[0088] The reward value R is specifically valued according to the following strategy:

[0089] For non-continuous tasks, if the task is completed, a first positive reward value R1 is given, if switching occurs, a first negative reward value R2 is given, and if the task fails, a second negative reward value R3 is given;

[0090] For continuous tasks, if the task is completed, a second positive reward value R4 is given, if switching occurs, a third negative reward value R5 is given, if the task fails, a fourth negative reward value R6 is given, and if delayed access occurs, a fifth negative reward value R7 is given; wherein, R1, R4>0; R2, R3, R5, R6, R7<0.

[0091] In a specific embodiment, step S103 includes the following steps:

[0092] Step S3.1, sort the tasks according to the start time window size to make a decision queue, define the decision time step v = 1, 2, 3, … done, the vth decision task is i v , define the state S v =

[0093] [O v ,U v ], O v represents the relay satellite beam state, U v represents the task state. For example, tasks T1 (start time window 8:00-10:00), T2 (start time window 8:30-11:00), T3 (start time window 9:00-12:00) are arranged in the order [T1, T2, T3]. Define the decision time step as t (t = 1, 2, 3, …), the task handled at the tth step is the tth task in the queue (such as t = 1, handle T1).

[0094] The O v is a relay satellite beam state matrix, which is represented as:

[0095]

[0096] Wherein represents how much time is left for the nth relay satellite kth beam to be released, and whenever a task selects a relay satellite and the corresponding beam, the state will be transferred according to the expected transmission time of the task.

[0097] The U v represents a task state vector, which is represented as:

[0098]

[0099] The is the vth decision task state vector, which is represented as:

[0100]

[0101] Wherein, represents the number of times the task accesses, represents the task type, represents the task time window, represents the expected transmission time of the task in the low-orbit satellite under various beam modes, represents the remaining visible time of the low-orbit satellite and each relay satellite to which the task belongs.

[0102] Step S3.2, perform actions according to the state:

[0103]

[0104] Step S3.3, update the environment state according to the executed action:

[0105]

[0106] wherein, denotes the release time of the current access, t v denotes the decision time of the current access.

[0107] The is denoted as:

[0108]

[0109] wherein, denotes the expected release time, the vth task makes a decision at t v , selects the kth beam of the nth relay satellite, the beam is released after , and the expected transmission time is The expected release time and the visible time of the relay satellite are taken as the minimum value to obtain the actual release time of the current access.

[0110] Step S3.4, when the decision made by the task results in that the transmission cannot be completed within the visible service time of the corresponding relay satellite, the remaining tasks are inserted into the subsequent task decision queue, and the subsequent access is re-performed.

[0111] Step S3.5, set the reward R according to the state and action and the task completion. R denotes the reward after the action interacts with the environment. For non-continuous tasks, after making a decision, it can be judged whether the task can be completed within the task time window. If it can be completed, R1 is given, if switching will occur, reward R2 is given, and if the task cannot be completed, reward R3 is given. For continuous tasks, after making a decision, it can be judged whether the task can be completed within the task time window. If it can be completed, R4 is given, if switching will occur, reward R5 is given, if the task cannot be completed, reward R6 is given, and if there is a delayed access, negative reward R7 is given.

[0112] Through the above steps, the access and switching planning problem of the relay satellite system is converted into a Markov decision process in this embodiment, so that the intelligent agent can learn the optimal decision strategy through the interaction with the environment, and the task completion rate is improved and the resources are efficiently utilized.

[0113] Further, the iterative optimization and updating of the Actor neural network parameters and the Critic neural network parameters by the agent in the process of interacting with the environment to obtain a target strategy for outputting access and switching decisions comprises:

[0114] Through the interaction of the agent and the environment, experience data is collected into a replay pool, and the parameters of the Actor neural network and the Critic neural network are iteratively optimized and updated to obtain a target strategy for outputting access and switching decisions by calculating discount rewards and value estimates, constructing a cut importance sampling target and an entropy reward item, using the experience data.

[0115] The interaction of the agent and the environment comprises:

[0116] Obtain state information of the relay satellite and perform feature encoding to obtain relay satellite feature encoding;

[0117] Obtain state information of the to-be-decided task and the subsequent task and perform feature encoding to obtain task feature encoding;

[0118] Fuse the relay satellite feature encoding and the task feature encoding through a multi-head attention mechanism to obtain a fusion feature vector;

[0119] Input the fusion feature vector into the Actor neural network to obtain an action probability distribution, and input the fusion feature vector into the Critic neural network to obtain a state value estimate.

[0120] The calculation of the discount reward and the value estimate, the construction of the cut importance sampling target and the entropy reward item specifically comprises:

[0121] The discount reward is calculated based on the following formula:

[0122] G t =R (z) (t+1)+γG t+1 ;

[0123] Wherein, γ is a discount factor, R (z) (t+1) represents a reward value obtained after the agent performs an action at the t+1 time step of the z rounds;

[0124] The advantage estimate is calculated based on the following formula:

[0125] A t =R (z) (t+1)+γ×value (z) (t+1)-value (z) (t);

[0126] Wherein, value (z)(t) represents the state value of the zth round and the tth time step estimated by the Critic network;

[0127] The new and old policy probability ratio is calculated based on the following formula:

[0128] r t = exp(logit (z) (t) - logπ θ (a t | S t ));

[0129] Wherein, logit (z) (t) represents the probability of selecting an action by the Actor network policy composed of old parameters, π θ (a t | S t ) represents the probability of selecting an action a t by the Actor network policy composed of current parameters under state S t .

[0130] The clipping target function is constructed as follows:

[0131]

[0132] Wherein, ε is the clipping parameter

[0133] The entropy reward term is calculated as follows:

[0134]

[0135] Wherein, β is the entropy coefficient.

[0136] In a specific embodiment, referring to Figure 3 , step S104 includes the following steps:

[0137] Step S4.1, initialize the parameters θ actor of the Agent Actor neural network and the parameters θ critic of the Critic neural network, empty the experience replay pool M, set the learning rate lr, the number of rounds, the maximum number of steps per round max_step, the number of parameter updates K, the gradient clipping threshold ε, and the discount reward γ.

[0138] Step S4.2, obtain the GEO and MEO state S tdrs (z) (t) from the environment, input the satellite state encoding module, extract the satellite resource features, and obtain the unified relay satellite feature encoding H tdrs (z) (t).

[0139] Step S4.3: Obtain the task to be decided and subsequent tasks. Input the task status coding submodule to extract the basic attributes of the task and the correlation attributes of subsequent tasks with the current task decision, and obtain a unified task feature code H. task (z) (t).

[0140] Step S4.4: Use the multi-head attention module to process the H from step S4.2. tdrs (z) (t) and H in step S4.3 task (z) (t) A multi-head attention fusion layer is used to capture satellite-mission cross-modal correlations, outputting a state vector H containing deep correlation features. fusion (z) (t).

[0141] Step S4.5: Take H from step S4.4... fusion (z) (t) is input into the Actor network to obtain the probability distribution of each action, where the action is used. (z) (t) represents the probability distribution of each action at time step t in round z. We select the action with the highest probability from the action distribution and its corresponding log probability logit. (z) (t)

[0142] Step S4.6: Take H from step S4.4 fusion (z) (t) is input into the Critic network to obtain the value of the state at time step t in the z-th episode. (z) (t). And according to steps S3.3, S3.4, and S3.5, the state space set S for the next time step is obtained. (z) (t+1), Reward R (z) (t+1), update the decision queue.

[0143] Step S4.7: Transfer experience points (z) (t), action (z) (t),logit (z) (t),R (z) (t+1)> is stored in the playback pool M.

[0144] Step S4.8: After each complete sampling of the environment, use the empirical tuples in the buffer pool to iteratively optimize and train the Actor network and Critic network until the set number of iterations is reached.

[0145] ​Step S4.8, after each time the environment is sampled once, the experience tuples in the buffer pool are used to iteratively optimize and train the Actor network and the Critic network.

[0146] Step S4.9, calculate the discounted return G t For the samples in the experience pool, iterate in reverse order according to time sequence, calculate the discounted return:

[0147] G t = R (z) (t+1) + γG t+1 ;

[0148] Step S4.9, calculate the advantage estimate A t :

[0149] A t = R (z) (t+1) + γ × value (z) (t+1) - value (z) (t);

[0150] Perform normalization on all A t to obtain to improve the numerical stability of gradient update.

[0151] Step S4.10, input each experience into the current Actor network again to obtain the new log probability

[0152] logπ θ (a t |S t );

[0153] Calculate the ratio combining the stored old probability logit (z) (t):

[0154] r t = exp(logit (z) (t) - logπ θ (a t |S t ));

[0155] Step S4.11, construct the clipped importance sampling target:

[0156]

[0157] Calculate the entropy reward:

[0158]

[0159] Get the Actor loss:

[0160]

[0161] Step S4.12, constructing Critic mean square error loss:

[0162]

[0163] Step S4.14, randomly shuffling all samples and dividing them into batches of size B, simultaneously minimizing L on each batch Actor and L Critic . Repeat the traversal of the dataset for a preset number of update rounds K using the Adam optimizer.

[0164] Step S4.15, after completing K rounds of parameter updates, copying the latest parameters (θ actor ,θ critic ) as the old parameters (θ actor,old ,θ critic,old ) and emptying the replay pool M to start the next round of sampling interactions.

[0165] In another embodiment, the application also provides a reinforcement learning-based relay satellite system access and switching planning system, which is used to implement the method described in the above embodiments.

[0166] In a specific embodiment, the system comprises:

[0167] a network model construction module, configured to construct a satellite network model comprising low-orbit satellites and medium-high-orbit relay satellites, determine the positions and quantities of the satellites, the number of beam resources of the relay satellites, and the time window, task type, and task data volume of the low-orbit satellite tasks;

[0168] a target function construction and decision module, configured to construct a target function aiming to maximize the sum of the numbers of completed non-continuous transmission tasks and continuous transmission tasks, and determine the optimal time of access and switching, the optimal relay satellite, and the optimal relay satellite beam of the task belonging low-orbit satellite according to the target function;

[0169] an agent modeling module, configured to construct an agent and model the access and switching planning problem as a Markov decision process;

[0170] a neural network training module, configured to randomly initialize the Actor neural network parameters and Critic neural network parameters of the agent, collect experience data through the interaction between the agent and the environment, and update the parameters using a proximal policy optimization algorithm to obtain a target policy outputting access and switching decisions.

[0171] The modules work collaboratively to efficiently implement the relay satellite system access and switching planning method through modular design, improve the scalability and maintainability of the system, and ensure the optimization effect of the low-orbit satellite task access and switching decisions.

[0172] In summary, the application models the satellite communication access and switching problem as a Markov decision process, uses the adaptive learning ability of deep reinforcement learning, can dynamically optimize the access and switching decision of the task, and significantly improves the task completion rate; Based on the state definition of the task time window and the resource occupation condition, the algorithm can consider the subsequent task information comprehensively, realize the cooperative optimization planning of multiple tasks, and improve the utilization rate of the relay satellite beam resources; At the same time, the differences between continuous tasks and discontinuous tasks are modeled and processed, which can better meet the transmission requirements of different types of tasks.

[0173] It should be noted that in this paper, the term "include", "contain" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such process, method, article or device. Without more limitations, the element defined by the sentence "including a" does not exclude the existence of other identical elements in the process, method, article or device including the element. In addition, it should be pointed out that the scope of the method and device in the embodiments of the application is not limited to the order of the functions shown or discussed, and can also include the functions performed in a substantially simultaneous manner or in the opposite order according to the functions involved, for example, the described method can be performed in an order different from that described, and various steps can also be added, omitted or combined. In addition, the features described with reference to some examples can be combined in other examples.

[0174] Of course, the present application can have other various embodiments, and those skilled in the art can make various corresponding changes and modifications according to the present application without departing from the spirit and essence of the present application, but these corresponding changes and modifications should all belong to the protection scope of the claims attached to the present application.

Claims

1. A relay satellite system access and handover planning method based on reinforcement learning, characterized in that, Including the following steps: Construct a satellite network model that includes low-Earth orbit satellites and medium-to-high-Earth orbit relay satellites, determine the position and number of each satellite in the satellite network model, the number of beam resources for relay satellites, and determine the time window, mission type, and mission data volume for low-Earth orbit satellite missions; Construct an objective function that maximizes the sum of the number of completed discontinuous transmission tasks and continuous transmission tasks; and determine the optimal time for access and handover of the low-Earth orbit satellite to which the task belongs, the optimal relay satellite, and the optimal relay satellite beam based on the objective function. Construct intelligent agents to model the access and handover planning problem as a Markov decision process; The Actor neural network parameters and Critic neural network parameters of the agent are randomly initialized, the objective function is solved and optimized, and the Actor neural network parameters and Critic neural network parameters are iteratively optimized and updated by the agent during its interaction with the environment to obtain the target strategy for outputting access and handover decisions.

2. The method according to claim 1, characterized in that, The objective function is expressed as: Among them, t i Indicates access and handover time decisions, Indicates satellite access and handover decisions, This represents the beam decision for access and handover, where i represents the i-th task, and y represents the beam decision for access and handover. i This indicates whether the i-th discontinuous transmission task has completed transmission. Indicates whether the continuous transmission task has been completed.

3. The method according to claim 2, characterized in that, Whether the discontinuous transmission task has completed transmission is indicated by: in, This represents the time when the i-th task was last accessed. This represents the expected transmission time of the last access of the i-th task, end. i This represents the task completion window for the i-th task; f(t) is defined as an indicator function for whether the task is completed, i.e., when t≥0, i.e. The task cannot be completed at this point, f(t) = 0; when t < 0, i.e. The task can be completed at this point, f(t) = 1.

4. The method according to claim 2, characterized in that, Whether the continuous transmission task has been completed is indicated by: in, Indicates the i-th continuous task in The relay satellite that performs the h-th access selection at any given time; This means that for continuous tasks, the end time of the last access should be equal to the start time of this access.

5. The method according to claim 1, characterized in that, Let v be the decision time step in the Markov decision process, and let i be the v-th decision task. v State S v =[O v U v Including relay satellite beam state matrix O v and task state vector U v The intelligent agent, according to state S v Execute access action The access action includes selecting a relay satellite and beam; the environment updates its state based on the executed action and awards a reward value R to the agent; Among them, the relay satellite beam state matrix O v Let be an N×K matrix, where the matrix elements are... This represents the remaining occupancy time of the k-th beam of the n-th relay satellite at the v-th decision time; the state vector U v It includes the status information of the current task to be decided and the subsequent L tasks to be decided.

6. The method according to claim 5, characterized in that, The reward value R is assigned according to the following strategy: For non-continuous tasks, if the task is completed, a first positive reward value R1 is given; if a switch occurs, a first negative reward value R2 is given; if the task fails, a second negative reward value R3 is given. For consecutive tasks, a second positive reward value R4 is given if the task is completed, a third negative reward value R5 is given if a switch occurs, a fourth negative reward value R6 is given if the task fails, and a fifth negative reward value R7 is given if delayed access occurs; where R1 and R4 > 0; R2, R3, R5, R6, and R7 < 0.

7. The method according to claim 1, characterized in that, The step of iteratively optimizing and updating the parameters of the Actor neural network and the Critic neural network through the interaction between the agent and the environment to obtain the target strategy for outputting access and handover decisions includes: Through interaction between the agent and the environment, experiential data is collected and stored in the replay pool. Using this experiential data, the parameters of the Actor neural network and the Critic neural network are iteratively optimized and updated by calculating discount rewards and value estimates, constructing shearing importance sampling targets and entropy reward terms, so as to obtain the target strategy for outputting access and switching decisions.

8. The method according to claim 7, characterized in that, The interaction between the intelligent agent and the environment includes: The status information of the relay satellite is obtained and its features are encoded to obtain the relay satellite feature code; Obtain the status information of the task to be decided and subsequent tasks, and perform feature encoding to obtain the task feature code; The relay satellite feature encoding and the mission feature encoding are fused using a multi-head attention mechanism to obtain a fused feature vector; The fused feature vector is input into the Actor neural network to obtain the action probability distribution; and then input into the Critic neural network to obtain the state value estimate.

9. The method according to claim 7, characterized in that, The process of calculating discount returns and value estimates, constructing a shearing importance sampling objective, and an entropy reward term specifically includes: The discount return is calculated based on the following formula: G t =R (z) (t+1)+γG t+1 ; Where γ is the discount factor, R (z) (t+1) represents the reward value obtained by the agent after performing an action at the (t+1)th time step in the z rounds; The advantage estimate is calculated based on the following formula: A t =R (z) (t+1)+γ×value (z) (t+1)-value (z) (t); Where, value (z) (t) represents the state value at time step t in round z, estimated by the Critic network; The probability ratio between the old and new strategies is calculated using the following formula: r t =exp(logit (z) (t)-logπ θ (a t |S t )); Among them, logit (z) (t) represents the probability of the Actor network policy choosing an action based on the old parameters, π θ (a t |S t ) represents the Actor network policy composed of the current parameters in state S. t Choose action a t The probability of; The objective function for shearing is constructed as follows: Where ε is the shear parameter Entropy calculation reward items: Where β is the entropy coefficient.

10. A reinforcement learning-based relay satellite system access and handover planning system for implementing the method as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Intelligent resource joint scheduling method under environment uncertainty remote sensing satellite network

    CN112422171A

  • Satellite resource dynamic allocation method and system under cloud-native hybrid cloud architecture

    CN120582679A

  • Systems and methods for swarm intelligence in industrial robotics

    WO2025085841A1