Low-orbit large-scale constellation communication path planning method for burst interference

By splitting the files to be transferred into sub-files and building an agent, and dynamically adjusting the communication path of the low-orbit constellation based on the group generation of the agent and dynamically adjusting the low-orbit constellation, the problems of low-orbit transmission efficiency, insufficient security and inability to dynamically adjust the path in the existing technology are solved, and efficient and secure large-scale low-orbit satellite files are achieved.

CN120074626AActive Publication Date: 2025-05-30SICHUAN UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510117542.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-30
Estimated Expiration
2045-01-24

AI Technical Summary

Technical Problem

The existing low-orbit constellation communication path planning method is inefficient in the face of burst interference and large file transmission, and the communication path cannot be dynamically adjusted.

Method used

By splitting the files to be transferred into sub-files and building an agent, a policy network, an evaluation network, a target policy network and a target evaluation network are built based on the group of agents, a preliminary communication path is generated, and the path is dynamically adjusted according to the burst interference information.

Benefits of technology

It improves the efficiency and security of large file transmission, can dynamically adjust the communication path to deal with sudden interference, and improves the stability and efficiency of low-orbit satellite communication systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120074626A_ABST
    Figure CN120074626A_ABST
Patent Text Reader

Abstract

The invention discloses a low-orbit large-scale constellation communication path planning method for burst interference. The method comprises the following steps: splitting a to-be-transmitted file to obtain to-be-transmitted sub-files, and constructing an intelligent agent according to the to-be-transmitted sub-files; constructing a strategy network, an evaluation network, a target strategy network and a target evaluation network of the intelligent agent based on the intelligent agent group, and generating a preliminary low-orbit large-scale constellation communication path according to the strategy network, the evaluation network, the target strategy network and the target evaluation network of the intelligent agent; and acquiring burst interference information, and dynamically adjusting the initial low-orbit large-scale constellation communication path based on the burst interference information so as to acquire the burst interference-oriented low-orbit large-scale constellation communication path. According to the method, the large file transmission efficiency is improved, the possibility that information is comprehensively intercepted is reduced, sudden interference can be dealt with, and the communication path can be dynamically adjusted or reselected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication path planning, and particularly to a communication path planning method for low-earth orbit large-scale constellations facing burst interference. Background Art

[0002] Currently, the number of low-earth orbit satellites has begun to gradually increase, but existing methods are still lacking for ultra-long-distance transmission of large files using low-earth orbit satellites. At present, an existing method proposes an optimization of a communication path planning method for low-earth orbit constellations based on the Dijkstra algorithm. This method transforms the problem of finding the optimal communication path into the problem of finding the shortest weighted path. Specifically, the low-earth orbit satellite network is abstracted into a weighted graph, and then the Dijkstra algorithm is used to plan the shortest path for each task, which has a certain effect on the planning of communication paths. However, this method has the following disadvantages:

[0003] (1) This method transmits each task on the same communication path. When the task contains a large file, the transmission efficiency is slow; at the same time, since all the information of a task is on the same communication path, it will face the risk of being completely intercepted, and the security is low;

[0004] (2) This method is based on the shortest path weighting method and only considers distance, load, and task priority. In the face of burst interference, it cannot dynamically adjust the communication path.

[0005] Therefore, there is an urgent need in this field for a communication path planning method that can transmit large files on low-earth orbit satellites and solve transmission efficiency and security problems in the face of burst interference. Summary of the Invention

[0006] In view of the above deficiencies in the prior art, the present invention provides a communication path planning method for low-earth orbit large-scale constellations facing burst interference.

[0007] To achieve the above invention object, the technical solution adopted by the present invention is:

[0008] A communication path planning method for low-earth orbit large-scale constellations facing burst interference, comprising the following steps:

[0009] S1. Split the file to be transmitted to obtain sub-files to be transmitted, and construct agents according to the sub-files to be transmitted;

[0010] S2. Based on the agent group, construct a policy network, an evaluation network, a target policy network, and a target evaluation network of the agents, and generate a preliminary communication path for low-earth orbit large-scale constellations according to the policy network, evaluation network, target policy network, and target evaluation network of the agents;

[0011] S3. Obtain the burst interference information, and dynamically adjust the preliminary low-Earth orbit large-scale constellation communication path based on the burst interference information to obtain a low-Earth orbit large-scale constellation communication path facing burst interference.

[0012] Further, in step S1, the basic information carried by the agent includes position coordinate information, speed information, and environmental information.

[0013] Further, in step S2, generate the preliminary low-Earth orbit large-scale constellation communication path according to the agent's policy network, evaluation network, target policy network, and target evaluation network, including the following steps:

[0014] A1. Randomly initialize the network parameters respectively, and construct an experience pool;

[0015] A2. Obtain the state information of the agent to form the observation information of the agent at the current time step;

[0016] A3. Based on the observation information of the agent at the current time step, use the agent's policy network to generate the execution action of the agent at the current time step, and obtain the reward of the agent at the current time step and the observation information of the agent at the next time step according to the execution action of the agent at the current time step;

[0017] A4. Store the observation information, execution action, reward, and observation information at the next time step of the agent at the current time step into the experience pool;

[0018] A5. Repeat steps A2 - A4 until the experience pool is completely stored, and randomly sample the experience pool to obtain a sample set;

[0019] A6. Based on the sample set, use the evaluation network of the agent to calculate the reward estimation value of the agent, and use the target policy network and target evaluation network of the agent to calculate the reward target value of the agent;

[0020] A7. Based on the observation information, execution action, reward estimation value, and reward target value of the agent, update the policy network parameters, evaluation network parameters, target policy network parameters, and target evaluation network parameters of the agent until the set update threshold condition is met to obtain the final policy network of the agent;

[0021] A8. Use the final policy network of the agent to generate the preliminary low-Earth orbit large-scale constellation communication path.

[0022] Further, in step A7, calculate the loss function of the evaluation network of the agent based on the reward estimation value and reward target value of the agent to update the evaluation network parameters of the agent, expressed as:

[0023]

[0024] where: \(L(\theta\) t+1,m,critic ) is the loss function of the evaluation network of the \(m\)-th agent at the next time step \(t + 1\), \(\theta\) t+1,m,critic is the network parameter of the evaluation network of the \(m\)-th agent at the next time step \(t + 1\), \(E[\cdot]\) is the expectation operator, is the reward estimate of the evaluation network of the \(m\)-th agent at the current time step \(t\), \(r\) t,m is the reward of the \(m\)-th agent at the current time step \(t\), \(\gamma\) is the discount factor, is the reward target value of the target evaluation network of the \(m\)-th agent at the next time step \(t + 1\).

[0025] Furthermore, in step A7, based on the observation information, executed actions, and reward estimates of the agent, the policy network parameters of the agent are updated, including the following steps:

[0026] B1. Calculate the policy gradient of the agent based on the observation information, executed actions, and reward estimates of the agent, expressed as:

[0027]

[0028] where: is the policy gradient of the policy \(\mu\) of the \(m\)-th agent with respect to the policy network parameter \(\theta\) t,m,Actor , \(E[\cdot]\) is the expectation operator, \(o, a \sim D\) are the observation information \(o\) and executed action \(a\) sampled from the experience pool \(D\), is the policy gradient of the policy \(\mu(a\) t,m | \(o\) t,m ) with respect to the policy network parameter \(\theta\) t,m,Actor , is the reward estimate of the \(m\)-th agent with respect to the executed action \(a\) t,m , \(J(\mu)\) is the objective function of the agent, the expectation of the agent's long-term cumulative return;

[0029] B2. Update the policy network parameters of the agent according to the policy gradient of the agent, expressed as:

[0030]

[0031] where: \(\theta\) t+1,m,Actor is the network parameter of the policy network of the \(m\)-th agent at the next time step \(t + 1\), \(\theta\) t,m,Actor is the network parameter of the policy network of the \(m\)-th agent at the current time step \(t\), \(\beta\) is the learning rate.

[0032] Further, in step A7, based on the agent's observation information, executed actions, reward estimate values, and reward target values, update the target policy network parameters and target evaluation network parameters of the agent, expressed as:

[0033] θ Target ← τθ + (1 - τ)θ Target

[0034] Where: θ Target is the target network parameter of the agent, θ is the network parameter of the agent, and τ is the soft update parameter.

[0035] Further, in step S3, dynamically adjust the preliminary low-Earth orbit large-scale constellation communication path based on the burst interference information to obtain a low-Earth orbit large-scale constellation communication path for burst interference, including the following steps:

[0036] C1. Calculate the reward expectation of the agent on the preliminary low-Earth orbit large-scale constellation communication path based on the burst interference information and the preliminary low-Earth orbit large-scale constellation communication path;

[0037] C2. Calculate the overall reward of the agent on the preliminary low-Earth orbit large-scale constellation communication path based on the reward expectation of the agent on the preliminary low-Earth orbit large-scale constellation communication path;

[0038] C3. Sort the overall rewards of the agent on the preliminary low-Earth orbit large-scale constellation communication path to obtain a dynamic adjustment priority sequence of the communication path, and use the dynamic adjustment priority sequence of the communication path to dynamically adjust the preliminary low-Earth orbit large-scale constellation communication path to obtain a low-Earth orbit large-scale constellation communication path for burst interference.

[0039] Further, in step C1, calculate the reward expectation of the agent on the preliminary low-Earth orbit large-scale constellation communication path based on the burst interference information and the preliminary low-Earth orbit large-scale constellation communication path, expressed as:

[0040] E m (R m ) = ∑P m R m

[0041] Where: E m (R m ) is the reward expectation of the m-th agent on the preliminary low-Earth orbit large-scale constellation communication path, P m is the probability that the m-th agent is under burst interference information, and R m is the transmission reward of the m-th agent under burst interference information.

[0042] Further, in step C2, based on the reward expectation of the agent on the preliminary low-Earth orbit large-scale constellation communication path, the overall reward of the agent on the preliminary low-Earth orbit large-scale constellation communication path is calculated, expressed as:

[0043]

[0044] Where: Q is the overall reward of the m-th agent on the preliminary low-Earth orbit large-scale constellation communication path, M is the total number of agents, and μ m is the weight factor of the m-th agent relative to the overall reward.

[0045] The present invention has the following beneficial effects:

[0046] (1) By splitting the file to be transmitted to obtain sub-files to be transmitted and constructing agents based on the sub-files to be transmitted, the present invention avoids the problem of slow transmission efficiency caused by placing the file to be transmitted on the same communication path when the file to be transmitted is large, and at the same time avoids the risk of complete interception of information due to all the information of the file to be transmitted being on the same communication path, thereby improving the security of information transmission;

[0047] (2) By constructing a policy network, an evaluation network, a target policy network, and a target evaluation network of the agent based on the agent group, and generating a preliminary low-Earth orbit large-scale constellation communication path according to the policy network, evaluation network, target policy network, and target evaluation network of the agent; then obtaining burst interference information and dynamically adjusting the preliminary low-Earth orbit large-scale constellation communication path based on the burst interference information to obtain a low-Earth orbit large-scale constellation communication path facing burst interference, the present invention avoids the problem that the shortest path weighting method only considers distance, load, and task priority and cannot dynamically adjust the communication path in the face of burst interference. Description of the Drawings

[0048] Figure 1 is a schematic flow chart of a method for planning a low-Earth orbit large-scale constellation communication path facing burst interference;

[0049] Figure 2 is a schematic diagram of the transmission path planning after splitting the file to be transmitted;

[0050] Figure 3 is a schematic diagram of generating a preliminary low-Earth orbit large-scale constellation communication path using the 4 networks. Detailed Embodiments

[0051] The specific implementation manners of the present invention will be described below to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation manners. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0052] As Figure 1 shown, a communication path planning method for a low-earth orbit large-scale constellation facing burst interference includes steps S1 - S3, specifically as follows:

[0053] S1. Split the file to be transmitted to obtain sub-files to be transmitted, and construct agents according to the sub-files to be transmitted.

[0054] In an optional embodiment of the present invention, in large-scale low-earth orbit satellites, for a large file to be transmitted from satellite A to satellite B, it is evenly divided into M small sub-files to be transmitted, and then the present invention constructs each sub-file to be transmitted into an agent. The operation of each agent is decided on each satellite of the current transmission node. As Figure 2 shown, after the splitting is completed, the system starts a parallel transmission mechanism so that these sub-files to be transmitted can start transmission on different communication paths simultaneously. The basic information carried by the agent includes position coordinate information, speed information, and environmental information.

[0055] S2. Construct a policy network, an evaluation network, a target policy network, and a target evaluation network of the agent based on the agent group, and generate a preliminary low-earth orbit large-scale constellation communication path according to the policy network, evaluation network, target policy network, and target evaluation network of the agent.

[0056] In an optional embodiment of the present invention, the present invention constructs a policy network, an evaluation network, a target policy network, and a target evaluation network of the agent based on the agent group. The four networks respectively correspond to Actor, Critic, Target_Actor, and Target_Critice in MADDPG (Multi-Agent Deep Deterministic Policy Gradient).

[0057] As Figure 3 shown, the present invention generates a preliminary low-earth orbit large-scale constellation communication path according to the policy network (Actor), evaluation network (Critic), target policy network (Target_Actor), and target evaluation network (Target_Critice) of the agent, including the following steps:

[0058] A1. Randomly initialize the network parameters respectively and construct an experience pool.

[0059] Specifically, in the present invention, the policy network parameters θ of the agent Actor , the evaluation network θ critic , the target policy network θ Target_actor and the target evaluation network θ Targer_critic are randomly initialized respectively, so that the network has a certain degree of randomness and diversity at the beginning of training, which helps the subsequent learning and optimization process. Construct an experience pool D and set the experience pool D to be empty.

[0060] A2. Obtain the state information of the agent to form the observation information of the agent at the current time step.

[0061] Specifically, in the present invention, each agent obtains the position coordinate information and speed information of itself and other agents, and then jointly constructs the observation information of the agent, which is expressed as:

[0062] (o t,1 , L, o t,m )

[0063] where: o t,1 is the observation information of the first agent at the current time step t, o t,m is the observation information of the mth agent at the current time step t, t is the current time step, and m is the agent number.

[0064] A3. Based on the observation information of the agent at the current time step, use the policy network of the agent to generate the execution action of the agent at the current time step, and obtain the reward of the agent at the current time step and the observation information of the agent at the next time step according to the execution action of the agent at the current time step.

[0065] Specifically, in the present invention, based on the observation information of the agent at the current time step, the policy network of each agent outputs an action respectively, that is, selects the next satellite node. Each agent will enter a new state, and then the reward of the agent at the current time step and the observation information of the agent at the next time step can be obtained.

[0066] A4. Store the observation information, execution action, reward of the agent at the current time step, and the observation information at the next time step into the experience pool.

[0067] A5. Repeat steps A2 - A4 until the experience pool is fully stored, and randomly sample the experience pool to obtain a sample set.

[0068] A6. Based on the sample set, use the evaluation network of the agent to calculate the reward estimation value of the agent, and use the target policy network and target evaluation network of the agent to calculate the reward target value of the agent.

[0069] The evaluation network of the agent calculates the reward estimation value of each agent according to the observation information, action information and parameter θ of all agents in the sample set k at the current time step. t,m,critic Calculate the reward estimation value of each agent. Among them, μ represents the determined policy, which means that the Q value is calculated based on a certain determined policy. At the same time, the target evaluation network of the agent receives the next state (o t+1,1 , L, o t+1,m ), and uses the next action output by the target policy network of the agent (a t+1,1 , L, a t+1,m ) to calculate the reward target value of each agent.

[0070] A7. Based on the observation information, executed action, reward estimation value and reward target value of the agent, update the policy network parameters, evaluation network parameters, target policy network parameters and target evaluation network parameters of the agent until the set update threshold condition is met to obtain the final policy network of the agent.

[0071] The present invention calculates the loss function of the evaluation network of the agent based on the reward estimation value and reward target value of the agent to update the evaluation network parameters of the agent, which is expressed as:

[0072]

[0073] Among them: L(θ t+1,m,critic ) is the loss function of the evaluation network of the m-th agent at the next time step t + 1, and θ t+1,m,critic is the network parameter of the evaluation network of the m-th agent at the next time step t + 1. E[·] is the expectation operator. is the reward estimation value of the evaluation network of the m-th agent at the current time step t, and r t,m is the reward of the m-th agent at the current time step t. γ is the discount factor. is the reward target value of the target evaluation network of the m-th agent at the next time step t + 1.

[0074] The present invention updates the policy network parameters of the agent based on the observation information, executed action and reward estimation value of the agent, including the following steps:

[0075] B1. Calculate the policy gradient of the agent based on the observation information, executed action and reward estimation value of the agent, which is expressed as:

[0076]

[0077] Among them: is the policy gradient of the policy μ of the m-th agent with respect to the policy network parameter θ t,m,Actor , E[·] is the expectation operator, o,a:D are the observed information o and the executed action a sampled from the experience pool D, is the policy μ(a t,m |o t,m ) with respect to the policy network parameter θ t,m,Actor of the policy gradient, is the reward estimate value of the m-th agent with respect to the executed action a t,m of the policy gradient, J(μ) is the objective function of the agent, the expectation of the long-term cumulative return of the agent.

[0078] B2. Update the policy network parameters of the agent according to the policy gradient of the agent, which is expressed as:

[0079]

[0080] where: θ t+1,m,Actor is the network parameter of the policy network of the m-th agent at the next time step t+1, θ t,m,Actor is the network parameter of the policy network of the m-th agent at the current time step t, and β is the learning rate.

[0081] The present invention updates the target policy network parameter and the target evaluation network parameter of the agent based on the observed information, executed action, reward estimate value and reward target value of the agent, which is expressed as:

[0082] θ Target ←τθ+(1-τ)θ Target

[0083] where: θ Target is the target network parameter of the agent, θ is the network parameter of the agent, and τ is the soft update parameter.

[0084] Specifically, τ in the present invention is a small soft update parameter. In this way, the target policy network parameter and the target evaluation network parameter of the agent will slowly approach the parameters of the current evaluation network and policy network respectively, making its update relatively stable and avoiding excessive fluctuations in the training process. Then the present invention optimizes the network parameters of the four networks until the set update threshold condition is met to obtain the final policy network of the agent.

[0085] A8. Generate a preliminary low-earth-orbit large-scale constellation communication path using the final policy network of the agent.

[0086] S3. Obtain the burst interference information, and dynamically adjust the preliminary low-Earth orbit large-scale constellation communication path based on the burst interference information to obtain a low-Earth orbit large-scale constellation communication path facing burst interference.

[0087] In an alternative embodiment of the present invention, the process of obtaining the burst interference information is as follows: The agent generates information about burst interference based on the preliminary low-Earth orbit large-scale constellation communication path and environmental information to obtain the burst interference information. The present invention dynamically adjusts the preliminary low-Earth orbit large-scale constellation communication path based on the burst interference information to obtain a low-Earth orbit large-scale constellation communication path facing burst interference, including the following steps:

[0088] C1. Based on the burst interference information and the preliminary low-Earth orbit large-scale constellation communication path, calculate the reward expectation of the agent on the preliminary low-Earth orbit large-scale constellation communication path, expressed as:

[0089] E m (R m ) = ∑P m R m

[0090] Where: E m (R m ) is the reward expectation of the m-th agent on the preliminary low-Earth orbit large-scale constellation communication path, P m is the probability that the m-th agent is under the burst interference information, and R m is the transmission reward of the m-th agent under the burst interference information.

[0091] C2. Based on the reward expectation of the agent on the preliminary low-Earth orbit large-scale constellation communication path, calculate the overall reward of the agent on the preliminary low-Earth orbit large-scale constellation communication path, expressed as:

[0092]

[0093] Where: Q is the overall reward of the m-th agent on the preliminary low-Earth orbit large-scale constellation communication path, M is the total number of agents, and μ m is the weight factor of the m-th agent relative to the overall reward.

[0094] C3. Sort the overall rewards of the agents on the preliminary low-Earth orbit large-scale constellation communication path to obtain a dynamic adjustment priority sequence of the communication path, and use the dynamic adjustment priority sequence of the communication path to dynamically adjust the preliminary low-Earth orbit large-scale constellation communication path to obtain a low-Earth orbit large-scale constellation communication path facing burst interference.

[0095] Specifically, the present invention selects the communication path corresponding to the maximum Q value by comparing the Q values of different paths in the dynamic adjustment priority sequence of the communication path, obtains the low-earth-orbit large-scale constellation communication path for burst interference, so that the intelligent agent can dynamically respond to various burst interferences, continuously adjust and optimize the communication path, and achieve the efficient and stable operation of the low-earth-orbit satellite large-file transmission system.

[0096] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0097] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0098] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0099] Specific embodiments of the present invention are used to elaborate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

[0100] Those of ordinary skill in the art will realize that the embodiments described herein are provided to assist the reader in understanding the principles of the present invention, and it should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the scope of protection of the present invention.

Claims

1. A method for low-orbit large-scale constellation communication path planning for sudden interference, characterized in that: The following steps are involved: S1, split the file to be transmitted to obtain sub-files to be transmitted, and construct an agent according to the sub-files to be transmitted; S2. Building a strategy network, evaluation network, target strategy network and target evaluation network of the intelligent agent based on the intelligent agent group, and generating a preliminary low-orbit large-scale constellation communication path according to the strategy network, evaluation network, target strategy network and target evaluation network of the intelligent agent; S3. Obtain burst interference information, and dynamically adjust the preliminary low-orbit large-scale constellation communication path based on the burst interference information to obtain a low-orbit large-scale constellation communication path resistant to burst interference.

2. The method for low-orbit large-scale constellation communication path planning for sudden interference according to claim 1, characterized in that: In step S1, the basic information carried by the agent includes position coordinate information, speed information and environment information.

3. The method for low-orbit large-scale constellation communication path planning for sudden interference according to claim 1, characterized in that: In step S2, a preliminary low-orbit large-scale constellation communication path is generated according to the agent's policy network, evaluation network, target policy network, and target evaluation network, including the following steps: A1. Randomly initialize the network parameters and build an experience pool; A2. Obtain the state information of the agent to form the observation information of the agent at the current time step; A3. Based on the observation information of the agent at the current time step, the agent's policy network is used to generate the agent's execution action at the current time step, and the agent's reward at the current time step and the agent's observation information at the next time step are obtained according to the agent's execution action at the current time step; A4. Store the agent's observation information, execution actions, rewards at the current time step, and observation information at the next time step into the experience pool; A5. Repeat steps A2-A4 until the experience pool is fully stored, and randomly sample the experience pool to obtain a sample set; A6. Based on the sample set, the agent's evaluation network is used to calculate the agent's reward estimate, and the agent's target strategy network and target evaluation network are used to calculate the agent's reward target value. A7. Based on the agent's observation information, execution actions, reward estimation values, and reward target values, the agent's policy network parameters, evaluation network parameters, target policy network parameters, and target evaluation network parameters are updated until the set update threshold conditions are met to obtain the agent's final policy network. A8. Use the agent’s final strategy network to generate preliminary low-orbit large-scale constellation communication paths.

4. The method for low-orbit large-scale constellation communication path planning for sudden interference according to claim 3 is characterized in that: In step A7, the loss function of the evaluation network of the agent is calculated based on the estimated reward value and the target reward value of the agent to update the evaluation network parameters of the agent, which is expressed as: Where: L(θ t+1,m,critic ) is the loss function of the evaluation network of the mth agent at the next time step t+1, θ t+1,m,critic is the network parameter of the evaluation network of the mth agent at the next time step t+1, E[·] is the expectation operator, is the reward estimate of the evaluation network of the mth agent at the current time step t, r t,m is the reward of the mth agent at the current time step t, γ is the discount factor, is the reward target value of the target evaluation network of the mth agent at the next time step t+1.

5. The method for low-orbit large-scale constellation communication path planning for sudden interference according to claim 3, characterized in that: In step A7, based on the agent's observation information, execution actions, and reward estimates, the agent's policy network parameters are updated, including the following steps: B1. Based on the agent's observation information, execution actions and reward estimates, the agent's policy gradient is calculated, expressed as: in: is the strategy μ of the mth agent with respect to the policy network parameter θ t,m,Actor The policy gradient of E[·] is the expected operator, o,a:D is the observation information o sampled from the experience pool D and the execution action a, is the strategy μ(a) of the mth agent t,m ∣o t,m ) About the policy network parameters θ t,m,Actor The policy gradient of is the estimated reward value of the mth agent |About executing action a t,m The policy gradient of , J(μ) is the objective function of the agent, and the expectation of the long-term cumulative return of the agent; B2. Update the policy network parameters of the agent according to the agent's policy gradient, expressed as: Where: θ t+1,m,Actor is the network parameter of the policy network of the mth agent at the next time step t+1, θ t,m,Actor is the network parameter of the policy network of the mth agent at the current time step t, and β is the learning rate.

6. The method for low-orbit large-scale constellation communication path planning facing sudden interference according to claim 3, characterized in that: In step A7, based on the agent's observation information, execution actions, reward estimation values, and reward target values, the agent's target strategy network parameters and target evaluation network parameters are updated, expressed as: i Target ←τθ+(1-τ)θ Target Where: θ Target is the target network parameter of the agent, θ is the network parameter of the agent, and τ is the soft update parameter.

7. The method for planning a low-orbit large-scale constellation communication path for sudden interference according to claim 1, characterized in that: In step S3, dynamically adjusting the preliminary LEO large-scale constellation communication path based on the burst interference information to obtain a LEO large-scale constellation communication path oriented to burst interference includes the following steps: C1. Calculate the expected reward of the agent on the preliminary low-orbit large-scale constellation communication path based on the burst interference information and the preliminary low-orbit large-scale constellation communication path; C2. Calculate the overall reward of the agent in the preliminary low-orbit large-scale constellation communication path based on the agent's expected reward in the preliminary low-orbit large-scale constellation communication path; C3. Sort the overall rewards of the intelligent agents in the preliminary low-orbit large-scale constellation communication path to obtain a dynamically adjusted priority sequence of the communication path, and dynamically adjust the preliminary low-orbit large-scale constellation communication path using the dynamically adjusted priority sequence of the communication path to obtain a low-orbit large-scale constellation communication path resistant to sudden interference.

8. The method for planning a low-orbit large-scale constellation communication path for sudden interference according to claim 7, characterized in that: In step C1, based on the burst interference information and the preliminary low-orbit large-scale constellation communication path, the reward expectation of the agent in the preliminary low-orbit large-scale constellation communication path is calculated, which is expressed as: E m (R m )=∑P m R m Where: E m (R m ) is the reward expectation of the mth agent under the initial low-orbit large-scale constellation communication path, P m is the probability that the mth agent is under sudden interference information, R m is the transmission reward of the mth agent under burst interference information.

9. The method for low-orbit large-scale constellation communication path planning facing sudden interference according to claim 7, characterized in that: In step C2, based on the expected reward of the agent in the preliminary low-orbit large-scale constellation communication path, the overall reward of the agent in the preliminary low-orbit large-scale constellation communication path is calculated, which is expressed as: Where: Q is the overall reward of the mth agent under the initial low-orbit large-scale constellation communication path, M is the total number of agents, μ m is the weight factor of the mth agent relative to the overall reward.

Citation Information

Patent Citations

  • Multi-agent autonomous navigation method based on reinforcement learning

    CN112132263A

  • Multi-agent-based collaborative transportation method and system thereof

    CN113724123A

  • Multi-agent path planning method based on deep reinforcement learning

    CN114815840A

  • Flow prediction satellite path selection method and system based on reinforcement learning

    CN116781139A

  • Multi-agent-based low-orbit satellite network routing decision-making method and device

    CN117614882A