A training method for a satellite routing prediction model and a low-earth orbit satellite routing method

By training satellite routing prediction models based on multi-objective reinforcement learning, the problem of traditional routing algorithms being difficult to adapt to dynamic link changes and being unable to balance delays and load balancing is solved, and efficient multi-objective optimization and differentiated service requirements are achieved.

CN119783554BActive Publication Date: 2025-06-20BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510272529.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-20
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

Traditional low-orbit satellite routing algorithms are difficult to adapt to highly dynamic inter-star link changes, and are usually optimized for only a single target, unable to balance delay and load balancing, and unable to meet differentiated business needs.

Method used

Using a low-orbit satellite routing method based on multi-objective reinforcement learning, the satellite routing prediction model is trained, and the observation information of each satellite node and the weight combination of the multi-objective optimization model is obtained, and the Q value of multiple actions is generated. The Mixing network is jointly trained, and the parameters of the routing prediction model are adjusted until the preset iteration conditions are met.

Benefits of technology

Multi-objective optimization of routing algorithms is realized, which can balance delay and load balancing, meet diversified business needs, and improve network performance and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783554B_ABST
    Figure CN119783554B_ABST
Patent Text Reader

Abstract

The present application provides a training method for a satellite routing prediction model and a low-earth orbit satellite routing method. The method includes obtaining the observation information of each satellite node in the current satellite network environment state and the weight combinations of each target randomly selected based on a multi-objective optimization model; inputting the observation information of each satellite node and the weight combinations of each randomly selected target into their respective satellite routing prediction models to obtain the Q values of multiple actions; jointly inputting the Q values output by each satellite routing prediction model into a Mixing network for joint training, calculating a loss value based on the main joint Q and the secondary joint Q values and a preset loss function, and adjusting the parameters of the satellite routing prediction model based on the loss value. The above training process is repeated until a preset iteration condition is reached. By using the weights of different routing metrics of the multi-objective optimization model as observations and inputting them into the satellite routing prediction model for training, the present invention can generate diverse strategies that meet different service preferences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and in particular, to a method for training a satellite routing prediction model and a low-earth orbit satellite routing method. Background Art

[0002] With the progress of technologies such as 5G communication, artificial intelligence, big data, and the Internet of Things, the scope of user requirements and communication services has been continuously growing, developing towards diversification and complexity. Business transmission requires high-performance metrics such as higher bandwidth and lower latency to support user needs. Traditional terrestrial networks cannot meet the above requirements due to limitations such as limited coverage and expensive equipment. However, the rapidly developing low-earth orbit satellite network in recent years provides new ideas for solving the above problems. The low-earth orbit satellite network is usually deployed in the sky at an altitude of 400 km - 2000 km above the ground. Therefore, compared with geostationary orbit satellites and high-orbit satellites, routing based on low-earth orbit satellites has lower latency and better communication quality.

[0003] The low-earth orbit satellite network usually consists of dozens or hundreds of periodically moving satellites. The load of satellite nodes changes rapidly, the inter-satellite link state is unstable, and the satellite network topology changes periodically. Traditional routing algorithms, such as the shortest path algorithm or the link state algorithm, usually have difficulty adapting to the highly dynamic changes of inter-satellite links. In addition, these methods usually optimize only for a single objective (such as minimizing latency or maximizing throughput), for example, they cannot balance latency and load balancing, resulting in the inability to meet differentiated service requirements. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a low-earth orbit satellite routing method based on multi-objective reinforcement learning to improve the efficiency of the routing algorithm and balance multiple routing metrics to meet diversified service requirements.

[0005] In a first aspect, a method for training a satellite routing prediction model is provided, which is applied to a low-earth orbit satellite routing scenario. Each satellite node corresponds to a satellite routing prediction model. The method includes:

[0006] Obtaining the observation information of each satellite node in the current satellite network environment state and the weight combination of each objective randomly selected based on a multi-objective optimization model; the observation information includes the current data queue length of each satellite node and the current data queue lengths of its neighbor nodes, as well as the distances from the neighbor nodes to the target satellite node; the multi-objective optimization model is a target optimization model pre-constructed based on each routing metric, and the routing metrics at least include routing transmission latency and routing load balancing;

[0007] Input the observation information of each satellite node and the weights of randomly selected targets into their respective satellite routing prediction models to obtain the Q-values of multiple actions; an action is to select the next hop of a data packet among neighbor nodes under the current input, and the Q-value is the expected return value obtained after selecting this action;

[0008] Jointly input the Q-values output by each satellite routing prediction model into the Mixing network for joint training, where the Mixing network consists of a main network and a secondary network; the main network outputs the main joint Q-value, and the secondary network outputs the secondary joint Q-value;

[0009] Calculate the loss value based on the main joint Q and secondary joint Q-values and a preset loss function;

[0010] Adjust the parameters of the satellite routing prediction model based on the loss value, and repeat the above training process until a preset iteration condition is reached.

[0011] Optionally, the method further includes:

[0012] Interact the action corresponding to the maximum Q-value among the Q-values output by each satellite routing prediction model with the current satellite network environment to generate the next satellite network environment state, and obtain the next observation information in the next satellite network environment state;

[0013] Calculate the reward value for executing the action with the maximum Q-value based on the reward function;

[0014] Construct a set of experience data from the current satellite network environment state, the current observation information, the reward value, the next satellite network environment state, the next observation information, and the weight combination;

[0015] Store the experience data in a pre-constructed experience replay pool;

[0016] When the amount of data in the experience replay pool reaches a preset threshold, randomly select several experience samples from the experience replay pool to train the Mixing network.

[0017] Optionally, the method further includes:

[0018] Establish a monotonic constraint condition for the Mixing network and each satellite routing prediction model, so that the action corresponding to the maximum Q-value of each routing prediction model is consistent with the action corresponding to the maximum joint Q-value output by the Mixing network, where the monotonic constraint condition is:

[0019]

[0020] Among them, the left side represents the best action selection of the overall Mixing network; the right side represents the best action selection of each routing prediction model; represents in the given state , and a randomly selected weight combination Under this condition, the selection action The expected value of; Indicates the number of routing prediction models.

[0021] Optionally, the method further includes:

[0022] At intervals of a preset number of training times, increase the probability of the unselected weight combination.

[0023] Optionally, the multi-objective optimization model is:

[0024]

[0025] Among them, , Both represent the routing delay index; , Both represent the routing load balancing index; Indicates the total number of data packets; Represents the local area formed by the current satellite node and its neighbor nodes; Represents a satellite node; Represents the th satellite node in the local area; Represents the average queue length of all satellite nodes in the local area.

[0026] Optionally, the method further includes:

[0027] During the training process, at intervals of a preset time duration, obtain the global state information of all satellite nodes;

[0028] Input the global state information into the super network to generate the parameters of the Mixing network.

[0029] In a second aspect, a low-earth orbit satellite routing method is provided, which is applied to the training method of any satellite routing prediction model in the first aspect. The method includes:

[0030] Obtain the observation information of the current satellite node; the current satellite node includes an initial satellite node and an intermediate satellite node. The initial satellite node is the satellite node that generates data packets; the intermediate satellite node is the node that forwards data packets; the observation information includes the current data queue length of the current satellite node and the current data queue lengths of its neighbor nodes, as well as the distance from the neighbor nodes to the target satellite node;

[0031] Input the observation information of the current satellite node into the satellite routing prediction model deployed in the current satellite node to obtain the Q values of multiple actions; the action is to select the next hop of the data packet among the neighbor nodes under the current observation information;

[0032] Select the next satellite node for data packet transmission based on the action corresponding to the maximum Q value;

[0033] Repeat the above routing process for the next satellite node until the target satellite node is reached.

[0034] In a third aspect, a low-earth orbit satellite routing device is provided, which is applied to the training method of any satellite routing prediction model in the first aspect. The device includes:

[0035] An acquisition unit for acquiring the observation information of the current satellite node; the current satellite node includes an initial satellite node and an intermediate satellite node. The initial satellite node is the satellite node that generates data packets; the intermediate satellite node is the node that forwards data packets; the observation information includes the current data queue length of the current satellite node and the current data queue lengths of its neighbor nodes, as well as the distances from the neighbor nodes to the target satellite node;

[0036] A prediction unit for inputting the observation information of the current satellite node into the satellite routing prediction model deployed in the current satellite node to obtain the Q values of multiple actions; the action is to select the next hop of the data packet among the neighbor nodes under the current observation information;

[0037] A selection unit for selecting the next satellite node for data packet transmission based on the action corresponding to the maximum Q value; and repeating the above routing process for the next satellite node until the target satellite node is reached.

[0038] In a fourth aspect, an electronic device is provided, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus;

[0039] The memory is used to store a computer program;

[0040] The processor is used to implement the method steps described in any one of the first aspect or the second aspect when executing the program stored in the memory.

[0041] In a fifth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method steps described in any one of the first aspect or the second aspect.

[0042] A training method for a satellite routing prediction model and a low-earth orbit satellite routing method provided by the present invention. The method obtains the observation information of each satellite node in the current satellite network environment state and the weight combinations of various objectives randomly selected based on a multi-objective optimization model; inputs the observation information of each satellite node and the randomly selected weight combinations of various objectives into their respective satellite routing prediction models to obtain the Q values of multiple actions; jointly inputs the Q values output by each satellite routing prediction model into a Mixing network for joint training, calculates the loss value based on the main joint Q and the secondary joint Q values and a preset loss function, and adjusts the parameters of the satellite routing prediction model based on the loss value. Repeat the above training process until the preset iteration condition is reached. By inputting the weights of different routing metrics of the multi-objective optimization model as observations into the satellite routing prediction model for training, the present invention makes the weight sampling uniform, can generate diversified strategies that meet different service preferences, and performs centralized training based on the Q values of the entire satellite network through the Mixing network, rather than being limited to local information, to meet the highly dynamic inter-satellite link changes, making the routing decision more intelligent and flexible, being able to quickly respond to changes in the network state, and improving the overall network performance and reliability.

[0043] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given in conjunction with the accompanying drawings and described in detail as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 Shows the flowchart of the training method for the satellite routing prediction model provided by the embodiments of the present invention;

[0046] Figure 2 Shows the schematic structural diagram of the low-earth orbit satellite network and its training and execution provided by the embodiments of the present invention;

[0047] Figure 3 Shows the flowchart of a low-earth orbit satellite routing method provided by the embodiments of the present invention;

[0048] Figure 4 Shows the schematic structural diagram of a low-earth orbit satellite routing device provided by the embodiments of the present invention;

[0049] Figure 5 Shows the schematic structural diagram of an electronic device provided by the embodiments of the present invention. Specific Embodiments

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0051] Considering that a low-Earth orbit satellite network usually consists of dozens or hundreds of periodically moving satellites, the load of satellite nodes changes rapidly, the inter-satellite link state is unstable, and the satellite network topology changes periodically. Traditional routing algorithms, such as the shortest path algorithm or the link state algorithm, usually have difficulty adapting to the highly dynamic changes of inter-satellite links. In addition, these methods usually only optimize for a single objective (such as minimizing latency or maximizing throughput), for example, they cannot balance latency and load balancing, resulting in the inability to meet differentiated service requirements.

[0052] Based on this, the embodiments of the present invention provide a training method for a satellite routing prediction model, a low-Earth orbit satellite routing method and device based on multi-objective reinforcement learning, which will be described below through embodiments.

[0053] The embodiments of the present invention provide a training method for a satellite routing prediction model, which is applied to a low-Earth orbit satellite routing scenario. Each satellite node corresponds to a satellite routing prediction model, as Figure 1 shown. The method includes:

[0054] Step S101: Obtain the observation information of each satellite node in the current satellite network environment state and the weight combination of each objective randomly selected based on the multi-objective optimization model.

[0055] Among them, the observation information includes the current data queue length of each satellite node, the current data queue length of its neighbor nodes, and the distance from the neighbor nodes to the target satellite node; the multi-objective optimization model is a target optimization model pre-constructed based on various routing metrics, and the routing metrics at least include routing transmission delay and routing load balancing.

[0056] As Figure 2 shown, a constructed low-Earth orbit satellite network system divides the low-Earth orbit satellite constellation into orbital planes, and on each orbital plane, satellites are evenly distributed. The angle between adjacent orbital planes is , and the angle between adjacent satellites in the same orbit is . In polar regions, the relative motion speed between satellites in different orbits is relatively fast. In polar regions, it is difficult for the antenna system of satellites to track the positions of neighboring satellites. Therefore, regions with dimensions exceeding and are defined as polar regions. Considering communication costs and technical factors, inter-orbit satellite links are not established within polar regions.

[0057] In this low-Earth orbit satellite network system, each satellite node can generate and forward data packets. There are limited computing and storage resources on low-Earth orbit satellites. During peak traffic periods, data packets need to be stored in the local cache queue for waiting to be processed. Data packets follow the first-in, first-out principle. Similarly, the cache queue processes data packets according to the first-come, first-served principle. The cache queue reduces a certain amount of data packet loss but introduces queuing delay. In data packet transmission, it is necessary to wait for the queue to queue up. Therefore, the embodiments of the present invention establish a data packet queuing model, where the first represents the arrival process of data packets; it follows a Poisson distribution, the arrival interval time follows an exponential distribution, and the average arrival rate is ; the second represents the service time required for data packet forwarding, which follows an exponential distribution, and the average forwarding rate is ; 1 represents; is the maximum capacity of the data packet queue; represents that the number of data packet sources is infinite. When there are already N data packets in the queue at a certain moment, newly arrived data packets will be rejected from entering the queue.

[0058] In the prior art, in data packet transmission, when selecting the optimal next-hop satellite node, generally only a single metric is considered, such as only considering minimizing data transmission delay or only considering maximizing throughput, which limits its ability to meet the diverse quality of service requirements of future satellite Internet.

[0059] Therefore, the embodiments of the present invention construct a data packet delay model and a load balancing model;

[0060] First, regarding the construction process of the data packet delay model: Considering that the delay of a data packet consists of propagation delay and queuing delay. Therefore, the total delay of a data packet from one satellite node to another satellite node is:

[0061] (1);

[0062] Among them, the propagation delay is:

[0063] (2);

[0064] Wherein, represents the size of the forwarded data packet; represents the satellite node The data packet transmission rate between; can be expressed as:

[0065] (3);

[0066] Wherein, is the bandwidth of the inter-satellite link of the satellite node within the time slot t; is the signal-to-noise ratio of the inter-satellite link at this time, and can be expressed as:

[0067] (4);

[0068] In the above formula, is the transmission power of the satellite ; is the satellite The transmitting antenna gain of; is the satellite The receiving antenna gain of; is the noise power of the inter-satellite link, and (5). Wherein is the Boltzmann constant, is the satellite The noise temperature at the receiving end of; is the loss caused by distance attenuation during the propagation of the signal in free space, and can be expressed as follows:

[0069] (6);

[0070] Wherein, represents the distance between the satellite nodes ; is the carrier frequency of the signal.

[0071] In another feasible implementation, its queuing delay can be expressed as:

[0072] (7);

[0073] Wherein, represents the effective arrival rate of the data packets arriving at the satellite node per unit time; represents the average forwarding rate; represents the satellite node The average number of queued data packets of; can be expressed as:

[0074] (8);

[0075] Among them, is a constant; represents the maximum capacity of the data packet queue.

[0076] In a low-earth orbit satellite network, a source node generates data packets and transmits them to a destination satellite node d through continuous forwarding by intermediate satellite nodes. Therefore, data packets need to be transmitted through multiple hops by relay nodes. Then the path of data packet k from source node satellite s to destination satellite d can be expressed as , then the total inter-satellite routing delay of data packet pkt is: (9).

[0077] Secondly, considering the traffic balance problem in the network, when a satellite node forwards data packets, it needs to consider the load conditions of neighbor nodes and try to avoid passing through high-load nodes to alleviate the traffic congestion problem in the network. Therefore, the load balance model can be expressed as:

[0078] (10);

[0079] Among them, zone is the local area composed of the current satellite node and neighbor nodes; represents a satellite node; represents the current queue length of the th satellite node in this local area; represents the average queue length of all satellite nodes in this local area.

[0080] Finally, a multi-objective optimization model is constructed according to the above two models as:

[0081] (11);

[0082] Among them, , both represent the routing delay index; , both represent the routing load balance index; represents the total number of data packets; represents the local area composed of the current satellite node and neighbor nodes; represents a satellite node; represents the current queue length of the th satellite node in this local area; represents the average queue length of all satellite nodes in this local area.

[0083] In the embodiments of the present invention, by optimizing two metrics, namely the delay metric and the load balancing metric, the requirements of different services can be met.

[0084] Step S102: Input the observation information of each satellite node and the weight combinations of randomly selected targets into their respective satellite routing prediction models to obtain the Q-values of multiple actions.

[0085] Among them, the action is to select the next hop of the data packet among the neighbor nodes under the current input, and the Q-value is the expected return value obtained after selecting this action.

[0086] During training, the satellite nodes are regarded as agents, and each agent has an Agent network, which is the satellite routing prediction model.

[0087] Specifically, each agent makes decisions through the state-action value function . The decision-making can be expressed as follows:

[0088] (12);

[0089] In the above formula, select an initial value. For example, 0.1 is a common starting point, which means that there is about a 90% probability of selecting the action corresponding to the maximum Q-value, and a 10% probability of randomly exploring in the action space. Gradually decrease the value over time. This can ensure that the algorithm has sufficient exploration opportunities in the initial stage and relies more on the knowledge already obtained in the later stage. This mechanism enables the agent to continuously explore unknown areas to find potential better solutions while effectively using the known information to obtain immediate benefits, thereby continuously improving its behavior strategy during the learning process.

[0090] The weight combinations of randomly selected targets are the weight combinations of the two metrics of routing delay and load balancing, which can quantify the importance of different targets according to different tasks or service levels. By randomly selecting this weight combination as the input of the Agent network, these weights can be used to adjust the priorities when the agent makes decisions.

[0091] Step S103: Jointly input the Q-values output by each satellite routing prediction model into the Mixing network for joint training.

[0092] Among them, the Mixing network consists of a main network and a secondary network; the main network outputs the main joint Q-value, and the secondary network outputs the secondary joint Q-value.

[0093] Exemplarily, the primary network and the secondary network include two sets of networks, each set of networks corresponding to an optimization objective. Specifically, one set of networks corresponds to routing delay, and one set of networks corresponds to load balancing. The input is the Q vectors of n Agent networks, that is, , the Q values corresponding to the same objective are combined and input into the corresponding set of networks, and finally the outputs of the two sets of networks are integrated to output the multi-objective joint Q value, that is, the vector composed of m objectives . In the embodiment of the present invention, m is 2.

[0094] Step S104: Calculate the loss value based on the primary joint Q and the secondary joint Q values and a preset loss function; adjust the parameters of the satellite routing prediction model based on the loss value, and repeat the above training process until a preset iteration condition is reached.

[0095] In an example, the preset iteration condition is, for example, reaching a preset number of iterations or the loss value gradually converging to a local minimum.

[0096] In the embodiment of the present invention, since the agents are homogeneous, in order to promote cooperation among the agents and reduce the overhead of the neural network model, a parameter sharing mechanism is adopted, that is, all Agent networks share a set of parameters.

[0097] In an example, the loss function is as follows:

[0098] (13);

[0099] Wherein, represents the output of the secondary network, represents the primary joint Q value output by the primary network; are the parameters of the Mixing network; are the parameters of the Agent network; represents the weight vector space of the optimization objective; represents the experience pool data; the objective of the loss function is to minimize the difference between the output of the primary network and the output of the target network to update the Agent network parameters.

[0100] Specifically, can be expressed as follows:

[0101] (14);

[0102] Wherein, is the immediate reward; is the discount factor; represents the action space; represents the weight vector space of the optimization objective; represents the parameters of the secondary network; Represents the combined Q-value of the secondary network output.

[0103] In the embodiments of the present invention, all Agent networks are trained simultaneously, adopting centralized training. During centralized training, satellite nodes act as virtual agents to obtain observation information, and the global state is used as auxiliary information for training, so as to obtain the optimal routing strategy. It can not only reduce the training time and improve the training efficiency, but also consider the global information to obtain a better routing strategy.

[0104] In the above process of updating network parameters, if the parameter update frequency is too frequent, it will lead to instability of the model and difficulty in convergence. Therefore, the parameters of the secondary network are not updated every time of training, but are updated regularly after a certain number of training rounds, thus ensuring the stability during the training process.

[0105] In the present invention, the weights of different routing metrics of the multi-objective optimization model are used as observations and input into the satellite routing prediction model for training, so that the weight sampling is uniform, and diverse strategies that meet different service preferences can be generated. And through the Mixing network, centralized training is carried out based on the Q-values of the entire satellite network, rather than being limited to local information, which can meet the highly dynamic inter-satellite link changes, making the routing decision more intelligent and flexible, able to quickly respond to the changes of the network state, and improve the overall network performance and reliability.

[0106] Based on the above embodiments, the method further includes:

[0107] Interact the action corresponding to the maximum Q-value among the Q-values output by each satellite routing prediction model with the current satellite network environment to generate the next satellite network environment state, and obtain the next observation information in the next satellite network environment state;

[0108] Calculate the reward value of executing the action with the maximum Q-value based on the reward function;

[0109] Construct a set of experience data by combining the current satellite network environment state, the current observation information, the reward value, the next satellite network environment state, the next observation information and the weights;

[0110] Store the experience data in a pre-constructed experience replay pool;

[0111] When the amount of data in the experience replay pool reaches a preset threshold, randomly extract several experience samples from the experience replay pool to train the Mixing network.

[0112] After the agent interacts with the environment, record the state, observation, reward, action, and weight of each step, that is . When the number of experiences is greater than batch, random sampling training is performed, and the number of experiences sampled each time is batch. When the capacity of the experience replay pool is greater than the threshold, old experience data is deleted to ensure the effectiveness of the experience data.

[0113] Based on the above embodiments, the method further includes:

[0114] Establish monotonic constraint conditions for the Mixing network and each satellite routing prediction model, so that the action corresponding to the maximum Q value of each routing prediction model is consistent with the action corresponding to the maximum joint Q value output by the Mixing network, where the monotonic constraint condition is:

[0115] (15);

[0116] Among them, the left side represents the best action selection of the overall Mixing network; the right side represents the best action selection of each routing prediction model; represents the given state , and the randomly selected weight combination under which, the expected value of selecting the action ; represents the number of routing prediction models.

[0117] Based on the above embodiments, the method further includes:

[0118] At intervals of a preset number of training times, increase the probability of the weight combinations that have not been selected.

[0119] During training, randomly sampling weights may lead to uneven solution sets. Therefore, at intervals of a fixed number of training rounds, check the density and uniformity of the solution sets, and increase the probability of the weights that have not been selected, so that the weights that have not been selected are emphasized in the subsequent training process. Exemplarily, for example, in a multi-objective optimization model, the weight of routing delay is 0.1 and the weight of load balancing is 0.9. Then, if the weight combination is such that the weight of delay is less than the weight of load balancing, it will lead to uneven selection of actions. Then, the probability of the weights that have not been selected can be increased. Then, the weight of routing delay will probably become 0.6 and the weight of load balancing will be 0.4. Finally, a set of dense and uniform disposable solution sets are obtained, which meet the weight preferences of different services.

[0120] Based on the above embodiments, the method further includes:

[0121] During the training process, at intervals of a preset duration, obtain the global state information of all satellite nodes;

[0122] Input the global state information into the hypernetwork to generate the parameters of the Mixing network.

[0123] During the centralized training, the satellite nodes act as virtual agents to obtain observation information, and the global state is used as auxiliary information for training, so as to obtain the optimal routing strategy.

[0124] After the above centralized training is completed, the Agent network is deployed on the satellite. By inputting the local observation, the satellite can obtain its own state-action function , and select the action with the largest value as the next hop. Such distributed decision-making has real-time performance and avoids unnecessary communication overhead.

[0125] Based on the above embodiments, an embodiment of the present invention provides a low-earth orbit satellite routing method, which is applied to the training method of the above satellite routing prediction model, as Figure 3 shown, the method includes:

[0126] Step S301: Obtain the observation information of the current satellite node.

[0127] The current satellite node includes an initial satellite node and an intermediate satellite node. The initial satellite node is the satellite node that generates data packets; the intermediate satellite node is the node that forwards data packets; the observation information includes the current data queue length of the current satellite node and the current data queue lengths of its neighbor nodes, as well as the distances from the neighbor nodes to the target satellite node.

[0128] Step S302: Input the observation information of the current satellite node into the satellite routing prediction model deployed in the current satellite node to obtain the Q values of multiple actions.

[0129] The action is to select the next hop of the data packet among the neighbor nodes under the current observation information.

[0130] Step S303: Select the next satellite node for data packet transmission based on the action corresponding to the maximum Q value; repeat the above routing process for the next satellite node until the target satellite node is reached.

[0131] The embodiment of the present invention adopts a framework of centralized training and distributed execution. During the distributed execution, the satellite nodes make independent routing decisions only through their own observation information, saving the communication overhead between satellites.

[0132] Based on the same inventive concept, a low-earth orbit satellite routing device is provided, which is applied to the training method of the above satellite routing prediction model, as Figure 4 shown, the device includes:

[0133] An acquisition unit 401 is configured to acquire the observation information of the current satellite node. The current satellite node includes an initial satellite node and an intermediate satellite node. The initial satellite node is the satellite node that generates data packets, and the intermediate satellite node is the node that forwards data packets. The observation information includes the current data queue length of the current satellite node, the current data queue lengths of its neighbor nodes, and the distances from the neighbor nodes to the target satellite node.

[0134] A prediction unit 402 is configured to input the observation information of the current satellite node into the satellite routing prediction model deployed in the current satellite node to obtain the Q values of multiple actions. An action is to select the next hop of the data packet among the neighbor nodes under the current observation information.

[0135] A selection unit 403 is configured to select the next satellite node for data packet transmission based on the action corresponding to the maximum Q value, and repeatedly execute the above routing process for the next satellite node until the target satellite node is reached.

[0136] Based on the same technical concept, an embodiment of the present invention further provides an electronic device, as Figure 5 shown, including a processor 501, a communication interface 502, a memory 503, and a communication bus 504. Among them, the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.

[0137] The memory 503 is used to store a computer program.

[0138] When the processor 501 is configured to execute the program stored on the memory 503, it implements the steps of the training method of the satellite routing prediction model and the low-earth orbit satellite routing method.

[0139] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0140] The communication interface is used for communication between the above electronic device and other devices.

[0141] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0142] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0143] The computer program product for the method of training a satellite routing prediction model and the low-earth orbit satellite routing method provided by the embodiments of the present invention includes a computer-readable storage medium storing program code, and the instructions included in the program code can be used to execute the methods described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments and will not be elaborated here.

[0144] The device for the low-earth orbit satellite routing method provided by the embodiments of the present invention may be specific hardware on the device, or software or firmware installed on the device, etc. For the device provided by the embodiments of the present invention, its implementation principle and the technical effects produced are the same as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiments, reference can be made to the corresponding content in the foregoing method embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the foregoing described systems, devices, and units can all refer to the corresponding processes in the above method embodiments and will not be elaborated here.

[0145] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. Another example is that multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be electrical, mechanical, or other forms.

[0146] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0147] In addition, each functional unit in the embodiments provided by the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0148] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0149] It should be noted that: similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0150] Finally, it should be noted that: the above-described embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. All should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for training a satellite route prediction model, characterized in that: Applied to low-orbit satellite routing scenarios, each satellite node corresponds to a satellite routing prediction model, and the method includes: Obtain observation information of each satellite node under the current satellite network environment state and a weight combination of each objective randomly selected based on a multi-objective optimization model; the observation information includes the current data queue length of each satellite node and the current data queue length of its neighboring node, and the distance from the neighboring node to the target satellite node; the multi-objective optimization model is a target optimization model pre-constructed based on various routing indicators, and the routing indicators at least include routing transmission delay and routing load balancing; Among them, the number of training times is preset at intervals to increase the probability of weight combinations that are not selected; The observation information of each satellite node and the weight combination of each randomly selected target are input into the respective satellite route prediction models to obtain the Q values ​​of multiple actions; the action is to select the next hop of the data packet in the neighboring node under the current input, and the Q value is the expected reward value obtained after selecting the action; The Q values ​​output by each satellite routing prediction model are input into the Mixing network for joint training, wherein the Mixing network is composed of a main network and a secondary network; the main network and the secondary network each include two groups of networks, each group of networks corresponds to an optimization target; one group of networks corresponds to routing delay, and one group of networks corresponds to load balancing; the Q values ​​corresponding to the same target are combined and input into a corresponding group of networks, and finally the two groups of networks are integrated to output a multi-target joint Q value; that is, the main network outputs a main joint Q value, and the secondary network outputs a secondary joint Q value; Calculating a loss value based on the primary joint Q value and the secondary joint Q value and a preset loss function; The parameters of the satellite route prediction model are adjusted based on the loss value, and the above training process is repeated until a preset iteration condition is reached.

2. The method according to claim 1, characterized in that The method further comprises: The action corresponding to the maximum Q value among the Q values ​​output by each satellite routing prediction model interacts with the current satellite network environment to generate the next satellite network environment state, and obtain the next observation information under the next satellite network environment state; Calculate the reward value for executing the action with the maximum Q value based on the reward function; The current satellite network environment state, current observation information, reward value, next satellite network environment state, next observation information and weight are combined to construct a set of experience data; storing the experience data in a pre-built experience replay pool; When the amount of data in the experience replay pool reaches a preset threshold, a number of experience samples are randomly selected from the experience replay pool to train the Mixing network.

3. The method according to claim 1, characterized in that The method further comprises: A monotonic constraint condition is established for the Mixing network and each satellite route prediction model so that the action corresponding to the maximum Q value of each route prediction model is consistent with the action corresponding to the maximum joint Q value output by the Mixing network, wherein the monotonic constraint condition is: The left side shows the best action selection for the Mixing network as a whole; the right side shows the best action selection for each routing prediction model; Indicates that in a given state , and randomly selected weight combinations Next, select Action Expected value; Indicates the number of route prediction models.

4. The method according to claim 1, characterized in that: The multi-objective optimization model is: in, , Both represent routing delay indicators; , Both represent routing load balancing indicators; Indicates the total number of data packets; Represents the local area formed by the current satellite node and its neighboring nodes; represents a satellite node; Indicates the local area The current queue length of satellite nodes; Represents the average queue length of all satellite nodes in the local area.

5. The method according to claim 1, characterized in that The method further comprises: During the training process, the global status information of all satellite nodes is obtained at each preset interval; The global state information is input into the super network to generate the parameters of the Mixing network.

6. A low-orbit satellite routing method, characterized in that: A training method for a satellite route prediction model applied to any one of claims 1 to 5, the method comprising: Acquire observation information of a current satellite node; the current satellite node includes an initial satellite node and an intermediate satellite node, the initial satellite node is a satellite node that generates a data packet; the intermediate satellite node is a node that forwards a data packet; the observation information includes a current data queue length of the current satellite node and a current data queue length of its neighboring node, and a distance from the neighboring node to the target satellite node; Inputting the observation information of the current satellite node into the satellite routing prediction model deployed in the current satellite node to obtain Q values ​​of multiple actions; the action is to select the next hop of the data packet in the neighboring node under the current observation information; Select the next satellite node for data packet transmission based on the action corresponding to the maximum Q value; The above routing process is repeated for the next satellite node until the target satellite node is reached.

7. A low-orbit satellite routing device, characterized in that: A training method for a satellite route prediction model applied to any one of claims 1 to 5, the device comprising: An acquisition unit is used to acquire observation information of a current satellite node; the current satellite node includes an initial satellite node and an intermediate satellite node, the initial satellite node is a satellite node that generates a data packet; the intermediate satellite node is a node that forwards a data packet; the observation information includes a current data queue length of the current satellite node and a current data queue length of its neighboring node, and a distance from the neighboring node to the target satellite node; A prediction unit, used to input the observation information of the current satellite node into a satellite routing prediction model deployed in the current satellite node to obtain Q values ​​of multiple actions; the action is to select the next hop of the data packet in the neighboring node under the current observation information; The selection unit is used to select the next satellite node for data packet transmission based on the action corresponding to the maximum Q value; and repeatedly execute the above routing process for the next satellite node until the target satellite node is reached.

8. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, for implementing the method steps described in any one of claims 1 to 6 when executing a program stored in a memory.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method steps described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Satellite network transmission method and device and electronic equipment

    CN111490817A

  • Fully distributed routing method and system based on deep reinforcement learning

    CN116248164A