A Satellite Networking Routing Method Based on Deep Reinforcement Learning

Through deep reinforcement learning, the satellite network routing algorithm is optimized, and the data transmission efficiency and packet loss caused by high latency and high bit error rates of satellite network are solved, and efficient routing decisions are achieved in a dynamic environment.

CN119834870BActive Publication Date: 2025-07-11BEIJING INST OF TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510259107.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-07-11
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

The high latency and high bit error rate of satellite networking make it difficult for routing algorithms to adapt to frequently changing network environments, resulting in inefficient data transmission and serious packet loss.

Method used

The satellite network routing method based on deep reinforcement learning is adopted, and the routing algorithm is optimized by generating the first and second action value functions of the decision action, combining exploration and utilization strategies.

Benefits of technology

It improves the adaptability of satellite networks in dynamic environments, and can find the optimal routing path under multiple constraints, improves data transmission efficiency, and reduces packet loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119834870B_ABST
    Figure CN119834870B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of communication technologies, and particularly to a satellite networking routing method based on deep reinforcement learning, including: generating a first action value function corresponding to each decision-making action based on the feature vector of the current network state, determining the decision-making action selection probability in the current network state based on the current target model, and using a preset exploration and exploitation strategy, the decision-making action selection probability, and the first action value function to select each decision-making action and output a corresponding second action value function, calculating the target action value function corresponding to each decision-making action in the initialized target network and the difference value between the target action value function and the second action value function, so as to calculate the minimum loss function of the current target model according to the difference value and optimize the satellite networking routing algorithm. Thereby, the problems that the routing algorithm is difficult to adapt to the frequently changing network environment due to the high latency and high bit error rate of satellite networking, resulting in low data transmission efficiency, packet loss, etc. are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and particularly to a satellite networking routing method based on deep reinforcement learning. Background Art

[0002] Satellite networks, especially LEO (Low Earth Orbit) satellite networking, due to their high latency, high bit error rate, and frequently changing dynamic topology, make routing algorithms need to quickly adapt to this highly dynamic network environment, thus becoming a major challenge in network design and optimization.

[0003] In related technologies, due to the high latency and high bit error rate of satellite networking, which increase the uncertainty and complexity of data transmission, its dynamic topology requires routing algorithms to have high flexibility and fast response capabilities to adapt to frequently changing link states.

[0004] However, the routing algorithms in related technologies perform poorly in the face of the above situations and are difficult to adapt to the frequently changing network environment, resulting in low data transmission efficiency and even serious packet loss. In addition, routing decisions in satellite networks need to comprehensively consider various constraints, such as bandwidth, latency, hop count, load balancing, etc., further increasing the complexity of routing decisions, which urgently need to be solved. Summary of the Invention

[0005] The present invention provides a satellite networking routing method based on deep reinforcement learning to solve the problems that due to the high latency and high bit error rate of satellite networking, routing algorithms are difficult to adapt to the frequently changing network environment, resulting in low data transmission efficiency and even serious packet loss.

[0006] A first aspect embodiment of the present invention provides a satellite networking routing method based on deep reinforcement learning, including the following steps:

[0007] Obtain the current network state, and generate a first action value function corresponding to each decision action in the current network state based on the feature vector of the current network state;

[0008] Based on the current target model, determine the decision action selection probability in the current network state, use a preset exploration and exploitation strategy, and based on the decision action selection probability and the first action value function, select each decision action, and output a corresponding second action value function according to the selected decision action;

[0009] Calculate the target action value function corresponding to each decision action in the target network after initialization, calculate the minimum loss function of the current target model according to the difference value between the target action value function and the second action value function, and determine the optimal optimization strategy according to the minimum loss function, so as to optimize the satellite networking routing algorithm according to the optimal optimization strategy.

[0010] According to an embodiment of the present invention, the use of the preset exploration and exploitation strategy, and based on the decision action selection probability and the first action value function to select each decision action, and output the corresponding second action value function according to the selected decision action, includes:

[0011] Use the preset exploration and exploitation strategy to randomly select from each decision action according to the first decision action selection probability and the first action value function, and output the corresponding second action value function according to the selected decision action;

[0012] Or, use the preset exploration and exploitation strategy to determine the highest action value function from each decision action according to the second decision action selection probability and the first action value function, and output the corresponding second action value function according to the selected decision action.

[0013] According to an embodiment of the present invention, after outputting the corresponding second action value function according to the selected decision action, it further includes:

[0014] Based on the second action value function, determine the reward function of the current target model, generate a new network state based on the selected decision action, and generate an experience sample of the current target model according to the reward function, the current network state, the selected decision action, and the new network state;

[0015] Determine the target task of the current target model, set the action experience replay buffer of the current target model according to the target task, and store the experience sample of the current target model in the action experience replay buffer;

[0016] Judge whether the action experience replay buffer is full;

[0017] If the action experience replay buffer is full, replace the earliest stored experience sample with the latest stored experience sample.

[0018] According to an embodiment of the present invention, the determining the reward function of the current target model based on the second action value function includes:

[0019] Define the positive reward and negative reward of the initial reward function based on the feedback result of the target task and the current network state;

[0020] Define the weighted sum and reward function of the target task according to the target task, and adjust the positive reward and the negative reward by using the weight coefficients of the weighted sum and reward function to optimize the initial reward function and obtain the reward function of the current target model.

[0021] According to an embodiment of the present invention, after storing the experience samples of the current target model into the action experience replay buffer, it further includes:

[0022] Randomly extract a preset number of experience samples from the current target model in the action experience replay buffer at a preset period, and update the current target model based on the extracted experience samples.

[0023] According to an embodiment of the present invention, before obtaining the current network state, it further includes:

[0024] Construct an initial target model;

[0025] Initialize the initial target model, train the initialized initial target model to obtain the current target model, and obtain the current network state based on the current target model.

[0026] According to a satellite networking routing method based on deep reinforcement learning in an embodiment of the present invention, generate a first action value function corresponding to each decision action based on the feature vector of the current network state, determine the decision action selection probability in the current network state based on the current target model, and use a preset exploration and exploitation strategy, decision action selection probability, and first action value function to select each decision action and output a corresponding second action value function, calculate the target action value function corresponding to each decision action in the initialized target network, and calculate the difference value between the target action value function and the second action value function, so as to calculate the minimum loss function of the current target model according to the difference value to optimize the satellite networking routing algorithm. Thereby, the problems that the routing algorithm is difficult to adapt to the frequently changing network environment due to the high latency and high bit error rate of satellite networking, resulting in low data transmission efficiency and serious packet loss are solved.

[0027] An embodiment of the second aspect of the present invention provides a satellite networking routing device based on deep reinforcement learning, including:

[0028] A generation module, configured to obtain the current network state, and generate a first action value function corresponding to each decision action in the current network state based on the feature vector of the current network state;

[0029] A determination module, configured to determine the decision action selection probability in the current network state based on the current target model, utilize a preset exploration and exploitation strategy, select each decision action based on the decision action selection probability and the first action value function, and output the corresponding second action value function according to the selected decision action;

[0030] An optimization module, configured to calculate the target action value function corresponding to each decision action in the initialized target network, calculate the minimum loss function of the current target model according to the difference value between the target action value function and the second action value function, and determine the best optimization strategy according to the minimum loss function, so as to optimize the satellite networking routing algorithm according to the best optimization strategy.

[0031] According to an embodiment of the present invention, the determination module includes:

[0032] A selection unit, configured to randomly select from each decision action according to the first decision action selection probability and the first action value function by using the preset exploration and exploitation strategy, and output the corresponding second action value function according to the selected decision action;

[0033] Alternatively, determine the highest action value function from each decision action according to the second decision action selection probability and the first action value function by using the preset exploration and exploitation strategy, and output the corresponding second action value function according to the selected decision action.

[0034] According to an embodiment of the present invention, after outputting the corresponding second action value function according to the selected decision action, the determination module further includes:

[0035] A generation unit, configured to determine the reward function of the current target model based on the second action value function, generate a new network state based on the selected decision action, and generate an experience sample of the current target model according to the reward function, the current network state, the selected decision action, and the new network state;

[0036] A storage unit, configured to determine the target task of the current target model, set an action experience replay buffer for the current target model according to the target task, and store the experience sample of the current target model into the action experience replay buffer;

[0037] A judgment unit, configured to judge whether the action experience replay buffer is full;

[0038] A replacement unit, configured to, if the action experience replay buffer is full, replace the earliest stored experience sample with the latest stored experience sample.

[0039] According to an embodiment of the present invention, the generating unit includes:

[0040] A defining subunit, configured to define positive and negative rewards of an initial reward function based on a feedback result of the target task and the current network state;

[0041] An optimizing subunit, configured to define a weighted sum reward function of the target task according to the target task, and adjust the positive and negative rewards by using a weight coefficient of the weighted sum reward function to optimize the initial reward function to obtain a reward function of the current target model.

[0042] According to an embodiment of the present invention, after storing an experience sample of the current target model into the action experience replay buffer, the storing unit further includes:

[0043] An updating subunit, configured to randomly extract a preset number of experience samples in the current target model from the action experience replay buffer at a preset period, and update the current target model based on the extracted experience samples.

[0044] According to an embodiment of the present invention, before obtaining the current network state, the generating module further includes:

[0045] A constructing unit, configured to construct an initial target model;

[0046] An obtaining unit, configured to initialize the initial target model, train the initialized initial target model to obtain a current target model, and obtain the current network state based on the current target model.

[0047] A satellite networking routing device based on deep reinforcement learning according to an embodiment of the present invention generates a first action value function corresponding to each decision action based on a feature vector of the current network state, determines a decision action selection probability in the current network state based on the current target model, and uses a preset exploration and exploitation strategy, the decision action selection probability, and the first action value function to select each decision action to output a corresponding second action value function, calculates a target action value function corresponding to each decision action in the initialized target network, and calculates a difference value between the target action value function and the second action value function, so as to calculate a minimum loss function of the current target model according to the difference value to optimize the satellite networking routing algorithm. Thereby, the problems that the routing algorithm is difficult to adapt to a frequently changing network environment due to the high latency and high bit error rate of satellite networking, resulting in low data transmission efficiency and serious packet loss are solved.

[0048] An embodiment of the third aspect of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement a satellite networking routing method based on deep reinforcement learning as described in the above embodiments.

[0049] An embodiment of the fourth aspect of the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to execute a satellite networking routing method based on deep reinforcement learning as described in the above embodiments.

[0050] An embodiment of the fifth aspect of the present invention provides a computer program product including computer programs / instructions that, when executed by a processor, implement a satellite networking routing method based on deep reinforcement learning as described in the above embodiments.

[0051] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:

[0053] Figure 1 FIG. is a flowchart of a satellite networking routing method based on deep reinforcement learning according to an embodiment of the present invention;

[0054] Figure 2 FIG. is an overall flowchart of an optimized routing algorithm for satellite networking according to an embodiment of the present invention;

[0055] Figure 3 FIG. is an example diagram of a satellite networking routing device based on deep reinforcement learning according to an embodiment of the present invention;

[0056] Figure 4 FIG. is a schematic structural diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.

[0058] The following describes a satellite networking routing method based on deep reinforcement learning according to an embodiment of the present invention. Aiming at the problem in the above-mentioned background technology that due to the high latency and high bit error rate of satellite networking, the routing algorithm is difficult to adapt to the frequently changing network environment, resulting in low data transmission efficiency and even serious packet loss, the present invention provides a satellite networking routing method based on deep reinforcement learning. In this method, a first action value function corresponding to each decision action is generated based on the feature vector of the current network state, the selection probability of the decision action in the current network state is determined based on the current target model, and each decision action is selected by using a preset exploration and exploitation strategy, the decision action selection probability, and the first action value function to output a corresponding second action value function. The target action value function corresponding to each decision action in the initialized target network is calculated, and the difference value between the target action value function and the second action value function is calculated to calculate the minimum loss function of the current target model according to the difference value, so as to optimize the satellite networking routing algorithm. Thus, the problems of low data transmission efficiency and serious packet loss caused by the high latency and high bit error rate of satellite networking, which make it difficult for the routing algorithm to adapt to the frequently changing network environment, are solved.

[0059] Specifically, before introducing the embodiments of the present invention, the defects of the satellite networking routing algorithm in the related technology are first introduced. Due to the long distance and special transmission medium, satellite communication links usually have high latency and bit error rate, which pose a huge challenge to traditional routing algorithms. Traditional algorithms often cannot operate efficiently in such an environment of high latency and high bit error rate, resulting in low data transmission efficiency and even serious packet loss.

[0060] The topological structure of the satellite network often changes. For example, low Earth orbit satellites move rapidly in orbit, which will cause the link connections between satellites to change frequently. This dynamicity requires the routing algorithm to be able to quickly adapt to the new network topology, while traditional static or semi-static routing algorithms are difficult to work effectively in such an environment.

[0061] In addition, routing decisions in satellite networks need to consider multiple constraints, such as bandwidth, latency, hop count, load balancing, etc. Since these factors interact with each other, it will also make the routing decision very complex. Traditional routing algorithms usually can only consider a small number of constraints and are difficult to find the optimal routing path under multiple constraints. Therefore, the embodiments of the present invention propose a solution to optimize the satellite networking routing algorithm by combining deep reinforcement learning with GPT-4.0, so as to improve the efficiency of routing decisions. The specific implementation scheme will be described in the following embodiments.

[0062] Specifically, Figure 1 It is a schematic flowchart of a satellite networking routing method based on deep reinforcement learning provided by an embodiment of the present invention.

[0063] As Figure 1 shown, a satellite networking routing method based on deep reinforcement learning includes the following steps:

[0064] In step S101, obtain the current network state, and generate a first action value function corresponding to each decision action in the current network state based on the feature vector of the current network state.

[0065] According to an embodiment of the present invention, before obtaining the current network state, it further includes: constructing an initial target model; initializing the initial target model, training the initialized initial target model to obtain the current target model, and obtaining the current network state based on the current target model.

[0066] Specifically, in the embodiment of the present invention, an initial target model is first constructed, such as a DQN (Deep Q-Network) model. At this time, the DQN model is an untrained model. The DQN model combines the traditional Q-learning algorithm with deep learning and approximates the Q-value function by using a convolutional neural network. Secondly, the DQN model is initialized, and the first training is performed based on the initialized DQN model to obtain the current target model, that is, the trained DQN model, and the current network state is obtained based on the current target model. Finally, receive the feature vector of the current network state, and generate a first action value function corresponding to each possible output decision action in the current network state based on the received feature vector, that is, the Q-value.

[0067] It should be noted that the input feature vector represents a numerical array of the current network state, including node connection status, link bandwidth, link delay, packet loss rate, traffic load, etc., and the output Q-value represents a vector, where each element represents the value of performing a specific decision action in the current network state. In this process, the embodiment of the present invention can use the GPT-4.0 technology to perform natural language description and interpretation on the network state and decision-making process to improve the interpretability and transparency of the DQN model.

[0068] In step S102, based on the current target model, determine the decision action selection probability in the current network state, use the preset exploration and exploitation strategy, select each decision action based on the decision action selection probability and the first action value function, and output a corresponding second action value function according to the selected decision action.

[0069] According to an embodiment of the present invention, using a preset exploration and exploitation strategy, and based on the decision-making action selection probability and the first action value function, each decision-making action is selected, and the corresponding second action value function is output according to the selected decision-making action, including: randomly selecting from each decision-making action according to the first decision-making action selection probability and the first action value function by using the preset exploration and exploitation strategy, and outputting the corresponding second action value function according to the selected decision-making action; or, determining the highest action value function from each decision-making action according to the second decision-making action selection probability and the first action value function by using the preset exploration and exploitation strategy, and outputting the corresponding second action value function according to the selected decision-making action.

[0070] Among them, the preset exploration and exploitation strategy can be selected by those skilled in the art according to actual experimental requirements, and no specific limitation is made here.

[0071] Specifically, as Figure 2 shown, at each time step, the agent (such as the DQN model) observes the current network state based on the current target model updated this time. For example, by observing the dynamic changes of the current network link, traffic fluctuations, etc., the decision-making action selection probability in the current network state is determined, and the preset exploration and exploitation strategy, such as the ε-greedy strategy, is used to randomly select from each decision-making action according to the first decision-making action selection probability (such as ε) and the first action value function. By randomly selecting actions, the agent can explore new state-action pairs, which helps to discover action selections that are better than the current strategy, thereby improving the effect of the overall strategy. At the same time, exploration can prevent the agent from falling into local optimal solutions to ensure that the agent can globally optimize its strategy.

[0072] Optionally, the embodiment of the present invention can also use the ε-greedy strategy to determine the decision-making action corresponding to the highest action value function from each decision-making action according to the second decision-making action selection probability (such as 1 - ε) and the first action value function, and output the corresponding second action value function, that is, the highest Q value. By selecting the decision-making action corresponding to the current highest Q value, the agent can use the known strategy to obtain the maximum cumulative reward. Using the existing experience, the agent can achieve optimal decision-making in the short term to improve the performance of the system.

[0073] Among them, the two action selection formulas are as follows:

[0074]

[0075] It should be noted that the purpose of using the ε-greedy strategy in the embodiments of the present invention is to balance the exploration of new strategies and the exploitation of existing strategies when training the agent. For the selection of ε, the embodiments of the present invention use GPT-4.0 to generate natural language prompts for attenuation. In the initial stage of training, a relatively high ε value can be selected to encourage exploration to avoid prematurely falling into local optimal solutions. As the training progresses, the embodiments of the present invention can make more refined adjustments to the ε value according to the prompts. If the performance of the DQN model improves during the training process, the ε value can be decreased according to the linear attenuation strategy to increase the exploitation of existing strategies. If the performance of the DQN model tends to be stable during the training process, the ε value used remains stable, and at this time, more learned strategies are utilized to improve the stability and performance of the model.

[0076] Among them, the linear attenuation strategy formula can be expressed as:

[0077]

[0078] Among them, are the start and end time steps of the initial stage, are the start and end time steps of the intermediate stage, is the exploration probability of the initial stage, is the exploration probability of the intermediate stage.

[0079] According to an embodiment of the present invention, after outputting the corresponding second action value function according to the selected decision action, it further includes: determining the reward function of the current target model based on the second action value function, generating a new network state based on the selected decision action, generating an experience sample of the current target model according to the reward function, the current network state, the selected decision action, and the new network state; determining the target task of the current target model, setting the action experience replay buffer of the current target model according to the target task, and storing the experience sample of the current target model in the action experience replay buffer; determining whether the action experience replay buffer is full; if the action experience replay buffer is full, replacing the earliest stored experience sample with the latest stored experience sample.

[0080] According to an embodiment of the present invention, determining the reward function of the current target model based on the second action value function includes: defining the positive reward and negative reward of the initial reward function based on the feedback result of the target task and the current network state; defining the weighted sum reward function of the target task according to the target task, and adjusting the positive reward and negative reward using the weight coefficient of the weighted sum reward function to optimize the initial reward function to obtain the reward function of the current target model.

[0081] Specifically, as Figure 2As shown in the figure, in the embodiment of the present invention, based on the target task of the DQN model, the positive and negative rewards of the initial reward function are defined according to the feedback result of the target task and the current network state. For example, when the delay experienced by the data packet increases, the packet loss rate rises, the routing decision cost increases, the DQN model detects abnormal traffic or potential network attacks, etc., negative rewards are given; when the network throughput increases, the network load distribution is more balanced, the network resources are utilized more effectively, the stability of the network state is enhanced, and the intelligent agent can better adapt to the changes in the network state, etc., positive rewards are given.

[0082] Considering the above factors, the embodiment of the present invention can design a weighted sum reward function, and use the weight coefficients of the weighted sum reward function to adjust the positive and negative rewards to balance multiple target tasks, thereby optimizing the initial reward function to obtain the reward function of the current target model, that is, the reward function of the DQN model. Taking the four objectives of delay change, packet loss rate change, network throughput change, and network load distribution as examples:

[0083]

[0084] Among them, is the throughput reward, is the load balancing reward, is the delay penalty, is the packet loss rate penalty, is the weight coefficient of the throughput reward, is the weight coefficient of the load balancing reward, is the weight coefficient of the delay penalty, is the weight coefficient of the packet loss rate penalty, thereby adjusting the importance of different factors through the weight coefficients.

[0085] Furthermore, when initializing the DQN model, the initial values of the weight coefficients are set. During the training process, the weight coefficients are determined by GPT-4.0 according to the requirements of specific situations. If the delay of the current network is high, the weight coefficient of the delay penalty can be increased at this time to prompt the model to pay more attention to reducing the delay. By dynamically adjusting and interpreting the reward function, the adaptability of the model in a complex network environment is improved.

[0086] It should be noted that the reward function provides a clear learning goal for the intelligent agent, that is, to pursue the maximization of the cumulative reward. By quantifying the strategy of the intelligent agent into a numerical value, the algorithm can evaluate and optimize these performance indicators, thereby helping the intelligent agent to optimize the routing algorithm.

[0087] Furthermore, as Figure 2As shown, in the embodiment of the present invention, after the agent executes an action, it is necessary to store the experience samples of the DQN model, and update the DQN model by implementing experience replay regularly. First, determine the reward function of the current target model based on the second action value function, and generate a new network state after the execution of the action based on the selected decision-making action. Generate the experience samples of the current target model according to the reward function, the current network state, the selected decision-making action, and the new network state; Secondly, set the action experience replay buffer of the DQN model according to the specific target task to be executed, convert each experience sample into a quadruple (such as state, executed action, reward, new state), and store it in the action experience replay buffer; Finally, determine whether the action experience replay buffer is full; if the action experience replay buffer is full, replace the earliest stored experience sample with the latest stored experience sample to keep the experience replay buffer updated.

[0088] According to an embodiment of the present invention, after storing the experience samples of the current target model in the action experience replay buffer, it further includes: randomly extracting a preset number of experience samples in the current target model from the action experience replay buffer according to a preset period, and updating the current target model based on the extracted experience samples.

[0089] Among them, both the preset period and the preset number can be selected by those skilled in the art according to actual experimental requirements, and no specific limitation is made here.

[0090] Specifically, during the training process of experience replay, a certain number of experience samples (forming a mini-batch) can be randomly extracted from the experience replay buffer according to a preset period, and the current target model can be updated based on the extracted experience samples, so as to reduce the correlation between data, improve the stability of the training process, and then continuously obtain the decision-making action corresponding to the highest Q value according to the iteratively updated current target model, so as to obtain the maximized cumulative reward, and use the existing experience to enable the agent to achieve optimal decision-making in the short term to improve the performance of the system.

[0091] In step S103, calculate the target action value function corresponding to each decision-making action in the initialized target network, calculate the minimum loss function of the current target model according to the difference value between the target action value function and the second action value function, and determine the best optimization strategy according to the minimum loss function to optimize the satellite networking routing algorithm according to the best optimization strategy.

[0092] Specifically, in the embodiments of the present invention, a target network introduced is utilized to calculate a stable target action value function, i.e., the target Q-value, to reduce the fluctuations in the training process. The initial target network is synchronously initialized when initializing the DQN model, and the parameters of the subsequent target network are copied from the DQN model every fixed number of steps, thereby realizing the update of the target network.

[0093] Furthermore, in the embodiments of the present invention, MSE (Mean Squared Error) can be used to calculate the difference between the second Q-value (i.e., the second action value function) in the online current network state and the target Q-value in the target network. By measuring the difference between the second Q-value and the target Q-value and guiding the training of the DQN model, the loss function value is minimized to obtain the minimum loss function of the current target model. Then, through the minimum loss function, the DQN model can learn a policy that can select the best action under a given state to maximize the expected cumulative reward, thereby optimizing the satellite networking routing algorithm. Thus, the DQN training method based on the target network not only improves the stability of the learning process but also helps to improve the performance of the final policy.

[0094] Among them, MSE can be expressed as:

[0095]

[0096] where N is the number of experience samples, is the Q-value predicted by the DQN model, is the target Q-value.

[0097] The target Q-value can be expressed as:

[0098]

[0099] where is the obtained reward, is the discount factor, are the parameters of the target network, The determination of needs to comprehensively consider many factors. Since this task requires long-term planning, a relatively high discount factor can be initially selected without detailed description.

[0100] In summary, through the specific description of the above embodiments, the embodiments of the present invention can achieve the following beneficial effects:

[0101] (1) Strong dynamic adaptability: Through continuous learning and optimization, the deep reinforcement learning model can quickly adapt to new topological structures and network states in a dynamically changing satellite network environment. The agent can adjust the routing strategy according to real-time data to ensure the optimization of network performance. By combining the natural language descriptions and explanations generated by GPT-4.0, operators can better understand and manage the network decision-making process, improving the interpretability and transparency of the model.

[0102] (2) Capable of handling multiple constraint conditions: By designing a reasonable reward function, the DRL model can simultaneously consider multiple constraint conditions, such as bandwidth, latency, hop count, load balancing, etc., to achieve multi-objective optimization. GPT-4.0 can dynamically adjust and interpret these weight coefficients, generate adjustment suggestions according to the actual network situation, and improve the adaptability of the model in complex network environments. This multi-constraint handling ability is significantly superior to traditional routing algorithms, enabling the optimal routing path to be found even in complex network environments.

[0103] According to an embodiment of the present invention, a satellite networking routing method based on deep reinforcement learning generates a first action value function corresponding to each decision action based on the feature vector of the current network state, determines the decision action selection probability in the current network state based on the current target model, and uses a preset exploration and exploitation strategy, decision action selection probability, and first action value function to select each decision action and output the corresponding second action value function. Calculate the target action value function corresponding to each decision action in the initialized target network, and calculate the difference value between the target action value function and the second action value function, so as to calculate the minimum loss function of the current target model according to the difference value to optimize the satellite networking routing algorithm. Thus, the problems of low data transmission efficiency and serious packet loss caused by the high latency and high bit error rate of satellite networking, which make it difficult for routing algorithms to adapt to frequently changing network environments, are solved.

[0104] Next, a satellite networking routing device based on deep reinforcement learning proposed according to an embodiment of the present invention is described with reference to the accompanying drawings.

[0105] Figure 3 It is a block diagram of a satellite networking routing device based on deep reinforcement learning according to an embodiment of the present invention.

[0106] As Figure 3 shown, the satellite networking routing device 10 based on deep reinforcement learning includes: a generation module 100, a determination module 200, and an optimization module 300.

[0107] Among them, the generation module 100 is used to obtain the current network state and generate a first action value function corresponding to each decision action in the current network state based on the feature vector of the current network state;

[0108] A determination module 200, configured to determine the decision action selection probability in the current network state based on the current target model, utilize a preset exploration and exploitation strategy, select each decision action based on the decision action selection probability and the first action value function, and output the corresponding second action value function according to the selected decision action;

[0109] An optimization module 300, configured to calculate the target action value function corresponding to each decision action in the initialized target network, calculate the minimum loss function of the current target model according to the difference between the target action value function and the second action value function, and determine the optimal optimization strategy according to the minimum loss function, so as to optimize the satellite network routing algorithm according to the optimal optimization strategy.

[0110] According to an embodiment of the present invention, the determination module 200 includes:

[0111] A selection unit, configured to randomly select from each decision action according to the first decision action selection probability and the first action value function by using a preset exploration and exploitation strategy, and output the corresponding second action value function according to the selected decision action;

[0112] Alternatively, determine the highest action value function from each decision action according to the second decision action selection probability and the first action value function by using a preset exploration and exploitation strategy, and output the corresponding second action value function according to the selected decision action.

[0113] According to an embodiment of the present invention, after outputting the corresponding second action value function according to the selected decision action, the determination module 200 further includes:

[0114] A generation unit, configured to determine the reward function of the current target model based on the second action value function, generate a new network state based on the selected decision action, and generate an experience sample of the current target model according to the reward function, the current network state, the selected decision action, and the new network state;

[0115] A storage unit, configured to determine the target task of the current target model, set the action experience replay buffer of the current target model according to the target task, and store the experience sample of the current target model in the action experience replay buffer;

[0116] A judgment unit, configured to judge whether the action experience replay buffer is full;

[0117] A replacement unit, configured to, if the action experience replay buffer is full, replace the earliest stored experience sample with the latest stored experience sample.

[0118] According to an embodiment of the present invention, after outputting a corresponding second action value function according to the selected decision action, the determination module 200 includes:

[0119] According to an embodiment of the present invention, the generation unit includes: a positive reward and a negative reward for defining an initial reward function based on the feedback result of the target task and the current network state;

[0120] The optimization subunit is configured to define a weighted sum reward function of the target task according to the target task, and adjust the positive reward and the negative reward by using the weight coefficient of the weighted sum reward function to optimize the initial reward function to obtain the reward function of the current target model.

[0121] According to an embodiment of the present invention, after storing the experience samples of the current target model into the action experience replay buffer, the storage unit further includes:

[0122] The update subunit is configured to randomly extract a preset number of experience samples from the current target model from the action experience replay buffer at a preset period, and update the current target model based on the extracted experience samples.

[0123] According to an embodiment of the present invention, before obtaining the current network state, the generation module 100 further includes:

[0124] The construction unit is configured to construct an initial target model;

[0125] The acquisition unit is configured to initialize the initial target model, train the initialized initial target model to obtain the current target model, and obtain the current network state based on the current target model.

[0126] A satellite networking routing device based on deep reinforcement learning according to an embodiment of the present invention generates a first action value function corresponding to each decision action based on the feature vector of the current network state, determines the decision action selection probability in the current network state based on the current target model, and uses a preset exploration and exploitation strategy, the decision action selection probability, and the first action value function to select each decision action to output a corresponding second action value function, calculates the target action value function corresponding to each decision action in the initialized target network, and calculates the difference value between the target action value function and the second action value function, so as to calculate the minimum loss function of the current target model according to the difference value to optimize the satellite networking routing algorithm. Thereby, the problems that the routing algorithm is difficult to adapt to the frequently changing network environment due to the high latency and high bit error rate of satellite networking, resulting in low data transmission efficiency and serious packet loss are solved.

[0127] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device may include:

[0128] A memory 401, a processor 402, and a computer program stored on the memory 401 and executable on the processor 402.

[0129] When the processor 402 executes the program, it implements a satellite network routing method based on deep reinforcement learning provided in the above embodiments.

[0130] Furthermore, the electronic device further includes:

[0131] A communication interface 403 for communication between the memory 401 and the processor 402.

[0132] The memory 401 is used to store a computer program executable on the processor 402.

[0133] The memory 401 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0134] If the memory 401, the processor 402, and the communication interface 403 are implemented independently, the communication interface 403, the memory 401, and the processor 402 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 4 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0135] Optionally, in a specific implementation, if the memory 401, the processor 402, and the communication interface 403 are integrated on a chip, the memory 401, the processor 402, and the communication interface 403 can communicate with each other through an internal interface.

[0136] The processor 402 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0137] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements a satellite networking routing method based on deep reinforcement learning as described above.

[0138] The embodiment of the present invention also provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, they implement a satellite networking routing method based on deep reinforcement learning as described in the above embodiment.

[0139] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0140] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0141] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a customized logic function or process. The scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of the present invention.

[0142] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from and execute instructions of the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or N wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which a program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then storing it in a computer memory.

[0143] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0144] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0145] In addition, each functional unit in various embodiments of the present invention may be integrated into a processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0146] The above-mentioned storage medium may be a read-only memory, a magnetic disk or an optical disc, etc. Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A satellite networking routing method based on deep reinforcement learning, characterized in that, It includes the following steps: Obtain the current network state, and generate a first action value function corresponding to each decision action in the current network state based on the feature vector of the current network state; Based on the current target model, determine the decision action selection probability in the current network state, use a preset exploration and exploitation strategy, and select each decision action based on the decision action selection probability and the first action value function, and output a corresponding second action value function according to the selected decision action; Calculate the target action value function corresponding to each decision action in the initialized target network, calculate the minimum loss function of the current target model according to the difference value between the target action value function and the second action value function, and determine the best optimization strategy according to the minimum loss function, so as to optimize the satellite network routing algorithm according to the best optimization strategy; Wherein, the preset exploration and exploitation strategy is the ε-greedy strategy, including ε and 1-ε. When selecting ε, use GPT-4.0 to generate natural language prompts for attenuation. In the initial stage of training the current target model, select an ε value that meets the preset conditions to encourage exploration. If the performance of the current target model improves, then reduce the ε value according to the linear attenuation strategy. If the performance of the current target model is stable, then keep the ε value stable; After outputting the corresponding second action value function according to the selected decision action, it further includes: determining the reward function of the current target model based on the second action value function, defining the positive reward and negative reward of the initial reward function based on the feedback result of the target task and the current network state; defining the weighted sum reward function of the target task according to the target task, and using the weight coefficient of the weighted sum reward function to adjust the positive reward and the negative reward, optimizing the initial reward function, and obtaining the reward function of the current target model.

2. The satellite networking routing method based on deep reinforcement learning according to claim 1, wherein The using the preset exploration and exploitation strategy, and selecting each decision action based on the decision action selection probability and the first action value function, and outputting a corresponding second action value function according to the selected decision action includes: Randomly select from each decision action according to the first decision action selection probability and the first action value function using the preset exploration and exploitation strategy, and output a corresponding second action value function according to the selected decision action; Or, determine the highest action value function from each decision action according to the second decision action selection probability and the first action value function using the preset exploration and exploitation strategy, and output a corresponding second action value function according to the selected decision action.

3. A satellite networking routing method based on deep reinforcement learning according to claim 1, characterized in that After outputting the corresponding second action value function according to the selected decision action, it further includes: Generate a new network state based on the selected decision action, and generate an experience sample of the current target model according to the reward function, the current network state, the selected decision action, and the new network state. Determine the target task of the current target model, set the action experience replay buffer of the current target model according to the target task, and store the experience samples of the current target model into the action experience replay buffer; Judge whether the action experience replay buffer is full; If the action experience replay buffer is full, replace the earliest stored experience sample with the latest stored experience sample.

4. A satellite networking routing method based on deep reinforcement learning according to claim 3, characterized in that After storing the experience samples of the current target model into the action experience replay buffer, it further includes: Randomly extract a preset number of experience samples from the action experience replay buffer at a preset period, and update the current target model based on the extracted experience samples.

5. A satellite networking routing method based on deep reinforcement learning according to claim 1, characterized in that Before obtaining the current network state, it further includes: Construct an initial target model; Initialize the initial target model, train the initialized initial target model to obtain the current target model, and obtain the current network state based on the current target model.

6. A satellite networking routing device based on deep reinforcement learning, characterized in that, It includes: A generation module, configured to obtain the current network state and generate a first action value function corresponding to each decision-making action in the current network state based on the feature vector of the current network state; A determination module, configured to determine the decision-making action selection probability in the current network state based on the current target model, use a preset exploration and exploitation strategy, and select each decision-making action based on the decision-making action selection probability and the first action value function, and output a corresponding second action value function according to the selected decision-making action; An optimization module, configured to calculate the target action value function corresponding to each decision-making action in the initialized target network, calculate the minimum loss function of the current target model according to the difference value between the target action value function and the second action value function, and determine the best optimization strategy according to the minimum loss function, so as to optimize the satellite networking routing algorithm according to the best optimization strategy; Wherein, the preset exploration and exploitation strategy is the ε-greedy strategy, including ε and 1-ε. When selecting ε, use GPT-4.0 to generate natural language prompts for attenuation. In the initial stage of training the current target model, select an ε value that meets the preset conditions to encourage exploration. If the performance of the current target model improves, reduce the ε value according to the linear attenuation strategy. If the performance of the current target model is stable, keep the ε value stable; After outputting the corresponding second action value function according to the selected decision-making action, the determination module further includes: a generation unit, configured to determine the reward function of the current target model based on the second action value function; the generation unit includes: a definition subunit, configured to define the positive reward and negative reward of the initial reward function based on the feedback result of the target task and the current network state; an optimization subunit, configured to define the weighted sum reward function of the target task according to the target task, and use the weight coefficient of the weighted sum reward function to adjust the positive reward and the negative reward, and optimize the initial reward function to obtain the reward function of the current target model.

7. The satellite networking routing device based on deep reinforcement learning according to claim 6, wherein, The determination module includes: A selection unit, configured to randomly select from each of the decision-making actions according to the first decision-making action selection probability and the first action value function by using the preset exploration and exploitation strategy, and output a corresponding second action value function according to the selected decision-making action; Alternatively, determine the highest action value function from each of the decision-making actions according to the second decision-making action selection probability and the first action value function by using the preset exploration and exploitation strategy, and output a corresponding second action value function according to the selected decision-making action.

8. An electronic device, characterized in that, It includes: A memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement a satellite network routing method based on deep reinforcement learning according to any one of claims 1-5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to be used to implement a satellite network routing method based on deep reinforcement learning according to any one of claims 1-5.

Citation Information

Patent Citations

  • Multi-agent-based low-orbit satellite network routing decision-making method and device

    CN117614882A