A Routing Method, Device and System for a Space-Ground Integrated Network
The DDQN-based routing method optimizes satellite-terrestrial networks by selecting optimal satellite nodes for efficient data transmission, addressing storage and resource issues while adapting to dynamic topologies.
Patent Information
- Application Number
- CN202411229003.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-03
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-09-03
AI Technical Summary
In a satellite-ground converged network, as the number of satellites increases, satellites need to store a large number of routing tables and consume a large number of signaling resources. The existing routing tables cannot be flexibly adjusted to adapt to topological changes and cannot meet communication needs.
The dual-deep Q network (DDQN) model is used to filter the starting and target satellites based on the satellite-ground link quality and satellite network status, and the data forwarding path is decided through the DDQN model to achieve end-to-end data transmission.
It improves the quality of service (QoS) performance of the satellite-earth converged network, optimizes the data transmission path, reduces signaling resource consumption, and adapts to network topology changes.
Smart Images

Figure CN118869054B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of satellite communication technologies, and in particular, to a routing method, apparatus, and system for a satellite-ground integrated network. Background Art
[0002] In the process of data transmission and forwarding in a satellite-ground integrated network, related technologies divide the orbital period of a satellite into several time slices, generate a routing table based on the topological snapshot of the satellite network in each time slice, and upload it to each satellite. Each satellite forwards data according to the saved routing table. At the start time of the next time slice, the next routing table is uploaded to each satellite.
[0003] As the number of satellites in the satellite network continues to increase, each satellite needs to store a large number of routing tables, and continuously uploading routing tables will also consume a large amount of signaling resources. In addition, satellites are constantly moving, and the topology of the satellite network will change accordingly. However, the routing table is calculated offline after obtaining the constellation topology snapshot. After the actual constellation topology snapshot changes, the calculated routing table cannot be flexibly adjusted, and it cannot meet the current communication requirements for satellite-ground integrated networks. Summary of the Invention
[0004] The present disclosure provides a routing method, apparatus, and system for a satellite-ground integrated network, which can improve the service quality of the satellite-ground integrated network.
[0005] The technical solution of the present disclosure is implemented as follows:
[0006] In a first aspect, the present disclosure provides a routing method for a satellite-ground integrated network. The routing method includes:
[0007] Screen out corresponding starting satellites and target satellites for a sending terminal and a receiving terminal respectively according to satellite-ground link communication quality indicators;
[0008] Receive data to be transmitted sent by the sending terminal through the starting satellite;
[0009] Starting from the starting satellite, a current satellite decides the next satellite to forward based on the satellite network state through a trained double deep Q-network (DDQN) model, and sends the data to be transmitted to the next satellite to forward until the data to be transmitted is sent to the target satellite;
[0010] Transmit the data to be transmitted to the receiving terminal through the target satellite.
[0011] In a second aspect, the present disclosure provides a routing apparatus for a satellite-ground integrated network. The routing apparatus includes a screening part, a satellite-ground transmission part, and an inter-satellite transmission part. Among them,
[0012] The screening part is configured to screen out corresponding starting satellites and target satellites for the sending terminal and the receiving terminal respectively according to the satellite-ground link communication quality indicators;
[0013] The satellite-ground transmission part is configured to receive the data to be transmitted sent by the sending terminal through the starting satellite;
[0014] The inter-satellite transmission part is configured to start from the starting satellite, and based on the satellite network status, the current satellite decides the next satellite to forward through the trained double deep Q network (DDQN) model, and sends the data to be transmitted to the next satellite to be forwarded until the data to be transmitted is sent to the target satellite;
[0015] The inter-satellite transmission part is configured to transmit the data to be transmitted to the receiving terminal through the target satellite.
[0016] In a third aspect, the present disclosure provides a routing system for a satellite-ground integrated network. The routing system includes a ground station and a satellite network composed of multiple satellites. Among them,
[0017] The sending terminal and the receiving terminal in the ground station are used to screen out corresponding starting satellites and target satellites in the satellite network respectively according to the satellite-ground link communication quality indicators;
[0018] The starting satellite in the satellite network is used to receive the data to be transmitted sent by the sending terminal;
[0019] The satellite network is used to start from the starting satellite, and based on the satellite network status, the current satellite decides the next satellite to forward through the trained double deep Q network (DDQN) model, and sends the data to be transmitted to the next satellite to be forwarded until the data to be transmitted is sent to the target satellite;
[0020] The target satellite in the satellite network is used to transmit the data to be transmitted to the receiving terminal.
[0021] The present disclosure provides a routing method, device and system for a satellite-ground integrated network; selects appropriate satellite nodes for satellite-ground transmission in the integrated network according to the satellite-ground link quality, and during the process of data forwarding using the satellite network, determines the next satellite to forward through the DDQN model based on the satellite network status, which can consider the end-to-end data transmission requirements of the integrated network starting from ground users, and comprehensively improves the QoS performance of the integrated network. Description of the Drawings
[0022] Figure 1 It is a schematic diagram of a satellite-ground integrated network architecture provided by the present disclosure.
[0023] Figure 2 Schematic flowchart of a routing method for a satellite - terrestrial integrated network provided by the present disclosure.
[0024] Figure 3 Schematic diagram of a set of satellites covering a service cell provided by the present disclosure.
[0025] Figure 4 Schematic diagram of the training of the DDQN model provided by the present disclosure.
[0026] Figure 5 Schematic diagram of the composition of a routing device for a satellite - terrestrial integrated network provided by the present disclosure.
[0027] Figure 6 Schematic diagram of the composition of another routing device for a satellite - terrestrial integrated network provided by the present disclosure.
[0028] Figure 7 Schematic diagram of the hardware structure of a computing device provided by the present disclosure. Detailed implementation manners
[0029] Next, the technical solutions in the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the present disclosure.
[0030] Refer to Figure 1 , which shows a schematic diagram of a satellite - terrestrial integrated network architecture applicable to the present disclosure. In Figure 1 , the ground terminals include a sending terminal and a receiving terminal. In some examples, the ground terminal may specifically be at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, a virtual reality terminal, an augmented reality terminal, a wireless terminal, and a laptop computer. The ground terminal has a communication function and can access a wired network, a wireless network, and a satellite network.
[0031] Continue to refer to Figure 1 , taking the example of dividing the earth's surface into service cells of the same size according to hexagonal grids, each grid serves as a service cell, which contains the ground terminals within the covered range and the satellites whose beam coverage can cover the grid. In Figure 1 , taking the low - earth - orbit giant constellation as an example for the satellite network, each satellite serves as a node of the satellite network, can establish inter - satellite communication links with the two adjacent satellites in the same orbit before and after, and can also establish inter - satellite communication links with the adjacent satellites in the adjacent orbits. The data transmission between two satellites through the inter - satellite communication link can be called "one - hop". In the case where the data sent by the sending terminal shown in Figure 1 is transmitted to the receiving terminal through the satellite network, in the satellite network, exemplarily, it can be according to Figure 1The inter-satellite communication link shown by the dashed line in the figure is used for transmission. In this case, data will be transmitted and relayed among multiple satellites, which can be called "multi-hop".
[0032] It should be noted that during the data transmission process, after each satellite receives the data to be relayed, it is necessary to determine the next-hop relay path for it, usually by querying the routing table to select the relay path. In the related solutions, the routing table is calculated offline after obtaining the constellation topology snapshot. After the actual constellation topology snapshot changes, the calculated routing table cannot be flexibly adjusted, and thus cannot meet the current communication requirements for the space-ground integrated network.
[0033] In order to improve the service quality of the space-ground integrated network, combined with Figure 1 the example of the space-ground integrated network architecture shown in the figure, the present disclosure provides a routing method for the space-ground integrated network. Refer to Figure 2 , the routing method includes steps S201 to S204.
[0034] In step S201, the corresponding starting satellite and target satellite are selected for the sending terminal and the receiving terminal respectively according to the space-ground link communication quality index.
[0035] In the present disclosure, combined with Figure 1 the network architecture shown in the figure, the sending terminal and the receiving terminal on the ground are respectively in their respective service cells, as shown in the hexagonal grid in Figure 1 . For each grid, due to the large number of satellites in the giant constellation, there will be multiple satellite beams covering the grid at the same time. In the related solutions, usually a satellite is randomly selected from these satellites to receive the data sent by the sending terminal on the ground or send data to the receiving terminal on the ground. In order to improve the overall service quality of the space-ground integrated network, the present disclosure screens the satellites receiving the ground terminals and the satellites sending data to the ground terminals through constraint conditions and communication quality, so as to select the optimal satellite in the satellites covering the same service cell (grid) for data transmission with the ground terminal.
[0036] In some examples, the step of selecting the corresponding starting satellite and target satellite for the sending terminal and the receiving terminal respectively according to the space-ground link communication quality index includes:
[0037] Obtain a first satellite set that can cover the grid where the sending terminal is located and a second satellite set that can cover the grid where the receiving terminal is located;
[0038] In the first satellite set and the second satellite set, respectively screen out a plurality of first candidate satellites and a plurality of second candidate satellites that meet the set constraint conditions;
[0039] Among the first candidate satellite and the second candidate satellite, the space-ground link communication quality index LQI of each first candidate satellite and each second candidate satellite is obtained according to the satellite's transmit power, transmission loss, antenna gain, center frequency of communication, and the distance between the receiving terminal and the transmitting terminal;
[0040] The first candidate satellite and the second candidate satellite with the largest LQI are respectively determined as the starting satellite and the target satellite.
[0041] For the above example, specifically, taking Figure 3 as an example, the beams of satellites Sa1, Sa2, and Sa3 can all cover the service cell shown by the hexagonal grid. That is to say, these three satellites can all provide communication services for the terminals in this service cell. For example, receiving data sent by the terminals or sending data to the terminals.
[0042] In some exemplary implementation processes, taking Figure 3 the terminal in as the transmitting terminal as an example, these three satellites form the first satellite set. The present disclosure can exemplarily set constraint conditions according to the signal-to-noise ratio, free space attenuation, and bandwidth limitations, and respectively evaluate whether the three satellites in the first satellite set simultaneously meet these constraint conditions. In the first satellite set, the satellites that simultaneously meet these constraint conditions are screened out as the first candidate satellites. For example, in the satellite set formed by satellites Sa1, Sa2, and Sa3, according to the above three constraint conditions, it is screened out that both satellites Sa1 and Sa2 simultaneously meet these constraint conditions, then satellites Sa1 and Sa2 are taken as the first candidate satellites. For each first candidate satellite, the present disclosure takes the following formula as an example to evaluate the space-ground link communication quality to obtain the space-ground link communication quality index (LQI, Link Quality Indicator) corresponding to each first candidate satellite.
[0043]
[0044] In the above formula, P set represents the rated transmit power of the satellite, is the transmission loss, is the noise, G tr (ν i ) and G re (ν j ) are the gains of the satellite's transmit antenna and receive antenna respectively, c is the speed of light, is the communication center frequency, d t (v i , e j ) is the distance between the satellite's receive antenna and transmit antenna.
[0045] After obtaining the above satellite-ground link LQIs of each first candidate satellite, the candidate satellite corresponding to the maximum LQI is used as the starting satellite corresponding to the sending terminal.
[0046] Regarding the above implementation process, it should be noted that when Figure 3 the terminal in is a receiving terminal, Figure 3 the satellites in form the second satellite set described in the above example. According to the above implementation process, second candidate satellites can be screened out in the second satellite set, and the target satellite corresponding to the receiving terminal can be determined based on the above satellite-ground link LQI. Details are not described herein in the present disclosure.
[0047] In step S202, the starting satellite receives the data to be transmitted sent by the sending terminal.
[0048] In the present disclosure, after determining the starting satellite corresponding to the sending terminal through the above step S201, the sending terminal sends the data to be transmitted that needs to be transmitted to the receiving terminal to the starting satellite. After receiving the data to be transmitted, the starting satellite determines a routing decision to forward the data to be transmitted to the target satellite corresponding to the receiving terminal through satellite network in step S203.
[0049] In step S203, starting from the starting satellite, the current satellite makes a decision on the next satellite to forward based on the satellite network status through the trained Double Deep Q-Learning (DDQN) model, and sends the data to be transmitted to the next satellite to forward until the data to be transmitted is sent to the target satellite.
[0050] In the present disclosure, starting from the starting satellite, each satellite that receives the data to be transmitted is used as the current satellite. The current satellite network status is input into the trained DDQN model, and the next satellite to forward (also referred to as the next-hop satellite) is determined based on the output of the DDQN model, and the data to be transmitted is sent to the next satellite to forward. Then, the next satellite to forward is used as the current satellite, and the current satellite network status is continuously input into the trained DDQN model to decide the next-hop satellite of the next satellite to forward until the data to be transmitted is forwarded to the target satellite.
[0051] Specifically, the DDQN model includes a target network and an estimation network, and the structures of the two networks are the same. In the actual application process after the DDQN model is trained, the estimation network is used to make decisions, and the target network can be used to update the weight parameters in the estimation network during the training process or the actual application process.
[0052] In some possible implementation manners, the current satellite determines the next satellite for forwarding based on the satellite network state through the trained DDQN model, including:
[0053] The current satellite inputs the satellite network state parameter values into the trained DDQN model, and the estimation network in the trained DDQN model obtains the state-action value (Q value) of the satellite node corresponding to the alternative direction according to the satellite network state parameter values;
[0054] The satellite node with the largest state-action value among all alternative directions is determined as the next satellite for forwarding.
[0055] For the above implementation manner, in some examples, the satellite network state parameter values input into the DDQN model include: the congestion state, link bandwidth, and delay of each satellite node in the satellite network, as well as the current satellite node and the destination node of the data to be transmitted. The estimation network can obtain the Q values of the satellite nodes corresponding to each alternative direction based on the satellite network state parameter values.
[0056] For each satellite, there are a total of 4 alternative directions. Combining with the architecture example in Figure 1 the four satellite nodes corresponding to these four directions are respectively the two adjacent front and rear satellites on the same orbit and the two satellites adjacent to the adjacent orbit. For the satellite nodes corresponding to these 4 alternative directions, the present disclosure evaluates the immediate reward according to the action of assuming that the data to be transmitted is respectively transmitted to these 4 satellite nodes, and obtains the Q values corresponding to these 4 nodes according to the immediate reward. The satellite node corresponding to the maximum Q value is determined as the next satellite for the current satellite to forward the data to be transmitted.
[0057] In some examples, the obtaining of the state-action value of the satellite node corresponding to the alternative direction by the estimation network in the trained DDQN model according to the satellite network state parameter values includes:
[0058] According to the satellite network state parameter values, the estimation network in the trained DDQN model obtains the immediate reward of the satellite node corresponding to each alternative direction; the immediate reward includes the rewards regarding transmission delay, link bandwidth, and transmission path improvement;
[0059] Generate the state-action value of the satellite node corresponding to each alternative direction according to the rewards regarding transmission delay, link bandwidth, and transmission path improvement.
[0060] For the above example, specifically, in the present disclosure, the Q value of the satellite nodes in the alternative directions is comprehensively evaluated from three aspects: transmission delay, network bandwidth, and transmission path, so that in the process of selecting the next satellite for forwarding from these satellite nodes, congested nodes can be avoided as much as possible, and higher bandwidth and shorter transmission paths can be pursued, comprehensively optimizing the QoS performance of the satellite network and realizing the efficient flow of data.
[0061] Regarding the transmission delay reward r lat It can be exemplarily represented by the following formula:
[0062]
[0063] where lat best represents the path delay of the shortest spatial distance, lat act is the actual delay assuming the selection of the satellite node in the alternative direction, and α is the discount factor, usually set as a constant greater than 0 and less than 1.
[0064] Regarding the reward r for link bandwidth band It can be exemplarily represented by the following formula:
[0065]
[0066] where γ is the discount factor, band act and band best are respectively the actual bandwidth of the satellite node assuming the selection of the alternative direction and the maximum bandwidth that the satellite network can achieve.
[0067] Regarding the reward r for transmission path improvement imp It can be obtained based on the first shortest distance from the current satellite to the target satellite and the second shortest distance from the next forwarding satellite to the target satellite, and can be exemplarily represented by the following formula:
[0068]
[0069] where the shortest distance from the current satellite to the target satellite is l1, the shortest distance from the next-hop satellite to the target satellite is l2, and r exp represents the desired reward value, usually a fixed value.
[0070] After obtaining the above immediate reward, the Q value of the satellite nodes corresponding to each alternative direction can be obtained according to the following formula:
[0071] Q(s,a) = r + γQ target (s′,a max )
[0072] Wherein, a represents an action, γ represents a discount factor, r represents the immediate reward obtained after executing the current action, s represents a state, s′ represents the next state to which the satellite network transfers after executing the action a, a max represents the maximum predicted reward of the target network for the next state, and Q target represents the target Q-value of the target network in the DDQN model.
[0073] In some examples, the above-mentioned immediate reward can be considered as the reward obtained after executing the current action. In some specific implementation processes, this immediate reward can be obtained by weighted summing the rewards for transmission delay, link bandwidth, and transmission path improvement described above, and then adding it to the basic reward r b to get. For example, it can be obtained by the following formula:
[0074] r = R(a, s) = r b + α1·r lat + α2·r band + α3·r imp
[0075] wherein, α1, α2, and α3 respectively represent the weights corresponding to the rewards for transmission delay, link bandwidth, and transmission path improvement.
[0076] In step S204, the data to be transmitted is transmitted to the receiving terminal through the target satellite.
[0077] Specifically, starting from the starting satellite, in the satellite network, the next satellite to forward is determined through the DDQN model until the data to be transmitted is forwarded to the target satellite. After that, the target satellite can send the data to be transmitted to the receiving terminal, thereby realizing the complete process of data transmission from the sending terminal to the receiving terminal based on the satellite-ground integrated network.
[0078] Through the above technical solution, the present disclosure selects the appropriate satellite nodes for satellite-ground transmission in the integrated network according to the satellite-ground link quality, and in the process of data forwarding using the satellite network, determines the next satellite to forward based on the satellite network state through the DDQN model, which can consider the end-to-end data transmission requirements of the integrated network starting from the ground user and comprehensively improve the QoS performance of the integrated network.
[0079] Based on the foregoing technical solution, in the specific implementation process, the entire satellite-ground integrated network can be used as a reinforcement learning intelligent agent, and the DDQN model in the intelligent agent is trained. As Figure 4 shown, the training process may include:
[0080] S401: Initialize the topological structure, link state, and node state of a training satellite network, and the DDQN model.
[0081] For this step, specifically, before the start of training, initialize the topology of the training satellite network for model training, such as the positions of all satellite nodes, and initialize the connection conditions of each inter-satellite link and the state of each satellite node. This ensures that the agent can be trained in a simulated real network environment, making the model more effective in practical applications.
[0082] For the DDQN model, two main neural networks, namely the estimation network and the target network, as well as the experience replay buffer can be initialized. This experience replay buffer is used to store the agent's past experiences (such as states, actions, rewards, next states) for subsequent training. Through the experience replay mechanism, the agent can learn from historical experiences, effectively reducing data correlation and improving the training effect. The initialization of the neural network can be to initialize the weight parameters in the network.
[0083] Exemplarily, after completing the above initialization process, the present disclosure sets the total number of training rounds in the training process. Each training round represents the agent completing a data transmission task from the sending terminal to the receiving terminal. After setting the total number of rounds N, each training round is trained according to the training process described in steps S402 to S408.
[0084] S402: In the i-th training round, based on the randomly selected sending terminal and receiving terminal, filter out the corresponding starting satellite and target satellite in the training satellite network according to the satellite-ground link communication quality index.
[0085] In the present disclosure, in each training round, the data transmission task is randomly formed, that is, the sending terminal and the receiving terminal are randomly generated. This randomly generated task helps the agent adapt to various possible transmission scenarios, thereby improving the generalization ability of the model.
[0086] Exemplarily, after randomly generating the sending terminal and the receiving terminal, the starting satellite and the target satellite of this training round can be obtained through the foregoing Figure 2 shown technical solution.
[0087] S403: In the i-th training round, initialize the satellite network state of the training satellite network.
[0088] Exemplarily, the satellite network state includes satellite node positions, the congestion state of each satellite node, the link bandwidth state, etc. By initializing these satellite network states, it can ensure that the agent makes reasonable decisions based on the current network situation.
[0089] S404: In the i-th training round, starting from the initial satellite, for each current satellite, input the satellite network state of the training satellite grid into the DDQN model, and obtain the satellite nodes corresponding to the alternative directions through the estimation network in the DDQN model.
[0090] Exemplarily, with the above satellite network state as the input value of the DDQN model, the estimated Q-values corresponding to all actions can be obtained through the estimation network of the DDQN model. In the present disclosure, an "action" can also be expressed as the satellite nodes corresponding to the alternative directions where data forwarding needs to be performed.
[0091] S405: In the i-th training round, determine the next satellite for forwarding through the ε-greedy strategy from the satellite nodes corresponding to the alternative directions.
[0092] Exemplarily, in order to enable the agent to not only make better decisions using existing knowledge but also explore new paths to discover better solutions. During the training process, the satellite node with the maximum Q-value is not used as the next satellite for forwarding in the data forwarding process of this training round. Instead, the ε-greedy strategy is adopted to select the next satellite for forwarding, so that the node with the currently estimated maximum Q-value can be selected most of the time, but there is also a certain probability of selecting other actions to explore new paths.
[0093] S406: In the i-th training round, obtain the immediate reward regarding transmission delay, link bandwidth, and transmission path improvement, as well as the satellite network state of the next step, based on the next satellite for forwarding.
[0094] Exemplarily, after determining the next satellite for forwarding and forwarding the data to this satellite, it can be considered that the "action" is completed, and based on this completed "action", the satellite network enters a new state. Therefore, the immediate reward can be calculated according to this "action" and the new state. Similar to the foregoing technical solution, the immediate reward can include rewards regarding transmission delay, link bandwidth, and transmission path improvement.
[0095] It should be noted that in this training round, by continuously performing actions and state changes, the agent gradually approaches the optimal path and adjusts its decisions in a timely manner to adapt to changes in the network environment.
[0096] S407: In the i-th training round, generate an experience based on the satellite network state of the current satellite, the next satellite for forwarding of the current satellite, the immediate reward, and the satellite network state of the next satellite for forwarding, and store it in the experience replay buffer.
[0097] Exemplarily, the current satellite network state, actions, immediate rewards, and new states are stored in the experience replay buffer, and training data is continuously accumulated for subsequent updates. In this way, the experience replay mechanism enables the agent to learn from diverse data, reducing data dependence and bias in training, thereby enhancing the stability of the model.
[0098] S408: In the i-th training episode, generate the target Q-values of the target network in the DDQN model based on the experiences in the experience replay buffer, and update the weight parameters in the DDQN model through the estimated Q-values of the estimation network and the target Q-values of the target network.
[0099] Exemplarily, use the target network to evaluate the new state and calculate the target Q-values of the next "action", which are used to estimate the actual value of the current action.
[0100] After obtaining the target Q-values, calculate the deviation between the target Q-values and the estimated Q-values, and update the weight parameters in the estimation network based on this deviation through the backpropagation algorithm, so that the updated estimation network outputs estimated Q-values that are closer to the target Q-values.
[0101] In some examples, after the set number of training episodes ends, or after the set number of "hops" have been estimated, the weight parameters of the estimation network can be copied to the target network to keep the calculation of the target Q-values stable, prevent the target values from fluctuating frequently during training, and thus improve the training effect of the entire model.
[0102] S409: Determine whether the current training episode number i satisfies the set total number of training episodes N; if so, end the training process and obtain the trained DDQN model, otherwise, return to S402 and execute the (i + 1)-th training episode.
[0103] Exemplarily, during the above training process, the method further includes:
[0104] In the i-th training episode, when the determined next satellite for forwarding is congested or is a satellite node that has been repeatedly determined, terminate the current training episode;
[0105] When the determined next satellite for forwarding is the target satellite, obtain the total reward of the DDQN model based on the weight parameters in the DDQN model and the immediate reward, and terminate the training process.
[0106] For the above example, in any training round, when the next satellite to be forwarded determined by the DDQN model is congested or has already been determined, it means that the data to be transmitted in this training round cannot be transmitted to the target satellite. Therefore, it is necessary to terminate the current training round and enter the next training round. When the next satellite to be forwarded is the target satellite, it means that the data to be transmitted in this training round has reached the target satellite, and the current training round ends. Terminate the current training round and enter the next training round. The experience, rewards, and model weight parameters learned in the current training round are retained for continued training in the next training round.
[0107] Based on the same inventive concept as the foregoing technical solution, refer to Figure 5 , which shows a routing device 50 of a satellite-ground integrated network provided by the present disclosure. The routing device 50 includes: a screening part 501, a satellite-ground transmission part 502, and an inter-satellite transmission part 503; wherein,
[0108] The screening part 501 is configured to screen out corresponding starting satellites and target satellites for the sending terminal and the receiving terminal respectively according to the satellite-ground link communication quality index;
[0109] The satellite-ground transmission part 502 is configured to receive the data to be transmitted sent by the sending terminal through the starting satellite;
[0110] The inter-satellite transmission part 503 is configured to start from the starting satellite, and the current satellite decides the next satellite to be forwarded through the trained double deep Q network (DDQN) model based on the satellite network state, and send the data to be transmitted to the next satellite to be forwarded until the data to be transmitted is sent to the target satellite;
[0111] The inter-satellite transmission part 503 is configured to transmit the data to be transmitted to the receiving terminal through the target satellite.
[0112] In some examples, the inter-satellite transmission part 503 is configured to:
[0113] The current satellite inputs the satellite network state parameter value into the trained DDQN model, and the estimation network in the trained DDQN model obtains the state-action value of the satellite node corresponding to the alternative direction according to the satellite network state parameter value;
[0114] Determine the satellite node with the largest state-action value among all alternative directions as the next satellite to be forwarded.
[0115] In some examples, the satellite network state parameter values include: the congestion state, link bandwidth, and delay of each satellite node in the satellite network, as well as the current satellite node and the destination node of the data to be transmitted.
[0116] In some examples, the inter-satellite transmission section 503 is configured to:
[0117] According to the satellite network state parameter values, obtain the immediate reward of the satellite node corresponding to each alternative direction through the estimation network in the trained DDQN model; the immediate reward includes rewards regarding transmission delay, link bandwidth, and transmission path improvement;
[0118] Generate the state-action value of the satellite node corresponding to each alternative direction according to the rewards regarding transmission delay, link bandwidth, and transmission path improvement.
[0119] In some examples, the screening section 501 is configured to:
[0120] Obtain a first satellite set that can cover the grid where the sending terminal is located and a second satellite set that can cover the grid where the receiving terminal is located;
[0121] In the first satellite set and the second satellite set, respectively screen out a plurality of first candidate satellites and a plurality of second candidate satellites that meet the set constraint conditions;
[0122] In the first candidate satellites and the second candidate satellites, obtain the link quality indicator LQI of each first candidate satellite and each second candidate satellite according to the transmit power, transmission loss, antenna gain, center frequency of communication, and the distance between the receiving terminal and the sending terminal;
[0123] Determine the first candidate satellite and the second candidate satellite with the largest LQI as the starting satellite and the target satellite, respectively.
[0124] In some examples, refer to Figure 6 , the routing device 50 further includes a training section 504, which is configured to:
[0125] Initialize the topology structure, link state, and node state of a training satellite network, as well as the DDQN model;
[0126] In each training round, based on randomly selected sending and receiving terminals, screen out the corresponding starting satellite and target satellite in the training satellite network according to the link quality indicator of the satellite-ground link;
[0127] In each training round, initialize the satellite network state of the training satellite network;
[0128] In each training episode, starting from the initial satellite, for each current satellite, the satellite network state of the training satellite grid is input into the DDQN model, and the satellite nodes corresponding to the alternative directions are obtained through the estimation network in the DDQN model;
[0129] In each training episode, the satellite for the next hop is determined from the satellite nodes corresponding to the alternative directions through the ε-greedy strategy;
[0130] In each training episode, an immediate reward regarding transmission delay, link bandwidth, and transmission path improvement, as well as the satellite network state of the next step, is obtained based on the satellite for the next hop;
[0131] In each training episode, an experience is generated based on the satellite network state of the current satellite, the satellite for the next hop of the current satellite, the immediate reward, and the satellite network state of the next hop satellite, and is stored in the experience replay buffer;
[0132] In each training episode, the target Q-value of the target network in the DDQN model is generated based on the experiences in the experience replay buffer, and the weight parameters in the DDQN model are updated through the estimated Q-value of the estimation network and the target Q-value of the target network.
[0133] In some examples, the training section 504 is further configured to:
[0134] Obtain the immediate reward for the transmission path improvement based on the first shortest distance from the current satellite to the target satellite and the second shortest distance from the satellite for the next hop to the target satellite.
[0135] In some examples, the training section 504 is further configured to:
[0136] In each training episode, when the determined satellite for the next hop is congested or is a satellite node that has been repeatedly determined, this training episode is terminated;
[0137] When the determined satellite for the next hop is the target satellite, the total reward of the DDQN model is obtained based on the weight parameters in the DDQN model and the immediate reward, and the training process is terminated.
[0138] Please refer to Figure 7, which shows a structural block diagram of a computing device provided by an exemplary embodiment of the present disclosure. In some examples, the computing device 70 may be at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, a virtual reality terminal, an augmented reality terminal, a wireless terminal, and a laptop portable computer. The computing device 70 has a communication function and can access a wired network or a wireless network. The computing device 70 may generally refer to one of multiple terminals, and those skilled in the art can know that the number of the above terminals may be more or less. In some examples, the computing device 70 may receive data based on the accessed wired network, wireless network, or satellite network. It can be understood that the computing device 70 undertakes the computing and processing work of the technical solution of the present disclosure, and the present disclosure does not limit this.
[0139] As Figure 7 shown, the computing device in the present disclosure may include one or more of the following components: a processor 710 and a memory 720.
[0140] Optionally, the processor 710 uses various interfaces and lines to connect various parts within the entire computing device. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 720, and by calling data stored in the memory 720, it executes various functions of the computing device and processes data. Optionally, the processor 710 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 710 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), a neural-network processing unit (NPU), and a baseband chip. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the touch display screen; the NPU is used to implement artificial intelligence (AI) functions; the baseband chip is used to process wireless communication. It can be understood that the above baseband chip may not be integrated into the processor 710 and may be implemented separately by a single chip.
[0141] The memory 720 may include a Random Access Memory (RAM), or may also include a Read-Only Memory (ROM). Optionally, the memory 720 includes a non-transitory computer-readable storage medium. The memory 720 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 720 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above various method embodiments, etc.; the data storage area may store data created according to the use of the computing device, etc.
[0142] In addition, those skilled in the art can understand that the structure of the computing device shown in the above drawings does not constitute a limitation on the computing device. The computing device may include more or fewer components than shown in the drawings, or combine some components, or have different component arrangements. For example, the computing device also includes components such as a display screen, a camera component, a microphone, a speaker, a radio frequency circuit, an input unit, sensors (such as an acceleration sensor, an angular velocity sensor, a light sensor, etc.), an audio circuit, a WiFi module, a power supply, a Bluetooth module, etc., which will not be elaborated here.
[0143] The present disclosure also provides a computer-readable storage medium storing at least one instruction for being executed by a processor to implement the routing method of the satellite-ground integrated network as described in the above various embodiments.
[0144] The present disclosure also provides a computer program product including computer instructions stored in a computer-readable storage medium; a processor of a computing device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions so that the computing device executes to implement the routing method of the satellite-ground integrated network as described in the above various embodiments.
[0145] Based on the same inventive concept as the foregoing technical solution, in combination with Figure 1 , which shows a routing system of a satellite-ground integrated network provided by the present disclosure. The routing system includes a ground station and a satellite network composed of multiple satellites; wherein,
[0146] The sending terminal and the receiving terminal in the ground station are used to respectively screen out the corresponding starting satellite and target satellite in the satellite network according to the satellite-ground link communication quality index;
[0147] The starting satellite in the satellite network is used to receive the data to be transmitted sent by the sending terminal;
[0148] The satellite network is used to start from the starting satellite. Based on the satellite network status, the current satellite makes a decision on the next satellite to forward through the trained double deep Q-network (DDQN) model, and sends the data to be transmitted to the next satellite to be forwarded until the data to be transmitted is sent to the target satellite;
[0149] The target satellite in the satellite network is used to transmit the data to be transmitted to the receiving terminal.
[0150] It should be noted that for the implementation process of screening the starting satellite and the target satellite for the sending terminal and the receiving terminal respectively, the detailed process of routing and forwarding in the satellite network, and the training process of the DDQN model, reference can be made to the corresponding content in the foregoing technical solutions, and the present disclosure will not elaborate on this.
[0151] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present disclosure can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium accessible by a general-purpose or special-purpose computer.
[0152] It should be noted that: among the technical solutions recorded in the present disclosure, they can be combined arbitrarily without conflict.
[0153] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all of them should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A routing method for a satellite-terrestrial integrated network, characterized in that, The routing method includes: Based on the satellite-ground link communication quality indicators, corresponding starting satellites and target satellites are screened out for the sending terminal and the receiving terminal respectively; Receive the data to be transmitted sent by the sending terminal through the starting satellite; starting from the starting satellite, the current satellite makes a decision on the next satellite to forward based on the satellite network state through the trained Double Deep Q-Network (DDQN) model, and sends the data to be transmitted to the next satellite to be forwarded until the data to be transmitted is sent to the target satellite; Transmit the data to be transmitted to the receiving terminal through the target satellite; The method further includes: Initialize the topological structure, link state, and node state of a training satellite network, as well as the DDQN model; In each training episode, based on randomly selected sending and receiving terminals, corresponding starting satellites and target satellites are screened out in the training satellite network according to the satellite-ground link communication quality indicators; In each training episode, initialize the satellite network state of the training satellite network; In each training episode, starting from the starting satellite, for each current satellite, input the satellite network state of the training satellite network into the DDQN model, and obtain the satellite nodes corresponding to the alternative directions through the estimation network in the DDQN model; In each training episode, determine the next satellite to forward from the satellite nodes corresponding to the alternative directions through the ε-greedy strategy; In each training episode, obtain the immediate reward regarding transmission delay, link bandwidth, and transmission path improvement, as well as the satellite network state of the next step, based on the next satellite to forward; In each training episode, generate an experience based on the satellite network state of the current satellite, the next satellite to forward of the current satellite, the immediate reward, and the satellite network state of the next satellite to forward, and store it in the experience replay buffer; In each training episode, generate the target Q value of the target network in the DDQN model based on the experience in the experience replay buffer, and update the weight parameters in the DDQN model through the estimated Q value of the estimation network and the target Q value of the target network.
2. The routing method according to claim 1, characterized in that The current satellite makes a decision on the next satellite to forward based on the satellite network state through the trained Double Deep Q-Network (DDQN) model, including: The current satellite inputs the satellite network state parameter values into the trained DDQN model, and obtains the state-action value of the satellite nodes corresponding to the alternative directions through the estimation network in the trained DDQN model according to the satellite network state parameter values; Determine the satellite node with the maximum state-action value among all alternative directions as the next satellite to forward.
3. The method according to claim 2, wherein The satellite network state parameter values include: the congestion state, link bandwidth, and delay of each satellite node in the satellite network, as well as the current satellite node and the destination node of the data to be transmitted.
4. The method according to claim 2, wherein The obtaining of the state-action value of the satellite nodes corresponding to the alternative directions through the estimation network in the trained DDQN model according to the satellite network state parameter values includes: According to the satellite network state parameter values, obtain the immediate rewards of the satellite nodes corresponding to each alternative direction through the estimation network in the trained DDQN model; The immediate rewards include rewards regarding transmission delay, link bandwidth, and transmission path improvement; Generate the state-action values of the satellite nodes corresponding to each alternative direction according to the rewards regarding transmission delay, link bandwidth, and transmission path improvement.
5. The method according to claim 1, wherein The screening of the corresponding starting satellite and target satellite for the sending terminal and the receiving terminal respectively according to the satellite-ground link communication quality index includes: Obtain a first satellite set that can cover the grid where the sending terminal is located and a second satellite set that can cover the grid where the receiving terminal is located; In the first satellite set and the second satellite set, respectively screen out a plurality of first candidate satellites and a plurality of second candidate satellites that meet the set constraint conditions; Among the first candidate satellites and the second candidate satellites, obtain the satellite-ground link communication quality index LQI of each first candidate satellite and each second candidate satellite according to the transmit power, transmission loss, antenna gain, center frequency of communication, and the distance between the receiving terminal and the sending terminal of the satellite; Determine the first candidate satellite and the second candidate satellite with the largest LQI as the starting satellite and the target satellite respectively.
6. The method according to claim 1, wherein The method further includes: Obtain the immediate reward for the transmission path improvement according to the first shortest distance from the current satellite to the target satellite and the second shortest distance from the next-hop forwarding satellite to the target satellite.
7. The method according to claim 1, wherein The method further includes: In each training episode, when the determined next-hop forwarding satellite is congested or is a repeatedly determined satellite node, terminate this training episode; When the determined next-hop forwarding satellite is the target satellite, obtain the total reward of the DDQN model according to the weight parameters in the DDQN model and the immediate reward, and terminate the training process.
8. A routing device for a space-ground integrated network, characterized in that The routing device includes: a screening part, a satellite-ground transmission part, an inter-satellite transmission part, and a training part; wherein, The screening part is configured to screen out the corresponding starting satellite and target satellite for the sending terminal and the receiving terminal respectively according to the satellite-ground link communication quality index; The satellite-ground transmission part is configured to receive the data to be transmitted sent by the sending terminal through the starting satellite; The inter-satellite transmission part is configured to start from the starting satellite, and the current satellite decides the next-hop forwarding satellite based on the satellite network state through the trained double deep Q network DDQN model, and send the data to be transmitted to the next-hop forwarding satellite until the data to be transmitted is sent to the target satellite; The inter-satellite transmission part is configured to transmit the data to be transmitted to the receiving terminal through the target satellite; The training part is configured to: Initialize the topology structure, link state, and node state of a training satellite network, and the DDQN model; In each training episode, based on randomly selected sending terminals and receiving terminals, screen out the corresponding starting satellite and target satellite in the training satellite network according to the satellite-ground link communication quality index; In each training episode, initialize the satellite network state of the training satellite network; In each training episode, starting from the starting satellite, for each current satellite, input the satellite network state of the training satellite network into the DDQN model, and obtain the satellite nodes corresponding to the alternative directions through the estimation network in the DDQN model; In each training episode, determine the next satellite to forward through the ε-greedy strategy from the satellite nodes corresponding to the alternative directions; In each training episode, obtain the immediate reward regarding transmission delay, link bandwidth, and transmission path improvement, as well as the satellite network state of the next step, based on the next satellite to forward; In each training episode, generate an experience based on the satellite network state of the current satellite, the next satellite to forward of the current satellite, the immediate reward, and the satellite network state of the next satellite to forward, and store it in the experience replay buffer; In each training episode, generate the target Q value of the target network in the DDQN model based on the experiences in the experience replay buffer, and update the weight parameters in the DDQN model through the estimated Q value of the estimation network and the target Q value of the target network.
9. A routing system for a space-ground integrated network, characterized in that, The routing system includes a ground station and a satellite network composed of multiple satellites; among them, The ground station includes a sending terminal and a receiving terminal, which are used to screen out the corresponding starting satellite and target satellite in the satellite network respectively according to the satellite-ground link communication quality index; The starting satellite in the satellite network is used to receive the data to be transmitted sent by the sending terminal; The satellite network is used to start from the starting satellite. Based on the satellite network state, the current satellite decides the next satellite to forward through the trained double deep Q network (DDQN) model, and sends the data to be transmitted to the next satellite to forward until the data to be transmitted is sent to the target satellite; The target satellite in the satellite network is used to transmit the data to be transmitted to the receiving terminal; The training process of the double deep Q network (DDQN) model includes: Initialize the topological structure, link state, and node state of a training satellite network, as well as the DDQN model; In each training episode, based on randomly selected sending and receiving terminals, screen out the corresponding starting satellite and target satellite in the training satellite network according to the satellite-ground link communication quality index; In each training episode, initialize the satellite network state of the training satellite network; In each training episode, starting from the starting satellite, for each current satellite, input the satellite network state of the training satellite network into the DDQN model, and obtain the satellite nodes corresponding to the alternative directions through the estimation network in the DDQN model; In each training episode, determine the next satellite to forward through the ε-greedy strategy from the satellite nodes corresponding to the alternative directions; In each training episode, obtain the immediate reward regarding transmission delay, link bandwidth, and transmission path improvement, as well as the satellite network state of the next step, based on the next satellite to forward; In each training episode, an experience is generated based on the satellite network status of the current satellite, the satellite to which the current satellite will forward in the next step, the immediate reward, and the satellite network status of the satellite to which it will forward in the next step, and is stored in the experience replay buffer; In each training episode, the target Q-value of the target network in the DDQN model is generated based on the experiences in the experience replay buffer, and the weight parameters in the DDQN model are updated by the estimated Q-value of the estimation network and the target Q-value of the target network.
Citation Information
Patent Citations
Multi-agent-based low-orbit satellite network routing decision-making method and device
CN117614882A
Method and device for recommending satellite communication scheme
CN117879676A