SDN-based air-sea cross-domain network reinforcement learning routing algorithm

By adopting reinforcement learning routing algorithm based on SDN in the air-sea cross-domain network, the problems of insufficient adaptability of dynamic environments and weak QoS requirements in the existing technology are solved, and efficient and intelligent routing decisions for the air-sea cross-domain network are realized, which significantly improves the comprehensive performance and cross-domain collaborative efficiency of the network.

CN119996290APending Publication Date: 2025-05-13HARBIN ENGINEERING UNIVERSITY SANYA NANHAI INNOVATION & DEVELOPMENT BASE

Patent Information

Application Number
CN202510457291.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing air-sea cross-domain network routing algorithms are insufficient in adaptability to dynamic environments and weak in the protection of diversified QoS requirements due to static rules or fixed optimization models, making it difficult to meet the real-time data transmission and cross-domain collaboration needs of marine exploration tasks.

Method used

Using the reinforcement learning routing algorithm of air-sea cross-domain network based on SDN, a dynamic routing decision model is designed through the combination of Markov decision chain system modeling, soft maximum strategy action selection and SARSA model to achieve adaptive matching of dynamic network state perception and modeling and differentiated QoS requirements.

Benefits of technology

It significantly improves the comprehensive performance of the air-sea cross-domain network, can sense network topology changes and task requirements in real time, adaptively optimize routing strategies, meet diversified QoS needs, reduce end-to-end delays, and enhance cross-domain collaboration efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996290A_ABST
    Figure CN119996290A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of wireless communication, discloses an air-sea cross-domain network reinforcement learning routing algorithm based on an SDN (Software Defined Network), and mainly aims to solve the problems of poor dynamic environment adaptability, insufficient multi-task QoS (Quality of Service) demand guarantee capability and the like caused by a static rule or local optimization of an existing air-sea cross-domain routing algorithm. Global network state perception and resource centralized scheduling are realized through an SDN controller, and a dynamic routing decision model driven by deep reinforcement learning (DRL) is constructed. A network state is sensed in real time through an SDN controller, a multi-target reward function is designed, and a cross-domain path selection strategy is dynamically optimized. According to the method, the limitation of a traditional routing protocol in a cross-domain heterogeneous network is broken through, the end-to-end transmission efficiency, the resource utilization rate and the multi-task differentiated service quality guarantee capability are remarkably improved, and the method can be widely applied to the scenes of marine environment monitoring, emergency rescue, cross-domain cooperative detection and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technology, and in particular to an air-sea cross-domain network reinforcement learning routing algorithm based on SDN. Background Art

[0002] With the rapid development of marine resource development, environmental monitoring, emergency rescue and other fields, the demand for real-time data transmission and cross-domain collaboration in marine exploration activities is becoming increasingly urgent. Heterogeneous equipment such as underwater robots, buoy sensors, surface ships and low-altitude drones need to achieve multi-level collaboration between air, sea and submarine, which places extremely high demands on the coverage capability, dynamic adaptability and quality of service (QoS) of the communication network. The traditional single network architecture is limited by its coverage range and dynamic adaptability, and it is difficult to meet the needs of cross-domain collaboration. The air-sea cross-domain network integrates satellites, drones, surface relay nodes and underwater sensor units to form three-dimensional communication coverage and resource complementarity, providing a highly reliable and real-time transmission channel for marine exploration missions, significantly improving the efficiency of task execution in complex scenarios.

[0003] However, the existing air-sea cross-domain network routing algorithms are mostly based on static rules or traditional optimization models (such as shortest path, load balancing), which are difficult to flexibly respond to dynamic network environments and diversified task requirements. On the one hand, the high mobility of air-sea nodes, the time-varying nature of link quality, and the heterogeneity of cross-domain topology make it difficult for traditional algorithms to capture network status changes in real time; on the other hand, ocean exploration tasks cover scenarios such as real-time video backhaul, emergency command issuance, and large-capacity data collection. Different tasks have significant differences in QoS requirements for latency, bandwidth, and packet loss rate. Existing algorithms lack the ability to adjust dynamic strategies, which often leads to problems such as rigid resource allocation and uneven service quality, which seriously restricts the overall performance of cross-domain networks.

[0004] Therefore, the present invention proposes an SDN-based air-sea cross-domain network reinforcement learning routing algorithm to solve the above technical problems. Summary of the invention

[0005] In order to solve the technical problems of insufficient adaptability to dynamic environments and weak ability to guarantee diversified QoS requirements caused by static rules or fixed optimization models in existing air-sea cross-domain network routing algorithms, the present invention provides an air-sea cross-domain network reinforcement learning routing algorithm based on SDN. The present invention integrates the global network control capability of SDN and the dynamic decision-making advantage of reinforcement learning to achieve dynamic network state perception and modeling, adaptive matching of differentiated QoS requirements, and intelligent decision-making in highly dynamic environments, thereby improving the efficiency of cross-domain collaborative tasks.

[0006] The present invention is achieved through the following technical solutions: The invention provides an SDN-based air-sea cross-domain network reinforcement learning routing algorithm, which includes a reinforcement learning action selection strategy and a routing protocol description.

[0007] The reinforcement learning action selection strategy under the SDN architecture is as follows: First, a Markov decision chain is used to provide a mathematical framework for system modeling of reinforcement learning, represented by a multivariable group (S, A, P, R), where S represents a finite state set, A represents a finite action set, P represents a state transition probability set, and R represents a reward set. The correct action selection strategy and state-action quality function are selected to ensure that the best long-term return is obtained while accelerating convergence, thereby ensuring QoS satisfaction rate.

[0008] The action selection strategy is responsible for specifying the action mode of the agent and mapping the state to the action. The ideal action selection strategy should be to select high-benefit actions with high probability and low-benefit or negative-benefit actions with low probability. Currently, there are two widely used action selection strategies, greedy and softmax. Under the greedy strategy, the action with the highest benefit is always selected at each step, and only the past data based on the current state and the quality function are considered, but higher-benefit actions that cannot be judged based on past data may be ignored. Therefore, in unstable network environments such as air-sea cross-domain communication networks (especially underwater network parts), when the quality function changes over time, the greedy strategy is difficult to achieve ideal results. Therefore, the softmax strategy is chosen.

[0009] Under this strategy, the general expression of action selection probability is as follows: , Where n is the number of possible actions, represents the corresponding quality parameter, is a variable parameter, called the temperature coefficient, which controls the trade-off between the exploration and exploitation phases to avoid falling into local optimal solutions. The temperature coefficient is set to change linearly over time to achieve convergence of the learning process in a finite time, as shown below: , in, It represents the time required for the algorithm to converge. and Respectively indicate time The initial temperature coefficient and the final temperature coefficient within the time middle, and .

[0010] State-Action Quality Function Indicates the current state Next select action First, the agent needs to calculate the value of each possible action of , and then choose the next action based on it. The topic explores two ways to set the cost function: Q learning and SARSA. Q learning is an off-policy reinforcement learning algorithm that updates the value function as follows: , , The present invention adopts a policy reinforcement learning method, and the update method of its value function is as follows: , , in, is the discount factor used to weaken the future Q value, is the learning rate, is the reward at time t. In Q-learning, the agent updates the reward function according to the action with the maximum benefit among the optional actions, but in policy reinforcement learning, the agent updates strictly according to experience, that is, Q-learning assumes that the optimal strategy is achieved in the future, while the SARSA strategy adopted by the present invention adopts clear future rewards, making action selection safer and more time-saving, and more suitable for air-sea cross-domain network application scenarios that are sensitive to errors.

[0011] The air-sea cross-domain network SDN routing protocol based on reinforcement learning is described as follows: (1) QoS-aware reward design: In the air-sea cross-domain network communication environment, since it involves underwater, surface, air and even space network domains, the choice of communication paths is relatively rich, and the actual communication path selection often needs to be determined based on the actual mission requirements, that is, QoS requirements. Different communication tasks have different demand characteristics for QoS. Military reconnaissance tasks require millisecond-level end-to-end latency and more than 99.99% transmission reliability. Marine environment monitoring tasks focus on 10Gbps-level bandwidth guarantee, while emergency search and rescue tasks need to give priority to meeting business continuity indicators. Traditional static routing strategies are difficult to adapt to the dynamic coupling relationship of multi-dimensional QoS indicators, resulting in a decrease in network resource utilization.

[0012] Inspired by this, this paper proposes a QoS-aware reward function based on reinforcement learning for SDN-based air-sea cross-domain networks. Based on this function, the task request end adds the QoS requirement data to the data stream, and the agent can find the routing path with the overall maximum QoS-aware reward according to different task requirements. In order to meet different QoS requirements, the QoS-aware reward function is defined as follows: , in, Indicates selection The cost of the action is set to a constant value in the present invention in order to provide quality of service. , , , , , , is an adjustable parameter determined by the QoS requirements of each data flow. The various parts of the reward function are defined as follows: , , , , , In the above formula, Indicates the number of switches. It is a switch The number of neighbors of and Respectively represent switches arrive The link transmission delay and queue delay of , and Respectively represent The link's available bandwidth, total bandwidth, and packet loss rate. Yes Link The relative transmission delays to all available connection links, Yes Link The relative queuing delay to the average queuing delay of the network, Indicates link The packet loss rate, Indicates link The relative available bandwidth to the average link bandwidth of the network. All quality of service weights are in Between, according to When the value is selected, the agent will Actions with values ​​close to 1 are considered preferred actions, while actions close to -1 represent low-yield actions.

[0013] (2) The following formula is a schematic diagram of the data flow: , Indicates the number of switches. A single flow is represented by a four-tuple, including the starting node , destination node , QoS requirements and traffic type The present invention classifies traffic into three types: 1) delay-sensitive; 2) loss-sensitive; 3) jitter-sensitive. The protocol proposed in the present invention aims to determine the best routing path for each data flow with specific QoS requirements, and includes the following four steps: First, the switch uses The message sends the QoS requirements of the data flow to the master controller in its subdomain. If the master controller has insufficient capacity, some switches will use CASM to migrate to the slave controller. The controller receives After receiving the message, the classification module will classify the traffic into three types: include , and ), corresponding to the second, third, and fourth parts of the reward function (delay, loss, and bandwidth). Then, The function will compare it with , and The SARSA model of the present invention is used as input to calculate the optimal routing path. The third step is reinforcement learning. First, according to the QoS requirements and traffic type Adjust the parameters of the reward function. Then continue to select the next hop until the destination is reached and record the routing path taken. Finally, the controller sends the optimal routing path obtained by the reinforcement learning algorithm to the flow table of the switch through the OpenFlow protocol.

[0014] Compared with the related art, the SDN-based air-sea cross-domain network reinforcement learning routing algorithm provided by the present invention has the following beneficial effects: The present invention significantly improves the comprehensive performance of air-sea cross-domain networks by integrating the global control capability of SDN and the dynamic decision-making advantages of reinforcement learning. It can perceive network topology changes, link loads and task requirements in real time, and adaptively optimize routing strategies with the help of multi-objective reinforcement learning models. It can meet the diverse QoS requirements of various cross-domain collaborative tasks in a dynamic environment. At the same time, it simplifies cross-domain protocol interactions through the unified control plane of SDN, reduces end-to-end latency, enhances the collaborative efficiency of satellites, drones, air-sea cross-domain gateways and underwater nodes, supports rapid adaptation to diverse scenario requirements such as marine exploration and emergency rescue, and provides intelligent routing solutions for air-sea integrated communications, which has significant engineering application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a framework diagram of the Q learning algorithm mentioned in the present invention; Figure 2 This is a schematic diagram of the routing algorithm proposed by the present invention. DETAILED DESCRIPTION

[0016] The present invention will be further described below in conjunction with the accompanying drawings and implementation modes.

[0017] The invention provides an SDN-based air-sea cross-domain network reinforcement learning routing algorithm, which includes a reinforcement learning action selection strategy and a routing protocol description.

[0018] The reinforcement learning action selection strategy under the SDN architecture is as follows: First, a Markov decision chain is used to provide a mathematical framework for system modeling of reinforcement learning, represented by a multivariable group (S, A, P, R), where S represents a finite state set, A represents a finite action set, P represents a state transition probability set, and R represents a reward set. The correct action selection strategy and state-action quality function are selected to ensure that the best long-term return is obtained while accelerating convergence, thereby ensuring QoS satisfaction rate.

[0019] The action selection strategy is responsible for specifying the action mode of the agent and mapping the state to the action. The ideal action selection strategy should be to select high-benefit actions with high probability and low-benefit or negative-benefit actions with low probability. Currently, there are two widely used action selection strategies, greedy and softmax. Under the greedy strategy, the action with the highest benefit is always selected at each step, and only the past data based on the current state and the quality function are considered, but higher-benefit actions that cannot be judged based on past data may be ignored. Therefore, in unstable network environments such as air-sea cross-domain communication networks (especially underwater network parts), when the quality function changes over time, the greedy strategy is difficult to achieve ideal results. Therefore, the softmax strategy is chosen.

[0020] Under this strategy, the general expression of action selection probability is as follows: , Where n is the number of possible actions, represents the corresponding quality parameter, is a variable parameter, called the temperature coefficient, which controls the trade-off between the exploration and exploitation phases to avoid falling into local optimal solutions. The temperature coefficient is set to change linearly over time to achieve convergence of the learning process in a finite time, as shown below: , in, It represents the time required for the algorithm to converge. and Respectively indicate time The initial temperature coefficient and the final temperature coefficient within the time middle, and .

[0021] State-Action Quality Function Indicates the current state Next select action First, the agent needs to calculate the value of each possible action of , and then choose the next action based on it. The topic explores two ways to set the cost function: Q learning and SARSA. Q learning is an off-policy reinforcement learning algorithm that updates the value function as follows: , , The present invention adopts a policy reinforcement learning method, and the update method of its value function is as follows: , , in, is the discount factor used to weaken the future Q value, is the learning rate, is the reward at time t. In Q-learning, the agent updates the reward function based on the action with the maximum benefit among the available actions, such as Figure 1 As shown, in policy reinforcement learning, the agent is updated strictly according to experience, that is, Q learning assumes that the optimal strategy is achieved in the future. The SARSA strategy adopted by the present invention adopts clear future rewards, making action selection safer and more time-saving, and more suitable for air-sea cross-domain network application scenarios that are more sensitive to errors.

[0022] The air-sea cross-domain network SDN routing protocol based on reinforcement learning is described as follows: (1) QoS-aware reward design: In the air-sea cross-domain network communication environment, since it involves underwater, surface, air and even space network domains, the choice of communication paths is relatively rich, and the actual communication path selection often needs to be determined based on the actual mission requirements, that is, QoS requirements. Different communication tasks have different demand characteristics for QoS. Military reconnaissance tasks require millisecond-level end-to-end latency and more than 99.99% transmission reliability. Marine environment monitoring tasks focus on 10Gbps-level bandwidth guarantee, while emergency search and rescue tasks need to give priority to meeting business continuity indicators. Traditional static routing strategies are difficult to adapt to the dynamic coupling relationship of multi-dimensional QoS indicators, resulting in a decrease in network resource utilization.

[0023] Inspired by this, this paper proposes a QoS-aware reward function based on reinforcement learning for SDN-based air-sea cross-domain networks. Based on this function, the task request end adds the QoS requirement data to the data stream, and the agent can find the routing path with the overall maximum QoS-aware reward according to different task requirements. In order to meet different QoS requirements, the QoS-aware reward function is defined as follows: , in, Indicates selection The cost of the action is set to a constant value in the present invention in order to provide quality of service. , , , , , , is an adjustable parameter determined by the QoS requirements of each data flow. The various parts of the reward function are defined as follows: , , , , , In the above formula, Indicates the number of switches. It is a switch The number of neighbors of and Respectively represent switches arrive The link transmission delay and queue delay of , and Respectively represent The link's available bandwidth, total bandwidth, and packet loss rate. Yes Link The relative transmission delays to all available connection links, Yes Link The relative queuing delay to the average queuing delay of the network, Indicates link The packet loss rate, Indicates link The relative available bandwidth to the average link bandwidth of the network. All quality of service weights are in Between, according to When the value is selected, the agent will Actions with values ​​close to 1 are considered preferred actions, while actions close to -1 represent low-yield actions.

[0024] (2) The following formula is a schematic diagram of the data flow: , Indicates the number of switches. A single flow is represented by a four-tuple, including the starting node , destination node , QoS requirements and traffic type The present invention classifies traffic into three types: 1) delay-sensitive; 2) loss-sensitive; 3) jitter-sensitive. The protocol proposed in the present invention aims to determine the best routing path for each data flow with specific QoS requirements, such as Figure 2 As shown, it includes the following four steps: First, the switch uses The message sends the QoS requirements of the data flow to the master controller in its subdomain. If the master controller has insufficient capacity, some switches will use CASM to migrate to the slave controller. The controller receives After receiving the message, the classification module will classify the traffic into three types: include , and ), corresponding to the second, third, and fourth parts of the reward function (delay, loss, and bandwidth). Then, The function will compare it with , and The SARSA model of the present invention is used as input to calculate the optimal routing path. The third step is reinforcement learning. First, according to the QoS requirements and traffic type Adjust the parameters of the reward function. Then continue to select the next hop until the destination is reached and record the routing path taken. Finally, the controller sends the optimal routing path obtained by the reinforcement learning algorithm to the flow table of the switch through the OpenFlow protocol.

[0025] The specific implementation of the SDN-based air-sea cross-domain network reinforcement learning routing algorithm proposed in the present invention is: (1) Environment initialization and network perception. Deploy the SDN controller as the global decision center, establish connections with airspace (drones, satellites), sea surface (buoys, ships) and underwater (unmanned submersibles) nodes through the OpenFlow protocol, and discover and build cross-domain network topology maps in real time. Initialize the deep reinforcement learning model (DRL), define the state space (network link status, node location, task QoS requirements), action space (optional cross-domain path set) and multi-dimensional reward function.

[0026] (2) Real-time status information collection. The SDN controller periodically collects the following dynamic data: Link status: obtain the real-time delay, available bandwidth, packet loss rate and stability index of each link through detection messages; Node status: GPS / acoustic positioning module reports the node’s three-dimensional coordinates, moving speed, and remaining energy; Mission requirements: The mission management module provides data flow priorities (such as disaster emergency response is the highest level) and QoS requirements (upper limit of latency, lower limit of bandwidth, and reliability threshold).

[0027] (3) The reinforcement learning decision engine runs. The collected raw data is normalized into a state vector, including a link quality matrix and a task requirement vector. Subsequently, the state vector is input into the reinforcement learning model, which captures the dynamic timing characteristics of the network, outputs the action value of each candidate path, adopts a soft maximum strategy, and generates 1-2 suboptimal paths as backups. Based on the action selection results, the Dijkstra algorithm is called in combination with link constraints to generate a specific forwarding node sequence from the source node to the target node.

[0028] (4) Flow table delivery and execution verification. The SDN controller delivers flow table entries to all nodes on the path through the OpenFlow protocol, specifying the forwarding rules for data packets and the backup path switching conditions (such as the primary path packet loss rate > 5% or the delay suddenly increases by 50ms); the source node switches the data flow to the new path according to the flow table rules, and the controller simultaneously monitors the first packet transmission delay and throughput to verify the validity of the path.

[0029] (5) Online model update and optimization. Based on the path execution results, the immediate reward value is calculated, and the state-action-reward-new state tuple is stored in the experience pool. When the sample size exceeds the threshold, batch data is randomly extracted to train the reinforcement learning network, and the Q value is updated to approximate the objective function. Based on the historical task completion status, the exploration rate and reward function weight are adaptively adjusted to prioritize the QoS requirements of high-priority tasks.

[0030] Compared with the related art, the SDN-based air-sea cross-domain network reinforcement learning routing algorithm provided by the present invention has the following beneficial effects: The present invention provides an SDN-based air-sea cross-domain network reinforcement learning routing algorithm, which realizes global network state perception and centralized control through the SDN architecture, and designs a dynamic routing decision model in combination with reinforcement learning technology. The algorithm can analyze task QoS requirements in real time, optimize multi-objective routing strategies through autonomous exploration and feedback mechanisms, and realize high-throughput routing for bandwidth-sensitive tasks, optimal path selection for delay-sensitive tasks, and adaptive load balancing for burst traffic. Compared with traditional methods, this solution breaks through the limitations of static rules, significantly improves the utilization of network resources and the differentiated service quality assurance capabilities of cross-domain tasks, and provides intelligent communication support for air-sea collaborative operations.

[0031] The above only describes the preferred embodiments of the present invention. It should be understood that the present invention is not limited to the above specific implementation methods. Although the present invention has been disclosed as above in the preferred embodiments, it is not used to limit the present invention. Any technician familiar with the profession can make some changes or modify the technical contents disclosed above into equivalent embodiments with equivalent changes without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement made to the above embodiments without departing from the content of the technical solution of the present invention, based on the technical essence of the present invention, within the spirit and principles of the present invention, still fall within the protection scope of the technical solution of the present invention.

Claims

1. A SDN-based air-sea cross-domain network reinforcement learning routing algorithm, characterized in that: The following steps are involved: S1, realize global network status perception and centralized control through SDN controller, collect network topology, link status and task QoS requirements in real time; S2. Build a reinforcement learning model that takes the dynamic network state as input and optimizes the routing strategy through a multi-objective reward function; S3, uses a soft maximum strategy for action selection and dynamically adjusts path selection to meet differentiated QoS requirements; S4, update the Q value through the SARSA algorithm, and combine the Dijkstra algorithm to generate the optimal path and the suboptimal backup path; S5. Flow table entries are issued through the OpenFlow protocol to achieve path switching and dynamic optimization.

2. According to the SDN-based air-sea cross-domain network reinforcement learning routing algorithm according to claim 1, it is characterized in that: The reinforcement learning model is based on Markov decision chain modeling and is represented by a multi-variable group S, A, P, and R, where S represents a finite state set, A represents a finite action set, P represents a state transition probability set, and R represents a reward set. The finite state set S includes network link status, node location, and task QoS requirements, the finite action set A includes an optional cross-domain path set, and the reward set R is dynamically adjusted through a multi-objective reward function.

3. According to the SDN-based air-sea cross-domain network reinforcement learning routing algorithm of claim 1, it is characterized in that: The soft-max strategy controls the trade-off between exploration and exploitation through a temperature coefficient that varies linearly with time to accelerate convergence and avoid local optimal solutions.

4. According to the SDN-based air-sea cross-domain network reinforcement learning routing algorithm of claim 1, it is characterized in that: The multi-objective reward function includes weighted values ​​of latency, bandwidth, packet loss rate and task priority, and the weights are dynamically adjusted according to the task QoS requirements to meet differentiated service quality requirements.

5. According to the SDN-based air-sea cross-domain network reinforcement learning routing algorithm of claim 1, it is characterized in that: The routing strategy includes: (a) Real-time collection of link delay, available bandwidth, packet loss rate and node location information; (b) Capture the dynamic temporal characteristics of the network through the reinforcement learning model and output the action value of the candidate path; (c) Generate a specific forwarding node sequence from the source node to the target node based on the action selection result and send it to the flow table of the switch.

6. According to the SDN-based air-sea cross-domain network reinforcement learning routing algorithm of claim 1, it is characterized in that: The path optimization includes: (a) Dynamically adjust the path selection strategy through QoS-aware reward function; (b) Calculate the immediate reward value based on the path execution result and update the Q value of the reinforcement learning model; (c) Adaptively adjust the exploration rate and reward function weight to prioritize the QoS requirements of high-priority tasks.

7. The SDN-based air-sea cross-domain network reinforcement learning routing algorithm according to claim 1 is characterized in that: The SDN controller implements path delivery and execution verification through the OpenFlow protocol, including: (a) Specify the forwarding rules for data packets and the backup path switching conditions; (b) Monitor the first packet transmission delay and throughput to verify the path validity; (c) Update the reinforcement learning model online based on the path execution results.

Citation Information

Patent Citations

  • Underwater multi-agent-oriented Q learning ant colony routing method

    CN111065145A

  • SDN (Software Defined Network) core network QoS (Quality of Service) routing optimization algorithm based on reinforcement learning

    CN112822109A

  • Intelligent community energy optimization scheduling method and system based on reinforcement learning, and storage medium

    CN117172499A

  • Low earth orbit satellite network routing method based on reinforcement learning

    CN117792984A

  • QoS (Quality of Service) sensing adaptive routing method

    CN118250210A

Cited By

  • Load balancing strategy method adaptive to air-sea cross-domain network control plane under SDN (Software Defined Network) architecture

    CN120034493A

  • Cross-group routing optimization method and device, and storage medium

    CN120263716A

  • Marine fishery resource big data dynamic allocation method based on reinforcement learning

    CN120562994A

  • Distributed large-scale anti-traceability elastic network intelligent routing method and system

    CN120567751A

  • Underwater topology control and channel selection method based on hierarchical reinforcement learning

    CN120729721A