Deterministic network congestion avoidance flow routing scheduling method

Through the deterministic network congestion avoidance of traffic routing scheduling mechanism combined with deep reinforcement learning and P4 programmable data plane, the problems of dynamic traffic fluctuations and burst congestion are solved, efficient path scheduling and rapid response are achieved, and network performance is improved.

CN120499102APending Publication Date: 2025-08-15CHONGQING UNIV
View PDF 0 Cites 10 Cited by

Patent Information

Application Number
CN202510808065.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When existing deterministic networks face dynamic traffic fluctuations and burst congestion, it is difficult to achieve efficient path scheduling, resulting in increased delay jitter in packets and packet loss. Traditional methods cannot predict congestion trends in real time and path switching is not optimized in coordination.

Method used

A deterministic network congestion avoidance traffic routing scheduling mechanism based on deep reinforcement learning is adopted, combined with IFIT flow detection and P4 programmable data plane, a perception-decision-execution closed-loop system is built, and localized and rapid response is achieved through real-time congestion coefficient calculation and intelligent path decision-making.

Benefits of technology

It significantly improves network throughput, reduces latency and jitter, effectively balances load, provides high-reliability dynamic scheduling, and improves network robustness and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499102A_ABST
    Figure CN120499102A_ABST
Patent Text Reader

Abstract

The invention relates to a deterministic network congestion avoidance flow routing scheduling method, which belongs to the technical field of industrial internet, and comprises the following steps: S1, sensing network congestion based on IFIT flow detection and queue state, and calculating a congestion coefficient; s2, performing planning decision on a traffic routing path based on deep reinforcement learning to avoid congestion; and S3, sinking congestion coefficient calculation and strategy mapping calculation logic to switch hardware by using a P4 programmable data plane to realize localized closed-loop control. Through dynamic path optimization and localization execution, the throughput is remarkably improved, the time delay and jitter are reduced, the intelligent routing mechanism effectively balances the load, and a high-reliability dynamic scheduling solution is provided for a large-scale deterministic network in combination with the deterministic guarantee capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of industrial Internet and relates to a deterministic network congestion avoidance traffic routing scheduling method. Background Art

[0002] With the rapid development of emerging applications such as the Industrial Internet, autonomous driving, and telemedicine, the demand for low latency, high reliability, and determinism in network communications is increasing. Traditional networks, based on a "best-effort" transmission mode, struggle to meet the stringent time-sensitivity and determinism requirements of these applications. Against this backdrop, Time-Sensitive Networking (TSN) and Deterministic Networking (DetNet) have emerged, aiming to provide bounded latency, low jitter, and high-reliability communication quality of service (QoS) guarantees for critical services. DetNet has exposed numerous bottlenecks in actual deployment. Dynamic traffic fluctuations, link heterogeneity, and sudden congestion make it difficult for traditional static scheduling mechanisms to balance network real-time performance, reliability, and resource efficiency. A smarter, more adaptive solution is urgently needed.

[0003] Current DetNet core technologies focus on traffic shaping and scheduling, aiming to achieve deterministic transmission through precise resource allocation. Static scheduling methods based on the Time-Aware Shaper (TAS) allocate bandwidth resources through predefined time slots. While this method effectively reduces transmission jitter, it cannot adapt to dynamic traffic changes. Traffic bursts or localized congestion often lead to insufficient or wasted reserved resources, causing significant transmission delays and even service interruptions. Meanwhile, multipath redundancy technologies improve reliability through packet duplication and elimination. However, large amounts of redundant traffic exacerbate bandwidth pressure and load imbalance, potentially leading to secondary congestion in resource-constrained scenarios. Existing congestion-aware mechanisms often rely on historical metrics such as link utilization or packet loss rates, lacking the ability to monitor underlying network behavior in real time and making it difficult to predict short-term congestion trends. For example, when node queue loads rapidly increase due to traffic bursts, traditional methods often only trigger path switching after congestion occurs, exposing critical data flows to the risk of packet loss or timeouts. Furthermore, existing path switching mechanisms, lacking coordinated optimization with traffic scheduling, can lead to packet out-of-order or congestion migration. Summary of the Invention

[0004] Large-scale, deterministic network traffic convergence can easily lead to network congestion, increased packet latency and jitter, and even packet loss, making it difficult to guarantee deterministic network performance. In light of this, the present invention proposes a deterministic network congestion avoidance traffic routing scheduling mechanism (Intelligent Congestion Perception and Avoidance Traffic Scheduling, CATRS-DRL) based on deep reinforcement learning. By integrating with-flow detection, deep reinforcement learning, and P4 programmable data plane technology, a closed-loop "perception-decision-execution" system is constructed. Based on with-flow detection technology, queue depth, average packet queuing delay, and its rate of change are collected in real time. A multi-factor dynamic congestion coefficient model is designed to accurately quantify network status. Congestion coefficient and other related parameters are input into a deep Q network, and deep reinforcement learning is used to implement intelligent path decision-making. Low-latency routes are generated by integrating service priority, link congestion, and path efficiency. Finally, with the help of a P4 programmable pipeline, path switching policies are mapped into switch actions, and action execution is decentralized to the switch hardware layer, achieving localized and rapid response.

[0005] In order to achieve the above object, the present invention provides the following technical solutions:

[0006] A method for deterministic network congestion avoidance traffic routing scheduling, comprising the following steps:

[0007] S1: Detects network congestion based on in-situ flow information telemetry (IFIT) and queue status, and calculates the congestion coefficient.

[0008] S2: Plans traffic routing paths based on deep reinforcement learning to avoid congestion.

[0009] S3: Use the P4 programmable data plane to move the congestion coefficient calculation and policy mapping calculation logic down to the switch hardware to achieve localized closed-loop control.

[0010] Furthermore, step S1 specifically includes:

[0011] When ordinary data packets enter the switching device, the device dynamically inserts the IFIT header during forwarding and collects key metrics. On the data plane, it uses a lightweight header encapsulation format based on the IPv6 extension header. It also introduces a dynamic sampling mechanism to adaptively adjust the IFIT packet generation frequency based on the current network load. Periodic sampling is used when bandwidth utilization falls below a threshold. When queue depth or queuing delay exceeds a warning value, the device switches to event-triggered mode, capturing sudden congestion events in real time.

[0012] The real-time congestion coefficient C is used to reflect the congestion level of a node. A larger real-time congestion coefficient indicates more severe congestion. Four indicators, queue occupancy, queuing delay, queue depth change rate, and queuing delay change rate, are selected to calculate the congestion level of a node. The correlation between the queue occupancy and the dequeue interval of a node is also comprehensively considered.

[0013] Furthermore, in terms of bandwidth, IFIT uses lightweight IPv6 extension header encapsulation, which occupies 8 bytes per packet. Under a 10 Gbps link, the single-port bandwidth occupancy is 6.4 kbps in basic periodic sampling mode, and the maximum overhead in event triggering mode is 5.33 Mbps.

[0014] Furthermore, the calculation method of the real-time congestion coefficient is expressed as:

[0015]

[0016] where ρ Q represents the queue occupancy, η T It represents the proportion of queuing delay, which is calculated as the ratio of the average waiting time of the data packet in the queue to the maximum tolerable delay allowed by the queue, Δ Q represents the queue depth change rate, Δ T represents the growth rate of the queuing delay. The first half of the formula is a static factor, which balances the queue and queuing delay through geometric mean and weakens the impact of extreme values of a single parameter. The second half is a dynamic factor, which focuses only on the positive growth trend of the queue and queuing delay and captures the coordinated deterioration of the two through multiplication.

[0017] Furthermore, the step S2 includes constructing a network topology model. In the network topology model G(V,E), the vertex set V = {v1, v2, ..., v N} represents the switching device node, and the edge set E={e1,e2,...,e M} represents the physical link between node devices; each edge e∈E has a certain transmission delay d e and bandwidth capacity c e ; The traffic demand is represented as an N-dimensional matrix D, where any element d i,j Indicates that the slave switch v i is the source node, v j The bandwidth requirement for traffic aggregation at the destination node;

[0018] Request for specific traffic i,j , enumerate source nodes v i To the target node v jAll feasible transmission paths between the two paths; perform service constraint verification on the paths and eliminate candidate paths that fail the verification; link bandwidth resources directly determine the transmission feasibility of data streams; select K optimal candidate paths from the full set of paths P to form the service customized path set P i ; Traffic request d i,j Strictly follow the single-path allocation principle, and its mathematical constraints are defined as:

[0019]

[0020] Among them, the binary decision variable x i (d) Characterizes the path selection state: If traffic request d selects path p as the transmission channel, the corresponding variable is assigned the value 1; otherwise, the variable associated with the unselected path takes the value 0. This constraint mechanism ensures that each traffic request is transmitted only through a single path in the candidate path set, resulting in the traffic demand constraint matrix D' = D·x;

[0021] For any node, use Represents the node information, where ρ L Indicates the link utilization from the previous node to the current node.

[0022] Furthermore, in step S2, a Deep Q Network (DQN) is used to calculate the routing strategy. The state S input to the DQN is defined as a vector that reflects the current network topology, traffic information, and node information. The traffic information includes the traffic demand constraint matrix and traffic priority. The node information includes the link utilization, queue depth, and the proportion and change rate of queuing delay. Assuming that the current node has n neighbors, the state S is specifically expressed as:

[0023] S={G,D',P,s self ,s1,…,s n} (3)

[0024] Where P represents the traffic priority. self Indicates the real-time information of the current node, s1 to s n Represents the real-time information of all n neighbor nodes;

[0025] The input of DQN is the current state S, and the output is the action-value function Q of all actions in that state. The action is represented by selecting an optimal neighbor node and sending traffic in its direction. The action space is defined as the set of all feasible neighbor nodes. When the current node has n neighbor nodes, the action space is:

[0026] A={a1,a2,...,a n} (4)

[0027] The action function is expressed as selecting the largest Q value in the action space:

[0028]

[0029] Where θ is the DQN parameter, Q(s t ,a;θ) predict the benefit of selecting neighbor a; for neighbors that cannot reach the destination node through them, set their Q value to -∞ to avoid being selected;

[0030] The reward function comprehensively considers the local congestion status, service priority, and path efficiency, and is defined as:

[0031] R=R S +R C +R H +R D (6)

[0032] where R S 、R C 、R H 、R l They represent service reward, congestion penalty, path efficiency penalty, and delay penalty, respectively, and are defined as follows:

[0033] Business rewards are used to reward traffic that is successfully forwarded to the next hop or successfully delivered to the destination node. Rewards are given according to priority P:

[0034]

[0035] in, It is an indicator function, which has a value of 1 when the indicated action is completed and 0 when the action fails to complete;

[0036] Congestion penalty is used to penalize the consumption of bandwidth utilization, queue depth, and queuing delay occupancy of the next-hop node by the decision-making process:

[0037] R C =-(U next +Q next +C next ) (8)

[0038] Among them, U next Indicates the bandwidth utilization of the next hop node, Q next Indicates the queue depth ratio of the next hop node, which is calculated by the ratio of the queue depth to the maximum queue capacity. next Indicates the real-time congestion coefficient of the next-hop node;

[0039] Path efficiency is used to penalize the extra hops caused by the decision:

[0040]

[0041] Among them H taken Indicates the actual number of hops of the currently selected path, H max Indicates the maximum number of hops allowed from the source node to the destination node in the network topology, which is used to limit the path length.

[0042] Latency penalties are used to reduce latency accumulation:

[0043]

[0044] where l accumulated represents the cumulative transmission delay of the data packet from the source node to the current node, l next represents the expected transmission delay from the current node to the next hop node, l max Indicates the maximum end-to-end delay threshold allowed for the service.

[0045] Furthermore, in step S3, a programmable pipeline designed based on the P4 language moves key computing logic down to the data plane of the switching device. By embedding registers, meters, and state machines in the ingress and egress pipelines, link metric collection, real-time congestion factor calculation, and sampling mode switching are completed directly at the hardware level.

[0046] In the ingress pipeline stage, ordinary data packets are matched according to the preset sampling strategy, and a lightweight IFIT header is dynamically inserted for qualified traffic. According to the weight balance requirement of static factors and dynamic factors in the real-time congestion coefficient calculation formula (1), the arithmetic logic of the corresponding calculation unit can be modified to adjust the parameters; in the egress pipeline stage, the port status parameters are analyzed and updated in real time through the collaborative calculation of registers and meters. Combined with the calculation method of the congestion coefficient, the compound operation of static factors and dynamic factors is completed directly on the data plane, and the calculation results are finally written into the metadata field of the switch to provide real-time input for subsequent intelligent scheduling;

[0047] At the policy execution level, by deeply integrating the features of the P4 programmable data plane, the optimal neighbor node decision output by DQN reasoning is directly mapped to pipeline-level action instructions; the business priority weight or path efficiency penalty coefficient in the deep reinforcement learning reward function is dynamically tuned by updating the matching-action rules of the P4 table; the edge node encodes the policy decision as a P4 action function, dynamically modifies the forwarding table entries in the packet processing pipeline, and redirects the next hop of the packet to the selected neighbor node.

[0048] Furthermore, the deep Q network training framework is as follows: the edge computing node continuously accumulates historical state-action-reward triplet data through the locally deployed experience replay cache, and incrementally fine-tunes the distributed global model based on the transfer learning mechanism.

[0049] The beneficial effects of the present invention are: through dynamic path optimization and localized execution, the present invention significantly improves throughput, reduces latency and jitter, and its intelligent routing mechanism effectively balances the load. Combined with deterministic guarantee capabilities, it provides a highly reliable dynamic scheduling solution for large-scale deterministic networks. According to the simulation results below, under high-load scenarios, the throughput of CATRS-DRL reaches 9.8Gbps, which is 58% and 21% higher than PREOF and MPRM, respectively, significantly surpassing the performance bottleneck of the comparison scheme. Secondly, based on the real-time congestion perception mechanism driven by queue behavior, it accurately perceives local congestion trends by dynamically quantifying queue occupancy, queuing delay and its rate of change. Based on IFIT flow detection and sampling, the congestion detection delay is effectively reduced, effectively suppressing the bandwidth occupation of redundant traffic. In addition, through the P4 programmable data plane to achieve localized policy execution, the path switching delay is less than 1ms, and the packet loss rate is only 0.03% under 20% link failure, which is 80% lower than MPRM, greatly improving the robustness of the network.

[0050] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0052] Figure 1 A framework for traffic routing and scheduling mechanisms to avoid congestion in deterministic networks based on deep reinforcement learning;

[0053] Figure 2 To report IFIT detection data based on telemetry;

[0054] Figure 3 It is a deep Q network structure;

[0055] Figure 4 Programmable dynamic policy execution flow chart for programmable data plane;

[0056] Figure 5 This is a diagram of the deep Q network training framework;

[0057] Figure 6 It is a Fat-Tree network topology;

[0058] Figure 7 (a) and (b) are the throughput comparison charts of CATRS-DRL, PREOF and MPRM under different loads;

[0059] Figure 8 The throughput comparison chart of CATRS-DRL, PREOF and MPRM under different loads;

[0060] Figure 9 (a) and (b) are comparisons of end-to-end delay and delay jitter of CATRS-DRL, PREOF, and MPRM under different loads;

[0061] Figure 10 The packet loss rate comparison chart of CATRS-DRL, PREOF and MPRM under different link failure rates; DETAILED DESCRIPTION

[0062] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0063] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.

[0064] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.

[0065] Example 1:

[0066] like Figure 1As shown, the present invention provides a deterministic network congestion-avoidance traffic routing scheduling mechanism (Intelligent Congestion Perception and Avoidance Traffic Scheduling, CATRS-DRL) based on deep reinforcement learning. This technology consists of three parts: first, a network congestion perception and congestion coefficient calculation method based on flow detection and queue status; second, a congestion-avoidance traffic routing path decision mechanism based on deep reinforcement learning; and third, a traffic scheduling and forwarding strategy based on a programmable data plane. This technology implements a closed-loop "perception-decision-execution" model. By integrating flow detection technology, deep reinforcement learning algorithms, and data plane programmability, it achieves real-time network status perception, dynamic path optimization, and policy execution. The perception layer uses IFIT to implement flow detection and status collection, providing real-time input for intelligent scheduling. The decision layer uses deep reinforcement learning with a deep Q network to perform intelligent path planning. The execution layer, based on the programmable data plane P4 programmable pipeline, sinks congestion calculation, policy mapping, and other logic to the switch hardware layer, achieving localized closed-loop control.

[0067] 1. Network Congestion Perception and Congestion Coefficient Calculation Method Based on Flow Detection and Queue Status

[0068] like Figure 2 As shown in Figure 1, network data is collected using IFIT. When a standard data packet arrives at a switch in the in-band network telemetry system, the IFIT module matches and mirrors the packet and encapsulates the information in an IFIT header. The switch forwards the telemetry information to the telemetry server, which then reports it to the controller for data analysis and routing decisions.

[0069] When ordinary data packets enter the switching device, the device dynamically inserts the IFIT header during the forwarding process and collects key indicators, including the current port bandwidth utilization, output queue depth, queuing delay, packet timestamp, and latency. To ensure the reliability and real-time performance of the system and reduce network overhead, the solution adopts a lightweight header encapsulation format based on the IPv6 extension header on the data plane, and introduces a dynamic sampling mechanism to adaptively adjust the IFIT packet generation frequency based on the current network load. When the bandwidth utilization is lower than the threshold, periodic sampling is used. When the queue depth or queuing delay indicator is detected to exceed the warning value, the device switches to event trigger mode to capture sudden congestion events in real time.

[0070] In terms of bandwidth, IFIT uses a lightweight IPv6 extension header encapsulation, which occupies 8 bytes per packet. Under a 10Gbps link, the single-port bandwidth in basic periodic sampling mode occupies approximately 6.4kbps, accounting for only 0.000064% of the link bandwidth. In event-triggered mode, the maximum overhead is 5.33Mbps, accounting for 0.053% of the link bandwidth, which is significantly lower than traditional active detection or full sampling solutions.

[0071] In order to quantitatively guide IFIT decisions, it is proposed to use the real-time congestion coefficient C as an indicator to reflect the congestion level of the node. The larger the real-time congestion coefficient value, the more severe the congestion level. Its accuracy directly affects the performance of the load balancing strategy. At the same time, in order to ensure the timeliness of data acquisition, the real-time congestion coefficient needs to be directly calculated from the local queue behavior. Therefore, the four indicators of queue occupancy, queuing delay, queue depth change rate and queuing delay change rate are selected to calculate the congestion level of the node, and the correlation between the queue occupancy rate and the dequeue time interval of the node needs to be comprehensively considered. When the change rate is used as the basis for decision-making, the current and future congestion levels are actually comprehensively considered, and the calculation can be completed without relying on explicit information transmission between controllers or switches. In summary, the calculation method of the real-time congestion coefficient can be expressed as:

[0072]

[0073] where ρ Q represents the queue occupancy, η T It represents the proportion of queuing delay, which is calculated as the ratio of the average waiting time of the data packet in the queue to the maximum tolerable delay allowed by the queue, Δ Q represents the queue depth change rate, Δ T represents the growth rate of queuing delay. The first half of the equation is a static factor, which balances the queue and queuing delay through geometric mean, mitigating the impact of extreme values of a single parameter. The second half is a dynamic factor, focusing only on the positive growth trends of the queue and queuing delay, and capturing the synergistic deterioration of the two through multiplication.

[0074] For example, when a high-priority queue with low traffic volume is frequently preempted, the queue depth is low but the queuing delay is high, triggering optimized scheduling. On the other hand, when bursty traffic is present but rapidly dequeued, the queue depth is high but the queuing delay is low, thus avoiding overreaction. Algorithm 1 shows the pseudocode for this algorithm, which more accurately reflects network congestion and performs load balancing more effectively.

[0075] Algorithm 1

[0076]

[0077]

[0078] 2. Congestion Avoidance Traffic Routing Path Decision Mechanism Based on Deep Reinforcement Learning

[0079] In the network topology model G(V,E), the vertex set V={v1,v2,...,v N} represents the switching device node, and the edge set E={e1,e2,...,e M} represents the physical link between node devices. Each edge e∈E has a certain transmission delay d e and bandwidth capacity c e The traffic demand is represented as an N-dimensional matrix D, where any element d i,j Indicates that the slave switch v i is the source node, v j The bandwidth requirement for traffic aggregation at the destination node.

[0080] Request for specific traffic i,j , enumerate source nodes v i To the target node v j All feasible transmission paths between. Service constraint verification is performed on the paths. Candidate paths that fail the verification will be eliminated to ensure path validity. Link bandwidth resources are the core evaluation indicator of path availability and directly determine the transmission feasibility of data streams. K optimal candidate paths are selected from the full set of paths P to form the service customized path set P. i Traffic request d i,j Strictly follow the single-path allocation principle, and its mathematical constraints are defined as:

[0081]

[0082] Among them, the binary decision variable x i (d) Represents the path selection state: If traffic request d selects path p as the transmission channel, the corresponding variable is assigned the value 1; otherwise, the variables associated with the unselected paths are assigned the value 0. This constraint mechanism ensures that each traffic request is transmitted only through a single path in the candidate path set. The resulting traffic demand constraint matrix D' = D·x.

[0083] For any node, use Represents the node information, where ρ L Indicates the link utilization from the previous node to the current node.

[0084] In large-scale deterministic networks, due to frequent dynamic changes and complex and diverse traffic flows, traditional routing algorithms struggle to adapt to dynamic traffic fluctuations and sudden congestion scenarios. Most machine learning algorithms rely on large amounts of offline training data and struggle to cope with extreme, sudden scenarios. Therefore, we chose Deep Q Networking (DQN), a deep reinforcement learning technique, to calculate routing strategies. Compared to traditional routing algorithms and some machine learning algorithms, DQN's advantage lies in its integration of deep learning and the Q-Learning framework. By building a dual-network architecture and using experience caching technology, it effectively overcomes the stability and convergence limitations of traditional methods. By leveraging real-time perception of network status for dynamic path optimization, DQN achieves a better balance between local optimization and global load balancing, ensuring low latency and high reliability while improving resource utilization. Figure 3 The structure of a deep Q-network is shown. The state S of the input DQN is defined as a vector that reflects the current network topology, traffic information, and node information. Traffic information includes the traffic demand constraint matrix and traffic priority. Node information includes the percentage and rate of change of link utilization, queue depth, and queuing delay. Assuming the current node has n neighbors, the state S can be expressed as:

[0085] S={G,D',P,s self ,s1,…,s n} (3)

[0086] Where P represents the traffic priority. self Indicates the real-time information of the current node, s1 to s n Represents the real-time information of all n neighbor nodes.

[0087] For DQN, its input is the current state S, and its output is the action value function Q of all actions in that state.

[0088] Here, the action is expressed as selecting an optimal neighbor node and sending traffic in its direction. Therefore, the action space is defined as the set of all feasible neighbor nodes. When the current node has n neighbor nodes, the action space is:

[0089] A={a1,a2,...,a n} (4)

[0090] The action function is expressed as selecting the maximum Q value in the action space.

[0091]

[0092] Where θ is the DQN parameter, Q(s t,a;θ) predict the benefit of selecting neighbor a. For neighbors that cannot reach the destination node through them, their Q value is set to -∞ to avoid being selected.

[0093] To guide the agent to prioritize paths with low congestion and that meet business needs when switching paths, the reward function needs to comprehensively consider the local congestion status, business priority, and path efficiency. The reward function is defined as:

[0094] R=R S +R C +R H +R D (6)

[0095] where R S 、R C 、R H 、R l They represent service reward, congestion penalty, path efficiency penalty, and delay penalty respectively. They are defined as follows:

[0096] Business rewards are used to reward traffic that is successfully forwarded to the next hop or successfully delivered to the destination node. Rewards are given according to priority P:

[0097]

[0098] It is an indicator function. Its value is 1 when the indicated action is completed and 0 when the action fails to complete.

[0099] Congestion penalty is used to penalize the consumption of bandwidth utilization, queue depth, and queuing delay occupancy of the next-hop node by the decision-making process:

[0100] R C =-(U next +Q next +C next ) (8)

[0101] Among them, U next Indicates the bandwidth utilization of the next hop node, Q next Indicates the queue depth ratio of the next hop node, which is calculated by the ratio of the queue depth to the maximum queue capacity. next Indicates the real-time congestion coefficient of the next-hop node.

[0102] Path efficiency penalties cause extra hops, preventing excessively long detours to reduce congestion.

[0103]

[0104] Among them H taken Indicates the actual number of hops of the currently selected path, H maxIndicates the maximum number of hops allowed from a source node to a destination node in a network topology, used to limit the length of the path.

[0105] Finally, the delay penalty is used to reduce the delay accumulation:

[0106]

[0107] where l accumulated represents the cumulative transmission delay of the data packet from the source node to the current node, l next represents the expected transmission delay from the current node to the next hop node, l max Indicates the maximum end-to-end delay threshold allowed for the service.

[0108] 3. Traffic Scheduling and Forwarding Strategy Based on Programmable Data Plane

[0109] This solution significantly reduces system overhead and improves real-time performance through the local processing capabilities of programmable pipelines. In traditional network architectures, state acquisition and decision execution are highly dependent on frequent interactions between the control plane and the data plane, resulting in dual bottlenecks of communication delay and computing resource consumption. To solve this problem, Figure 4 As shown in Figure 1, the programmable pipeline designed based on the P4 language sinks key computing logic to the data plane of the switching device. By embedding registers, meters, and state machines in the ingress and egress pipelines, core operations such as link metric collection, real-time congestion factor calculation, and sampling mode switching are completed directly at the hardware level. This design completely localizes state awareness and policy decision-making during packet processing, avoiding the protocol parsing and serialization overhead of cross-layer communication.

[0110] To support the dynamic collection of flow detection information and real-time congestion assessment, this solution designs a lightweight pipeline in the programmable switch based on the P4 language, and implements efficient closed-loop control within the data plane through customized processing logic.

[0111] At the ingress pipeline stage, the system matches ordinary data packets according to a preset sampling strategy and dynamically inserts a lightweight IFIT header for eligible traffic. To meet the weight balance requirements of the static and dynamic factors in the real-time congestion coefficient calculation formula 1, developers only need to modify the arithmetic logic of the corresponding calculation unit to quickly adjust the parameters without having to reconstruct the entire pipeline. Network status information can be carried along the flow without significantly increasing the length of the data packet. On this basis, the egress pipeline stage uses the collaborative calculation of registers and meters to parse and update the port status parameters in real time. Combined with the queue-queueing delay joint real-time congestion coefficient model proposed in Algorithm 1, the compound calculation of static and dynamic factors is completed directly on the data plane, and the calculation results are finally written into the metadata field of the switch to provide real-time input for subsequent intelligent scheduling.

[0112] To adapt to the dynamic characteristics of the local network environment, use Figure 5 Edge computing nodes do not directly use global model parameters. Instead, they continuously accumulate historical state-action-reward triples through a locally deployed experience replay cache and incrementally fine-tune the distributed global model using a transfer learning mechanism. This distributed learning architecture inherits the generalization capabilities of the global model while autonomously optimizing policy weights for scenarios such as local topology changes and traffic pattern shifts, effectively alleviating policy rigidity.

[0113] At the policy execution level, the system directly maps the optimal neighbor node decision output by DQN reasoning into pipeline-level action instructions by deeply integrating the features of the P4 programmable data plane. The business priority weight or path efficiency penalty coefficient in the deep reinforcement learning reward function can also be dynamically tuned by updating the matching-action rules of the P4 table. Specifically, the edge node encodes the policy decision as a P4 action function, dynamically modifies the forwarding table entry in the packet processing pipeline, and redirects the next hop of the packet to the selected neighbor node. This completely bypasses the overhead of hop-by-hop query of the routing table in the traditional control plane, and achieves fast policy execution through the parallel processing capabilities of the hardware pipeline, ensuring immediate response capabilities in bursty traffic scenarios.

[0114] To comprehensively evaluate the effectiveness of CATRS-DRL technology in dynamic network environments, this paper focuses on technical indicators such as network throughput, end-to-end delay, delay jitter, and packet loss rate. These indicators can fully reflect the performance of CATRS-DRL in congestion control, path optimization, and policy enforcement.

[0115] To verify the effectiveness of the scheme, two comparison schemes, PREOF and Multipath Reliable Routing Mechanism (MPRM), were selected. PREOF ensures reliability through packet replication and elimination mechanisms, but lacks dynamic path optimization capabilities and relies on a fixed set of paths. The significance of the comparison lies in verifying the performance difference between intelligent scheduling and static multipath forwarding in a dynamic network environment. MPRM combines the SDN controller with the P4 data plane to implement multipath forwarding, but cannot adapt to local congestion changes in real time. PREOF has low latency under low loads, but is prone to increased congestion due to redundant traffic under high loads. MPRM alleviates congestion through multipath diversion, but the controller calculation delay may become a bottleneck. CATRS-DRL combines local execution with intelligent scheduling, which is expected to balance throughput and latency.

[0116] Build a simulation platform based on MATLAB and OMNet++. Figure 6As shown in Table 1, to fully validate the performance of the CATRS-DRL solution, a three-level Fat-Tree network topology with k=6 was used, with bandwidth configured as 100 bps at the core layer, 40 Gbps at the aggregation layer, and 10 Gbps at the access layer. This structure simulates the characteristics of large-scale deterministic networks. The three-level Fat-Tree, through its layered design across the core, aggregation, and access layers, retains the load balancing benefits of multipath redundancy while also creating a typical "bandwidth funnel" by decreasing bandwidth at each layer. This simulates traffic convergence scenarios in the Industrial Internet, from edge device access, regional aggregation, and core backbone transmission, triggering multi-level link contention. In the traffic model design, background traffic simulates regular data packet arrivals using a Poisson distribution with an interval parameter λ ranging from 0.5 to 2 packets / μs, covering light to near-full load scenarios. This simulates the randomness and traffic fluctuations of regular industrial network services, spanning light to medium-to-high load scenarios, to verify the algorithm's dynamic adaptability. Burst traffic is further simulated with a higher-density Poisson arrival and exponentially distributed duration. This short, high-intensity flow simulates sudden events in industrial scenarios, such as equipment failure alarms and emergency control commands. This aims to trigger transient congestion and test the ability to mitigate congestion migration. Furthermore, traffic is categorized into three priority levels: high, medium, and low, with weights of 0.6, 0.3, and 0.1, respectively. This is directly embedded in the DQN reward function to reflect the differentiated quality of service requirements of industrial scenarios. High weights are assigned to critical services, such as production line control commands, ensuring they receive absolute priority in resource competition. A progressive penalty mechanism is implemented to prevent low-priority traffic from starving. Network optimization utilizes the Adam algorithm, with an initial learning rate of 0.001 and an experience replay buffer of 1×10⁶ to store historical state-action-reward triplets. During training, the batch size is 64, and the discount factor γ is set to 0.99 to balance immediate rewards with long-term benefits. The exploration rate ∈ uses a linear decay strategy, with an initial value of 1.0, indicating completely random exploration, and gradually decaying to 0.1 as the number of training steps increases, indicating a primarily stable strategy. The decay period covers 80% of the total training steps. Furthermore, the target network update frequency is set to synchronize the main network parameters every 100 steps to mitigate the problem of Q-value overestimation. These parameters are determined through grid search and cross-validation to ensure rapid convergence while avoiding local optima.

[0117] Table 1

[0118]

[0119] First, we compared the E2E latency and jitter performance of traffic when using DQN to calculate paths and when using traditional routing strategies to calculate paths. Figure 7Figures (a) and (b) show that when the network load increases from 50% to 90%, the end-to-end latency of CATRS-DRL-DRL increases linearly from 4.35ms to 8.3ms, while the latency of CATRS-DRL-noDRL increases from 9.2ms to 13.8ms. This performance advantage stems from the synergy between DQN's dynamic decision-making mechanism and localized execution. In contrast, traditional strategies, due to their static path selection, cannot adapt to dynamic traffic fluctuations, resulting in a surge in queuing delays under high loads. In terms of delay jitter, CATRS-DRL-DRL's jitter increases from 0.13ms to 1.305ms, consistently lower than the 1.55ms to 4.4ms range of CATRS-DRL-noDRL. DRL uses a congestion dynamic factor to accurately predict congestion trends and leverages P4 to rapidly switch paths, effectively suppressing packet delay fluctuations. Traditional strategies, on the other hand, rely on historical metrics, resulting in delayed path switching and a dispersed latency distribution.

[0120] Figure 8 The results show that when the network load reaches 90%, the average throughput of CATRS-DRL can reach 9.8Gbps. In comparison, the average throughput of PREOF is 6.2Gbps and the average throughput of MPRM is 8.1Gbps, showing a significant performance advantage of CATRS-DRL. In dynamic path optimization, DRL uses real-time selection of less congested paths to reduce bandwidth usage by redundant traffic similar to PREOF frame replication. In localized execution, the P4 pipeline is used to directly calculate the congestion coefficient and forward it, avoiding the additional delay of approximately 3ms / hop caused by controller interaction in MPRM. In the face of burst traffic of 6-8Gbps, CATRS-DRL quickly switches paths by triggering sampling when the queue is >80%. The standard deviation of throughput fluctuation is only 1.2Gbps, which is lower than PREOF's 3.5Gbps and MPRM's 2.1Gbps. In terms of resource utilization, its link bandwidth utilization standard deviation is 9.7%, which is 54% and 34% lower than PREOF's 21.3% and MPRM's 14.8%, respectively. It has better load balancing capabilities, and the core switch buffer occupancy rate is always below 70%, without overflow and packet loss.

[0121] Figure 9Figures (a) and (b) show that under 50%-90% network load, its end-to-end latency increases from 4.2ms to 8.3ms, a 97% increase, lower than PREOF and MPRM. This is due to its dynamic path optimization through real-time perception of network status using DQN, localized execution based on the P4 pipeline to reduce interaction delays with the SDN controller, and high-priority service guarantees. In terms of latency jitter, CATRS-DRL increases from 0.3ms to 0.9ms, only 29% of PREOF and 50% of MPRM. This is due to the use of a congestion coefficient formula to accurately identify and suppress congestion, event-triggered sampling to capture congestion signals earlier, and effective control of redundant traffic. Under a high load of 90%, CATRS-DRL maintains low latency of 8.3ms and jitter of 0.9ms, while PREOF and MPRM face performance bottlenecks due to static mechanisms and functional fragmentation.

[0122] Figure 10 The results show that when injected with 20% random link failures, CATRS-DRL achieves a packet loss rate of only 0.03%, significantly lower than PREOF's 0.22% and MPRM's 0.15%. This is primarily due to two factors: First, CATRS-DRL leverages the local P4 pipeline for fast fault recovery, with a path switching delay of less than 1ms. In contrast, MPRM relies on controller rerouting, with an average delay of 12ms. Second, in terms of redundant flow control, PREOF's frame replication generates a large amount of redundant traffic on the faulty link, exacerbating congestion and causing packet loss, while CATRS-DRL avoids this situation.

[0123] Simulation results demonstrate that CATRS-DRL achieves high throughput, low latency, and strong robustness in large-scale, three-level Fat-Tree networks. The synergy between IFIT real-time detection and DRL dynamic scheduling enables a closed-loop perception-decision-execution architecture. The P4 pipeline offloads congestion calculation and policy mapping to switches, avoiding control plane interaction bottlenecks. Furthermore, differentiated weighting ensures quality of service for critical traffic, meeting the stringent requirements of the Industrial Internet and real-time communications.

[0124] By perceiving network topology, traffic priority, and link status in real time, CATRS-DRL is able to generate optimal path decisions within a tolerable time. Simulation results show that under high-load scenarios, CATRS-DRL achieves a throughput of 9.8 Gbps, which is 58% and 21% higher than PREOF and MPRM, respectively, significantly surpassing the performance bottlenecks of the comparison schemes. Secondly, based on a real-time congestion perception mechanism driven by queue behavior, local congestion trends can be accurately perceived by dynamically quantifying queue occupancy, queuing delay, and its rate of change. IFIT-based flow detection and sampling effectively reduces congestion detection delays and effectively suppresses bandwidth usage by redundant traffic. In addition, localized policy execution is achieved through the P4 programmable data plane, with path switching delays less than 1 ms and packet loss rates of only 0.03% under 20% link failures, which is 80% lower than MPRM, significantly improving the robustness of the network.

[0125] Example 2:

[0126] An electronic device comprising a memory and a processor;

[0127] The memory is used to store computer programs;

[0128] The processor is configured to implement the method described in Example 1 when executing the computer program.

[0129] Example 3:

[0130] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in Example 1 is implemented.

[0131] Example 4:

[0132] A computer program product includes a computer program, which implements the method described in embodiment 1 when executed by a processor.

[0133] In the above embodiments, references to "this embodiment" in the specification indicate that a particular feature, structure, or characteristic described in conjunction with the embodiment is included in at least some embodiments, but not necessarily all embodiments. Multiple occurrences of "this embodiment" do not necessarily refer to the same embodiment.

[0134] In the above embodiments, although the invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory structures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed. The embodiments of the present invention are intended to encompass all such alternatives, modifications, and variations that fall within the broad scope of the appended claims.

[0135] Regarding the computer-readable storage medium in this embodiment, those skilled in the art will appreciate that all or part of the steps in the aforementioned method embodiments can be implemented using hardware associated with the computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When executed, the program performs the steps in the aforementioned method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0136] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store computer programs, the communication interface is used for communication, and the processor and the transceiver are used to run computer programs so that the electronic terminal executes the various steps of the above method.

[0137] In this embodiment, the memory may include a random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk storage.

[0138] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.

[0139] The present invention can be used in a wide variety of general-purpose or special-purpose computing system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments that include any of the above.

[0140] The present invention may be described in the general context of computer-executable instructions, such as program modules, executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in both local and remote computer storage media, including storage devices.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.

Claims

1. A deterministic network congestion avoidance traffic routing scheduling method, characterized by: The following steps are involved: S1: Based on in-band flow information measurement (IFIT) and queue status, it senses network congestion and calculates the congestion coefficient. S2: Plans traffic routing paths based on deep reinforcement learning to avoid congestion. S3: Use the P4 programmable data plane to move the congestion coefficient calculation and policy mapping calculation logic down to the switch hardware to achieve localized closed-loop control.

2. The deterministic network congestion avoidance traffic routing scheduling method according to claim 1, characterized in that: Step S1 specifically includes: When ordinary data packets enter the switching device, the device dynamically inserts the IFIT header during forwarding and collects key metrics. On the data plane, it uses a lightweight header encapsulation format based on the IPv6 extension header. It also introduces a dynamic sampling mechanism to adaptively adjust the IFIT packet generation frequency based on the current network load. Periodic sampling is used when bandwidth utilization falls below a threshold. When queue depth or queuing delay exceeds a warning value, the device switches to event-triggered mode, capturing sudden congestion events in real time. The real-time congestion coefficient C is used to reflect the congestion level of a node. A larger real-time congestion coefficient indicates more severe congestion. Four indicators, queue occupancy, queuing delay, queue depth change rate, and queuing delay change rate, are selected to calculate the congestion level of a node. The correlation between the queue occupancy and the dequeue interval of a node is also comprehensively considered.

3. The deterministic network congestion avoidance traffic routing scheduling method according to claim 2, characterized in that: In terms of bandwidth, IFIT uses lightweight IPv6 extension header encapsulation, which occupies 8 bytes per packet. Under a 10Gbps link, the single-port bandwidth occupancy is 6.4kbps in basic periodic sampling mode, and the maximum overhead in event triggering mode is 5.33Mbps.

4. The deterministic network congestion avoidance traffic routing scheduling method according to claim 2, characterized in that: The calculation method of the real-time congestion coefficient is expressed as: where ρ Q represents the queue occupancy, η T It represents the proportion of queuing delay, which is calculated as the ratio of the average waiting time of the data packet in the queue to the maximum tolerable delay allowed by the queue, Δ Q represents the queue depth change rate, Δ T represents the growth rate of the queuing delay. The first half of the formula is a static factor, which balances the queue and queuing delay through geometric mean and weakens the impact of extreme values of a single parameter. The second half is a dynamic factor, which focuses only on the positive growth trend of the queue and queuing delay and captures the coordinated deterioration of the two through multiplication.

5. The deterministic network congestion avoidance traffic routing scheduling method according to claim 1, characterized in that: The step S2 includes constructing a network topology model. In the network topology model G(V,E), the vertex set V = {v1, v2, ..., v N } represents the switching device node, and the edge set E={e1,e2,...,e M } represents the physical link between node devices; each edge e∈E has a certain transmission delay d e and bandwidth capacity c e ; The traffic demand is represented as an N-dimensional matrix D, where any element d i,j Indicates that the slave switch v i is the source node, v j The bandwidth requirement for traffic aggregation at the destination node; Request for specific traffic i,j , enumerate source nodes v i To the target node v j All feasible transmission paths between the two paths; perform service constraint verification on the paths and eliminate candidate paths that fail the verification; link bandwidth resources directly determine the transmission feasibility of data streams; select K optimal candidate paths from the full set of paths P to form the service customized path set P i ; Traffic request d i,j Strictly follow the single-path allocation principle, and its mathematical constraints are defined as: Among them, the binary decision variable x i (d) Characterizes the path selection state: If traffic request d selects path p as the transmission channel, the corresponding variable is assigned the value 1; otherwise, the variable associated with the unselected path takes the value 0. This constraint mechanism ensures that each traffic request is transmitted only through a single path in the candidate path set, resulting in the traffic demand constraint matrix D' = D·x; For any node, use Represents the node information, where ρ L Indicates the link utilization from the previous node to the current node.

6. The deterministic network congestion avoidance traffic routing scheduling method according to claim 5, characterized in that: In step S2, the routing strategy is calculated using the deep Q network DQN, and the state S input to the DQN is defined as a vector that reflects the current network topology, traffic information, and node information. The traffic information includes the traffic demand constraint matrix and traffic priority; the node information includes the link utilization, queue depth, and the proportion and change rate of queuing delay. Assume that the current node has n neighbors. Specifically, the state S is expressed as: S={G,D',P,s self ,s1,…,s n } (3) Where P represents traffic priority; s self Indicates the real-time information of the current node, s1 to s n Represents the real-time information of all n neighbor nodes; The input of DQN is the current state S, and the output is the action-value function Q of all actions in that state; the action is represented by selecting an optimal neighbor node and sending traffic in its direction; The action space is defined as the set of all feasible neighbor nodes. When the current node has n neighbor nodes, the action space is: <h2 style=";text-align:left;direction:ltr">A = {a1,a2,…,a<h2 style=";text-align:left;direction:ltr"> n <h2 style=";text-align:left;direction:ltr">} (4) The action function is expressed as selecting the largest Q value in the action space: Where θ is the DQN parameter, Q(s t ,a;θ) predict the benefit of selecting neighbor a; for neighbors that cannot reach the destination node through them, set their Q value to -∞ to avoid being selected; The reward function comprehensively considers the local congestion status, service priority, and path efficiency, and is defined as: R=R S +R C +R H +R D (6) where R S 、R C 、R H 、R l They represent service reward, congestion penalty, path efficiency penalty, and delay penalty, respectively, and are defined as follows: Business rewards are used to reward traffic that is successfully forwarded to the next hop or successfully delivered to the destination node. Rewards are given according to priority P: in, It is an indicator function, which has a value of 1 when the indicated action is completed and 0 when the action fails to complete; Congestion penalty is used to penalize the consumption of bandwidth utilization, queue depth, and queuing delay occupancy of the next-hop node by the decision-making process: R C =-(U next +Q next +C next ) (8) Among them, U next Indicates the bandwidth utilization of the next hop node, Q next Indicates the queue depth ratio of the next hop node, which is calculated by the ratio of the queue depth to the maximum queue capacity. next Indicates the real-time congestion coefficient of the next-hop node; Path efficiency is used to penalize the extra hops caused by the decision: Among them H taken Indicates the actual number of hops of the currently selected path, H max Indicates the maximum number of hops allowed from the source node to the destination node in the network topology, which is used to limit the path length. Latency penalties are used to reduce latency accumulation: where l accumulated represents the cumulative transmission delay of the data packet from the source node to the current node, l next represents the expected transmission delay from the current node to the next hop node, l max Indicates the maximum end-to-end delay threshold allowed for the service.

7. The deterministic network congestion avoidance traffic routing scheduling method according to claim 6, characterized in that: In step S3, a programmable pipeline designed using the P4 language moves key computational logic down to the data plane of the switching device. By embedding registers, meters, and state machines in the ingress and egress pipelines, link metrics collection, real-time congestion factor calculation, and sampling mode switching are performed directly at the hardware level. At the ingress pipeline stage, ordinary data packets are matched according to the preset sampling strategy, and lightweight IFIT headers are dynamically inserted for qualified traffic. In order to balance the weights of static factors and dynamic factors in the real-time congestion coefficient calculation formula (1), the arithmetic logic of the corresponding calculation unit can be modified to adjust the parameters. The egress pipeline stage uses registers and meters to coordinate calculations, analyze and update port status parameters in real time, and, combined with the congestion coefficient calculation method, perform compound operations on static and dynamic factors directly on the data plane. The calculation results are ultimately written into the switch's metadata field, providing real-time input for subsequent intelligent scheduling. At the policy execution level, by deeply integrating the features of the P4 programmable data plane, the optimal neighbor node decision output by DQN reasoning is directly mapped to pipeline-level action instructions; the business priority weight or path efficiency penalty coefficient in the deep reinforcement learning reward function is dynamically tuned by updating the matching-action rules of the P4 table; the edge node encodes the policy decision as a P4 action function, dynamically modifies the forwarding table entries in the packet processing pipeline, and redirects the next hop of the packet to the selected neighbor node.

8. The deterministic network congestion avoidance traffic routing scheduling method according to claim 7, characterized in that: The deep Q network training framework is as follows: edge computing nodes continuously accumulate historical state-action-reward triples through the locally deployed experience replay cache, and incrementally fine-tune the distributed global model based on the transfer learning mechanism.

Citation Information

Cited By

  • Self-adaptive Mesh network architecture construction method for hybrid networking of industrial Internet of Things

    CN120769326A

  • P4-based early flow type identification and flow scheduling method, system and equipment

    CN121418908A

  • Key flow-based SRv6 explicit path intelligent scheduling method and system

    CN121644481A

  • Key flow-based SRv6 explicit path intelligent scheduling method and system

    CN121644481B

  • Congested large flow rerouting method and system based on network state perception

    CN121691185A