Intelligent computing center-oriented multi-level network congestion control method
Through the multi-level network congestion control method of real-time monitoring and dynamic adjustment, the congestion problem caused by time-varying network traffic in the intelligent computing center is solved, and rapid scheduling and global collaborative optimization are achieved, which improves the timeliness and stability of congestion control.
Patent Information
- Application Number
- CN202510717771.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-05-30
AI Technical Summary
The existing technology is difficult to effectively deal with the problem of poor global and single-level network congestion timeliness caused by strong time-variability of network traffic in smart computing centers and many nodes, especially in large network architectures, which are difficult to achieve accurate congestion control.
A multi-level network congestion control method for intelligent computing centers is adopted. By monitoring the state parameters of network traffic at each level in real time, calculating the congestion coefficient and setting dynamic thresholds, single-level congestion control and cross-level collaborative optimization are carried out, and cross-level collaborative optimization is used for cross-level collaborative optimization, and bandwidth allocation, queue weight and path weight are dynamically adjusted.
It realizes rapid scheduling and global collaborative optimization of the intelligent computing center network, can prevent diffusion when single-layer network congestion, accurately adapt to heterogeneous traffic in complex scenarios, and improves the timeliness and stability of congestion control.
Smart Images

Figure CN120378369A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network communication technologies, and particularly to a multi-level network congestion control method for an intelligent computing center. Background Art
[0002] With the rapid development of intelligent computing and cloud computing technologies, the intelligent computing center, as an important platform for centralized management and scheduling of various resources, plays a key role in aspects such as data processing, task scheduling, and resource utilization efficiency. With the rapid development of information technology and the deepening of digital transformation, the computing network scenarios are becoming increasingly diversified and refined, covering multiple fields such as cloud interconnection, large model applications, intelligent applications, and the Internet of Things. These diverse application scenarios pose different requirements for the computing network infrastructure, making the differential characteristics of infrastructure capabilities more obvious. These diverse application scenarios pose various and complex requirements for the computing network infrastructure, resulting in more obvious differential characteristics of infrastructure capabilities. Furthermore, when designing, deploying, and optimizing these infrastructures, it is necessary to fully consider the demand characteristics of specific application scenarios to achieve more efficient, flexible, and customized service support.
[0003] The congestion control service continuously monitors the congestion situation of the network and dynamically adjusts the data sending rate of the sender according to the overall performance of the network, including indicators such as throughput, latency, and packet loss rate, so as to keep the network load within a reasonable range. However, the network traffic in the intelligent computing center is highly time-varying. When the traffic suddenly stops transmitting or a large amount of new traffic joins the transmission queue, the existing rule-based congestion control algorithms are difficult to accurately adjust the congestion control strategy due to the lack of the ability to learn historical experience. The prior art lacks the ability to dynamically adjust paths from a global perspective. For example, multi-path load balancing (such as ECMP) is prone to flow conflicts when the hashing is uneven, exacerbating congestion.
[0004] For example, Chinese Patent with publication number CN118413489A discloses a network congestion scheduling method and system for an AI intelligent computing center. The method includes: Step 1: Duplicate packets and distinguish service packets and control packets; Step 2: Select a queue; Step 3: Update the queue length and calculate the average queue length; Step 4: Match the discard probability through the average queue length; Step 5: Store the discard probability in the ingress pipeline; Step 6: Distinguish admitted traffic and non-admitted traffic; Step 7: Discard non-admitted traffic according to the probability. This invention decouples the congestion detection and feedback mechanism of queue management. By duplicating to generate control packets, congestion detection is achieved by service packets, and the problem of feedback discard probability is achieved by control packets, which not only ensures the normal forwarding of service packets but also realizes fine-grained feedback.
[0005] As disclosed in the Chinese patent with the authorization announcement number CN114760252B, a data center network congestion control method and system are disclosed. The system includes: periodically obtaining the total current packet length of packets sent by each terminal connected to itself in the data center network to the same destination terminal; if it is monitored that the total packet length of the destination terminal exceeds the congestion threshold corresponding to the destination terminal, generating a congestion notification message for the destination terminal; and sending the congestion notification message to each of the terminals connected to itself, so that the terminal obtains the corresponding pause sending time and pauses sending packets during the counting period of the pause sending time. This patent can greatly shorten the feedback time of congestion signals, enable the terminal as the packet sender to respond to network congestion more quickly, and can effectively reduce the sending of additional packets during the convergence process, reduce the queue accumulation at the switch port, and thus improve the efficiency and reliability of the data center network congestion control process.
[0006] The above patents all have the problems raised in this background technology: The network traffic in the intelligent computing center has strong time-varying characteristics, and the number of nodes in each layer of the network is large. It cannot be well hierarchically controlled when encountering global network congestion and single-layer network congestion, resulting in poor timeliness of congestion control and difficulty in coping with the large network architecture of the intelligent computing center. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a multi-level network congestion control method for an intelligent computing center in view of the deficiencies of the prior art.
[0008] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0009] A multi-level network congestion control method for an intelligent computing center includes the following steps:
[0010] Real-time monitor the traffic state parameters of each layer of the intelligent computing center network, where the intelligent computing center network includes: an edge access layer, an aggregation layer, and a core layer;
[0011] Based on the traffic state parameters of each layer of the network, calculate the congestion coefficient of each layer respectively;
[0012] Set the dynamic congestion threshold for each layer of the network;
[0013] When the single-layer congestion coefficient exceeds the dynamic congestion threshold, perform congestion control on each layer of the network respectively;
[0014] When the congestion coefficients of two or more layers exceed the dynamic congestion threshold, perform cross-layer collaborative optimization.
[0015] Furthermore, the specific formula for calculating the congestion coefficient of each layer is:
[0016]
[0017] Among them, E represents the congestion index of the edge access layer, A represents the congestion index of the aggregation layer, C represents the congestion index of the core layer, and BW m and BW c They represent the peak bandwidth in the most recent 10ms window and the currently allocated link bandwidth, respectively. d and Q m Represent the current queue depth and maximum queue depth respectively, ▽P l Indicates the gradient of packet loss rate change, RTT n Represents the normalized round-trip time RTT value, N Q Indicates queue imbalance, CE indicates the proportion of ENC congestion marking, ΔDelay and Delay indicate the delay change in the last second and the current delay respectively, μ indicates delay variance, and BU indicates bandwidth utilization. represents the average bandwidth utilization, F c Indicates the flow collision rate of ECMP hash collisions.
[0018] Furthermore, the dynamic congestion threshold of each layer of the network is specifically formulated as follows:
[0019]
[0020] Among them, T E represents the dynamic congestion threshold of the edge access layer, T A represents the dynamic congestion threshold of the aggregation layer, T C Indicates the core layer dynamic congestion threshold, E h represents the historical ECI mean value for the same period, t represents the hourly time, i represents the i-th hour, σ E represents the ECI standard deviation in the last hour, W represents the sliding window size, represents the queue depth and σ in the sliding window Q represents the standard deviation of the queue depth in the sliding window, BW represents the total bandwidth, λ(o) represents the load factor during the period, σ D represents the latency variance in the last hour, and ▽BW represents the bandwidth fluctuation rate in the last hour.
[0021] Furthermore, the congestion control of each layer of the network specifically includes: dynamically allocating bandwidth of the edge access layer; adjusting weights of priority queues of the aggregation layer; and optimizing multipaths of the core layer.
[0022] Furthermore, the dynamic allocation of the bandwidth of the edge access layer specifically includes the following steps:
[0023] Prioritize traffic based on service type;
[0024] Dynamically allocate bandwidth based on priority tags, and the specific formula is:
[0025]
[0026] Among them, BW x represents the bandwidth allocated for the current service type, N represents the total number of service types, P x represents the priority of the current service type, P y represents the priority of the y-th service type, G represents the GPU utilization rate, and α = 0.8 represents the adjustment factor.
[0027] Furthermore, perform weight adjustment on the priority queue of the aggregation layer, and the specific formula is:
[0028]
[0029] Among them, M m represents the weight of the current queue, M represents the total number of queues, represents the priority of the current queue, represents the priority of the n-th queue, Q m represents the delay of the current queue, Q n represents the delay of the n-th queue, and k = 0.1 represents the delay sensitivity coefficient.
[0030] Furthermore, perform multi-path weight optimization on the core layer, specifically including:
[0031] Perform path weight calculation, and the specific formula is:
[0032]
[0033] Among them, R d represents the weight of the current path, R represents the total number of queues, BWR d represents the bandwidth of the current path, BWR g represents the bandwidth of the g-th path, D d represents the delay of the current path, D g represents the delay of the j-th path;
[0034] Update the path weight every 10 ms.
[0035] Furthermore, the cross-layer collaborative optimization specifically includes the following steps:
[0036] Obtain the network state characteristics of the edge access layer, the queue time series characteristics of the aggregation layer, and the core layer topology characteristics respectively;
[0037] Perform multi-modal feature embedding on the network state characteristics of the edge access layer, the queue time series characteristics of the aggregation layer, and the core layer topology characteristics, and the specific formula is:
[0038] Z = Concat(MLP(X E ), TCN(X A ), GAT(X C ))
[0039] Where Z represents the multimodal feature, Concat represents the concatenation function, MLP represents the fully connected layer, TCN represents the dilated convolutional layer, GAT represents the graph attention, and X E , X A and X C represent the network state feature matrix of the edge access layer, the queue time series feature matrix of the aggregation layer, and the core layer topology feature matrix respectively;
[0040] Input the multimodal feature into the spatio-temporal graph convolutional network;
[0041] Optimize the spatio-temporal graph convolutional network with the objective loss function. The specific formula is:
[0042] L = 0.5L Q + 0.3L f + 0.2L c
[0043] Where L represents the objective loss function, L Q represents the quality of service loss function, L f represents the fairness loss function, L c represents the consistency loss function;
[0044] Output the edge access layer bandwidth allocation matrix, the aggregation layer queue weight matrix, and the core layer path weight matrix respectively.
[0045] Furthermore, the specific formula of the quality of service loss function L Q is:
[0046]
[0047] Where D pred and D real represent the predicted delay and the actual delay respectively, B pred and B real represent the predicted bandwidth utilization rate and the actual bandwidth utilization rate respectively, ||·||2 represents the L2 norm, and cs represents the cosine similarity function;
[0048] The specific formula of the fairness loss function L f is:
[0049]
[0050] Where BW a and BWb represent the bandwidths of the a-th and b-th service types respectively, and P a and P b represent the priorities of the a-th and b-th service types respectively;
[0051] The specific formula of the consistency loss function L c is as follows:
[0052]
[0053] where ▽C and ▽E represent the gradients of the congestion coefficients in the edge access layer and the core layer respectively, represents the time derivative of the congestion coefficient in the aggregation layer, and ||·||1 represents the L1 norm.
[0054] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0055] 1. The present invention provides a dynamic parameter adjustment mechanism. During the calculation of the congestion coefficient, adaptive parameters are set according to historical traffic data, enabling the congestion calculation to be more sensitive to burst traffic.
[0056] 2. The present invention performs hierarchical control on the intelligent computing center network. When congestion occurs in a single-layer network, it can quickly perform scheduling to prevent congestion from spreading to other layers.
[0057] 3. Through cross-layer collaborative optimization, when the intelligent computing center network is globally congested, the present invention can better coordinate the network conditions of other layers, and obtain the collaborative optimization result through a neural network; at the same time, the loss function of the spatio-temporal graph convolutional network is specifically optimized based on the network characteristics of the intelligent computing center, making the output result more accurate and stable.
[0058] 4. Through the triple design of hierarchical calculation of the base value, dynamic real-time correction, and cross-layer collaborative optimization, while ensuring the evaluation accuracy, the present invention effectively balances sensitivity and stability, and can accurately adapt to the complex scenario of concurrent heterogeneous traffic in the intelligent computing center. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objectives, and advantages of the present invention will become more apparent:
[0060] Figure 1 is a schematic flowchart of an embodiment of the present invention;
[0061] Figure 2 is a cross-layer collaborative optimization structure diagram of an embodiment of the present invention;
[0062] Figure 3 is a network topology structure diagram of the intelligent computing center of an embodiment of the present invention;
[0063] Figure 4 This is the structure diagram of the spatio-temporal convolutional network for the embodiments of the present invention. Detailed implementation manners
[0064] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0065] As Figure 1 shown, the multi-level network congestion control method for an intelligent computing center includes the following steps:
[0066] Real-time monitor the traffic state parameters of each level network in the intelligent computing center, where each level network in the intelligent computing center includes: an edge access layer, an aggregation layer, and a core layer;
[0067] Calculate the congestion coefficient of each level based on the traffic state parameters of each level network;
[0068] Set the dynamic congestion threshold for each level network;
[0069] When the congestion coefficient of a single layer exceeds the dynamic congestion threshold, perform congestion control on each level network respectively;
[0070] When the congestion coefficients of two or more levels exceed the dynamic congestion threshold, perform cross-level collaborative optimization.
[0071] The calculation of the congestion coefficient of each level, the specific formula is:
[0072]
[0073] where E represents the congestion index of the edge access layer, A represents the congestion index of the aggregation layer, C represents the congestion index of the core layer, BW m and BW c respectively represent the peak bandwidth within the most recent 10 ms window and the currently allocated link bandwidth, Q d and Q m respectively represent the current queue depth and the maximum queue depth, ▽P l represents the packet loss rate change gradient, RTT n represents the normalized round-trip time RTT value, N Q represents the queue imbalance degree, CE represents the proportion of ENC congestion marks, ΔDelay and Delay respectively represent the delay change amount within the most recent second and the current delay, μ represents the delay variance, BU represents the bandwidth utilization rate, represents the average bandwidth utilization rate, F c represents the flow conflict rate of ECMP hash conflicts.
[0074] The dynamic congestion threshold of each level network, the specific formula is:
[0075]
[0076] Among them, T E represents the dynamic congestion threshold of the edge access layer, T A represents the dynamic congestion threshold of the aggregation layer, T C Indicates the core layer dynamic congestion threshold, E h represents the historical ECI mean value for the same period, t represents the hourly time, i represents the i-th hour, σ E represents the ECI standard deviation in the last hour, W represents the sliding window size, represents the queue depth and σ in the sliding window Q represents the standard deviation of the queue depth in the sliding window, BW represents the total bandwidth, λ(o) represents the load factor during the period, σ D represents the latency variance in the last hour, and ▽BW represents the bandwidth fluctuation rate in the last hour.
[0077] The load factor λ(o) is set according to the load period. When the service is at peak time, λ(o)=1.2 is set. When the service is at low load at night, λ(o)=0.8 is set. When the service is at low load at night, the core layer link delay variance σ D =0.3, bandwidth fluctuation ▽BW=0.15, F c Flow conflict rate = 5%, then the core layer dynamic congestion threshold When the congestion coefficient exceeds the threshold, core layer congestion control is triggered.
[0078] The congestion control of each layer of the network specifically includes: dynamically allocating bandwidth of the edge access layer; adjusting weights of priority queues of the aggregation layer; and optimizing multipaths of the core layer.
[0079] The dynamic allocation of bandwidth of the edge access layer specifically includes the following steps:
[0080] Prioritize traffic based on service type;
[0081] Dynamically allocate bandwidth based on priority marking. The specific formula is:
[0082]
[0083] Among them, BW x Indicates the bandwidth allocated to the current service type, N indicates the total number of service types, P x Indicates the priority of the current service type, P y represents the priority of the yth service type, G represents the GPU utilization, and α=0.8 represents the adjustment factor.
[0084] The operations of the intelligent computing center can be classified according to the types of computing tasks into:
[0085] AI model training: Requires large-scale distributed computing and high-bandwidth network; Typical loads: Training of large NLP models (such as GPT series), training of computer vision models.
[0086] AI inference service: Low latency requirement, elastic resource scheduling; Typical scenarios: Real-time image recognition, intelligent customer service.
[0087] Scientific computing (HPC): Double-precision floating-point operation ability, MPI communication-intensive; Typical applications: Weather simulation, molecular dynamics calculation.
[0088] Big data analysis: Storage-computation separation architecture, high IO throughput requirement; Typical scenarios: User behavior analysis, log processing.
[0089] When the current operations include real-time image recognition, user behavior analysis, training of large NLP models, and molecular dynamics calculation, set the priorities of user behavior analysis and training of large NLP models with high-bandwidth requirements to 2, and set real-time image recognition and molecular dynamics calculation to 1. When the GPU utilization rate is 50%, and the total bandwidth is 1000 GBps, the bandwidth is 1000 / (1 + 1 + 0.574 + 0.574) ≈ 317.7 GBps.
[0090] The weight adjustment of the priority queue of the aggregation layer is as follows, and the specific formula is:
[0091]
[0092] where M m represents the weight of the current queue, M represents the total number of queues, represents the priority of the current queue, represents the priority of the nth queue, Q m represents the latency of the current queue, Q n represents the latency of the nth queue, and k = 0.1 represents the latency sensitivity coefficient.
[0093] The multi-path weight optimization of the core layer includes:
[0094] Calculate the path weight, and the specific formula is:
[0095]
[0096] where R d represents the weight of the current path, R represents the total number of queues, BWR d represents the bandwidth of the current path, BWR g represents the bandwidth of the gth path, D dRepresents the current path delay, D g Represents the j-th path delay;
[0097] The path weight is updated every 10 ms.
[0098] As Figure 2 shown, the cross-layer collaborative optimization specifically includes the following steps:
[0099] Obtain the network state characteristics of the edge access layer, the queue time series characteristics of the aggregation layer, and the core layer topology characteristics respectively;
[0100] Perform multi-modal feature embedding on the network state characteristics of the edge access layer, the queue time series characteristics of the aggregation layer, and the core layer topology characteristics. The specific formula is:
[0101] Z = Concat(MLP(X E ), TCN(X A ), GAT(X C ))
[0102] Among them, Z represents the multi-modal feature, Concat represents the connection function, MLP represents the fully connected layer, TCN represents the dilated convolutional layer, GAT represents the graph attention, X E 、X A and X C respectively represent the network state feature matrix of the edge access layer, the queue time series feature matrix of the aggregation layer, and the core layer topology feature matrix;
[0103] Input the multi-modal feature into the spatio-temporal graph convolutional network;
[0104] Optimize the spatio-temporal graph convolutional network with the target loss function. The specific formula is:
[0105] L = 0.5L Q + 0.3L f + 0.2L c
[0106] Among them, L represents the target loss function, L Q represents the quality of service loss function, L f represents the fairness loss function, L c represents the consistency loss function;
[0107] Output the edge access layer bandwidth allocation matrix, the aggregation layer queue weight matrix, and the core layer path weight matrix respectively.
[0108] The specific formula of the quality of service loss function L Q is:
[0109]
[0110] Among them, D pred and D real respectively represent the predicted latency and the actual latency, B pred and B real respectively represent the predicted bandwidth utilization rate and the actual bandwidth utilization rate, ||·||2 represents the L2 norm, and cs represents the cosine similarity function;
[0111] The specific formula of the fairness loss function L f is as follows:
[0112]
[0113] Among them, BW a and BW b respectively represent the bandwidths of the a-th and b-th service types, P a and P b respectively represent the priorities of the a-th and b-th service types;
[0114] The specific formula of the consistency loss function L c is as follows:
[0115]
[0116] Among them, ▽C and ▽E respectively represent the gradient of the congestion coefficient of the edge access layer and the gradient of the congestion coefficient of the core layer, represents the time derivative of the congestion coefficient of the aggregation layer, and ||·||1 represents the L1 norm.
[0117] As Figure 3 shown, the three-layer network architecture of the intelligent computing center includes an aggregation switch group, a core switch group, and an access switch group, and the access switches are connected to the servers.
[0118] Among them, the core switch group: As the network traffic hub, it undertakes the horizontal high-speed interconnection requirements of cross-regional data centers, cloud computing resource pools, and external networks, and supports high-bandwidth services such as AI training and inference. The aggregation switch group serves as the policy execution center: realizes inter-VLAN routing, ACL access control, and service quality (QoS) marking, and divides independent logical channels for different AI services (such as isolation between the training cluster and the inference cluster). The access switch group is directly connected to the servers: connects devices such as GPU servers and storage nodes through high-speed ports, provides an access capacity of ≥200 Gbps per port, and meets the high-concurrency communication requirements of AI computing power nodes. And each access switch is connected to two aggregation switches through two independent uplink links, forming a topology structure without single points of failure.
[0119] As Figure 4As shown, the simple network structure of the spatio-temporal graph convolutional network consists of three parts: Normalization: Normalize the input data. Spatio-temporal variation: Through multiple STGCN blocks, where GCN and TCN are alternately used in each block. Output: Use average pooling and fully connected layers to classify the features. Finally, output the scheduling strategies for the aggregation layer, core layer, and edge access layer.
[0120] The computer-readable storage medium of this embodiment can be the internal storage unit of the terminal, such as the hard disk or memory of the terminal; the computer-readable storage medium of this embodiment can also be the external storage device of the terminal, such as the plug-in hard disk, smart memory card, secure digital card, flash card, etc. equipped on the terminal; further, the computer-readable storage medium can also include both the internal storage unit and the external storage device of the terminal.
[0121] The computer-readable storage medium of this embodiment is used to store computer programs and other programs and data required by the terminal, and the computer-readable storage medium can also be used to temporarily store the data that has been output or will be output.
[0122] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0123] The examples described in the present invention are only descriptions of the preferred embodiments of the present invention, and do not limit the concept and scope of the present invention. Without departing from the design concept of the present invention, various deformations and improvements made by those skilled in the art to the technical solutions of the present invention should fall within the protection scope of the present invention.
Claims
1. A multi-level network congestion control method for intelligent computing centers, characterized in that, It includes the following steps: Real-time monitor the traffic status parameters of each layer of the intelligent computing center network. Among them, each layer of the intelligent computing center network includes: the edge access layer, the aggregation layer, and the core layer; Calculate the congestion coefficient of each layer based on the traffic status parameters of each layer; Set the dynamic congestion threshold for each layer of the network; When the congestion coefficient of a single layer exceeds the dynamic congestion threshold, perform congestion control on each layer of the network respectively; When the congestion coefficients of two or more layers exceed the dynamic congestion threshold, perform cross-layer collaborative optimization.
2. The method according to claim 1, wherein The specific formula for calculating the congestion coefficient of each layer is: Among them, E represents the edge access layer congestion index, A represents the aggregation layer congestion index, C represents the core layer congestion index, BW m and BW c respectively represent the peak bandwidth and the currently allocated link bandwidth within the most recent 10 ms window, Q d and Q m respectively represent the current queue depth and the maximum queue depth, represents the packet loss rate change gradient, RTT n represents the normalized round-trip time RTT value, N Q represents the queue imbalance degree, CE represents the proportion of ENC congestion marks, ΔDelay and Delay respectively represent the amount of delay change and the current delay within the most recent second, μ represents the delay variance, BU represents the bandwidth utilization rate, represents the average bandwidth utilization rate, F c represents the flow conflict rate of ECMP hash conflicts.
3. The method according to claim 1, wherein The specific formula for the dynamic congestion threshold of each layer of the network is: Among them, T E represents the dynamic congestion threshold of the edge access layer, T A represents the dynamic congestion threshold of the aggregation layer, T C represents the dynamic congestion threshold of the core layer, E h represents the average value of historical ECI in the same period, t represents the hourly time, i represents the i-th hour, σ E represents the standard deviation of ECI in the most recent hour, W represents the sliding window size, represents the sum of queue depths within the sliding window, σ Q represents the standard deviation of queue depths within the sliding window, BW represents the total bandwidth, λ(o) represents the period load factor, σ D represents the variance of latency in the most recent hour, represents the bandwidth volatility in the most recent hour.
4. The method according to claim 1, characterized in that, The congestion control performed on each layer of the network respectively specifically includes: dynamically allocating the bandwidth of the edge access layer; adjusting the weights of the priority queues of the aggregation layer; performing multi-path optimization on the core layer.
5. The method according to claim 4, wherein The dynamic allocation of the bandwidth of the edge access layer specifically includes the following steps: Mark the traffic priority based on the service type; Dynamically allocate the bandwidth based on the priority mark. The specific formula is: Among them, BW x represents the bandwidth allocated for the current service type, N represents the total number of service types, P x represents the priority of the current service type, P y represents the priority of the y-th service type, G represents the GPU utilization rate, and α = 0.8 represents the adjustment factor.
6. The method according to claim 5, characterized in that The specific formula for adjusting the weights of the priority queues of the aggregation layer is: Among them, M m represents the weight of the current queue, M represents the total number of queues, represents the priority of the current queue, represents the priority of the nth queue, Q m represents the delay of the current queue, Q n represents the delay of the nth queue, and k = 0.1 represents the delay sensitivity coefficient.
7. The method according to claim 5, characterized in that, The multi-path weight optimization performed on the core layer specifically includes: Perform path weight calculation. The specific formula is: Among them, R d represents the weight of the current path, R represents the total number of queues, BWR d represents the bandwidth of the current path, BWR g represents the bandwidth of the g-th path, D d represents the delay of the current path, D g represents the delay of the j-th path; Update the path weight every 10 ms.
8. The method according to claim 1, characterized in that, The cross-layer collaborative optimization specifically includes the following steps: Obtain the network state characteristics of the edge access layer, the queue time series characteristics of the aggregation layer, and the core layer topology characteristics respectively; Perform multi-modal feature embedding on the network state characteristics of the edge access layer, the queue time series characteristics of the aggregation layer, and the core layer topology characteristics. The specific formula is: Z = Concat(MLP(X E ), TCN(X A ), GAT(X C )) Among them, Z represents multi-modal features, Concat represents the concatenation function, MLP represents the fully connected layer, TCN represents the dilated convolutional layer, GAT represents graph attention, and X E 、X A and X C respectively represent the network state feature matrix of the edge access layer, the queue time series feature matrix of the aggregation layer, and the core layer topology feature matrix; Input the multi-modal features into the spatio-temporal graph convolutional network; Optimize the target loss function of the spatio-temporal graph convolutional network. The specific formula is: L = 0.5L Q + 0.3L f + 0.2L c Among them, L represents the target loss function, L Q represents the quality of service loss function, L f represents the fairness loss function, L c represents the consistency loss function; Output the edge access layer bandwidth allocation matrix, the aggregation layer queue weight matrix, and the core layer path weight matrix respectively.
9. The method according to claim 8, characterized in that, The service quality loss function L Q has the following specific formula: Among them, D pred and D real represent the predicted delay and the actual delay respectively, B pred and B real represent the predicted bandwidth utilization rate and the actual bandwidth utilization rate respectively, ||·||2 represents the L2 norm, and cs represents the cosine similarity function; The fairness loss function L f has the following specific formula: Among them, BW a and BW b respectively represent the bandwidths of the a-th and b-th service types, and P a and P b respectively represent the priorities of the a-th and b-th service types; The consistency loss function L c has the following specific formula: Among them, and respectively represent the congestion coefficient gradient of the edge access layer and the congestion coefficient gradient of the core layer, represents the time derivative of the congestion coefficient of the aggregation layer, and ||·||1 represents the L1 norm.
Citation Information
Patent Citations
Data center network congestion control method and system
CN114760252B
Network congestion scheduling method and system of AI intelligent computing center
CN118413489A
Network congestion control method and related products
CN113726671A
Cross-layer end network cooperative congestion control method based on event driving
CN115865827A
Data congestion control in hierarchical sensor networks
US10298505B1
Cited By
CXL switching structure based on dynamic link prediction and adaptive data optimization method
CN120614310A
Distributed network flow intelligent scheduling system based on edge computing and AI cooperation
CN121664759A