A network congestion control method, system, device and medium
By generating target action vectors based on semantic understanding and reinforcement learning, dynamically adjusting network bandwidth allocation, the shortcomings of TCP congestion control algorithm in bandwidth utilization and rate adjustment are solved, and network resource utilization and throughput are improved.
Patent Information
- Application Number
- CN202510273511.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-10
AI Technical Summary
TCP's traditional congestion control algorithm cannot effectively utilize network bandwidth, and there is a lag in the transmission rate adjustment, and it is impossible to actively predict and avoid congestion, resulting in waste of network resources and reduced training efficiency.
By obtaining the current state vector and historical action vector of the network cluster, input the target model based on semantic understanding and reinforcement learning, generate the target action vector, and send it to the end-side network card in the network cluster to dynamically adjust the bandwidth allocation strategy.
Dynamically adjust traffic scheduling strategies, improve resource utilization, improve network throughput and response speed, and actively predict and avoid network congestion risks.
Smart Images

Figure CN119766736B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of network control technology, and in particular, to a network congestion control method, system, device, and medium. Background Art
[0002] With the rapid development of the Internet, the demand for computing resources has been increasing day by day, network traffic has shown explosive growth, and network congestion control has become the key to ensuring data transmission efficiency and quality. Especially in a distributed training environment, the communication requirements between computing nodes have a crucial impact on the overall training efficiency.
[0003] The Transmission Control Protocol (TCP) is one of the core protocols for reliable data transmission in the Internet, and it plays an important role in solving network congestion. TCP realizes the effective management of network congestion through a series of mechanisms. Although TCP has achieved remarkable achievements in network congestion control, it also faces some challenges and limitations.
[0004] Specifically, the traditional congestion control algorithm of TCP cannot effectively utilize network bandwidth and has a certain lag in quickly adjusting the sending rate. In addition, TCP mainly relies on reactive mechanisms to adjust traffic and cannot actively predict and avoid the occurrence of congestion, resulting in problems such as network resource waste and reduced training efficiency when facing complex network environments and dynamically changing traffic patterns.
[0005] Therefore, how to solve the problem of network congestion and improve resource utilization is an urgent problem for those skilled in the art. Summary of the Invention
[0006] In view of this, one aspect of this application provides a network congestion control method, and the method includes:
[0007] Obtain a state vector for characterizing the current network operation state of the network cluster and a historical action vector for guiding bandwidth resource allocation;
[0008] Input the state vector and the historical action vector into a target model based on semantic understanding and reinforcement learning to obtain a target action vector;
[0009] Send the target action vector to each end-side network card in the network cluster to control the end-side network card to perform data transmission according to the bandwidth allocation strategy corresponding to the target action vector.
[0010] Optionally, the inputting the state vector and the historical action vector into a target model based on semantic understanding and reinforcement learning to obtain a target action vector includes:
[0011] Obtain the current timestamp and the historical feedback reward value of the network cluster;
[0012] Combine the current timestamp, the historical feedback reward value, the state vector, and the historical action vector into a sample quadruple;
[0013] Use the sample quadruple as the input of the target model for calculation to obtain the target action vector.
[0014] Optionally, the using the sample quadruple as the input of the target model for calculation to obtain the target action vector includes:
[0015] Perform an embedding operation on the sample quadruple, and perform a positional embedding operation on the historical feedback reward value, the state vector, and the historical action vector to obtain an embedding vector;
[0016] After concatenating the embedding vectors, perform a Transformer processing and a fully connected layer processing in sequence to obtain the target action vector.
[0017] Optionally, after controlling the end-side network card to perform data transmission according to the bandwidth allocation policy corresponding to the target action vector, it further includes:
[0018] Determine whether the target model reaches a preset convergence condition;
[0019] If the preset convergence condition is not reached, then perform the following steps:
[0020] Obtain the bandwidth utilization percentage of the end-side network card and the queue length ratio of the switch port;
[0021] Determine the current feedback reward value of the network cluster according to the bandwidth utilization percentage and the queue length ratio;
[0022] Optimize the parameters of the target model according to the current feedback reward value.
[0023] Optionally, obtaining a state vector for characterizing the current network operation state of the network cluster includes:
[0024] Obtain the switch information and link information in the network cluster; the switch information at least includes the buffer occupancy rate, port bandwidth utilization rate, and port queue length; the link information at least includes the packet loss rate, transmission delay, and throughput;
[0025] Integrate the switch information and the link information to obtain the state vector.
[0026] Another aspect of the present application provides a network congestion control system, which includes an agent for implementing the network congestion control method described above, and a simulation platform for building the network cluster.
[0027] Optionally, the system further includes: a network topology management module, a traffic generator, and a network monitoring module;
[0028] The network topology management module is used to store the network topology structure information of multiple network clusters;
[0029] The traffic generator is used to generate the data traffic of the network cluster;
[0030] The network monitoring module is used to collect the switch information and link information in the network cluster in real time; and transmit the collected switch information and link information to the agent.
[0031] Another aspect of the present application provides a network congestion control device, which includes:
[0032] A vector acquisition module, which is used to acquire a state vector for characterizing the current network operation state of the network cluster and a historical action vector for guiding bandwidth resource allocation;
[0033] A target vector determination module, which is used to input the state vector and the historical action vector into a target model based on semantic understanding and reinforcement learning to obtain a target action vector;
[0034] A bandwidth allocation module, which is used to send the target action vector to each end-side network card in the network cluster to control the end-side network card to perform data transmission according to the bandwidth allocation policy corresponding to the target action vector.
[0035] Another aspect of the present application provides a network congestion control device, which includes a memory and a processor. A computer program that can run on the processor is stored on the memory. When the processor executes the program, the steps of the network congestion control method are implemented.
[0036] Another aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the network congestion control method are implemented.
[0037] A network congestion control method, system, device and medium provided by this application have the following beneficial effects: Therefore, according to different network states and historical action decisions, the traffic scheduling strategy is dynamically adjusted, improving resource utilization while increasing network throughput and response speed. In addition, through the integration of semantic understanding and reinforcement learning, semantic information in the state vector and historical action vector is extracted by the semantic understanding model, and the extracted information is used to guide the intelligent decision-making process of the reinforcement learning model, that is, the target action vector generated by the target model is used to guide the bandwidth resource allocation, realizing the active prediction and avoidance of network congestion risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic flowchart of a network congestion control method provided by an embodiment of this application;
[0039] Figure 2 It is a schematic structural diagram of a network cluster architecture provided by an embodiment of this application;
[0040] Figure 3 It is a schematic flowchart of a network congestion control method provided by another embodiment of this application;
[0041] Figure 4 It is a schematic architecture diagram of a target model based on semantic understanding and reinforcement learning provided by an embodiment of this application;
[0042] Figure 5 It is a schematic structural diagram of a network congestion control system provided by an embodiment of this application;
[0043] Figure 6 It is a schematic structural diagram of a network congestion control device provided by an embodiment of this application;
[0044] Figure 7 It is a schematic structural diagram of a network congestion control device provided by another embodiment of this application.
[0045] The reference numerals are as follows: 50 is an agent, 51 is a simulation platform, 52 is a network topology management module, 53 is a traffic generator, 54 is a network monitoring module, 70 is a memory, 71 is a processor, 72 is a display screen, 73 is an input / output interface, 74 is a communication interface, 75 is a power supply, 76 is a communication bus, 701 is a computer program, 702 is an operating system, and 703 is data. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0047] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0048] Figure 1 The flowchart of a network congestion control method provided for the embodiments of this application is as Figure 1 shown, and the method includes:
[0049] S10: Obtain a state vector for characterizing the current network operating state of the network cluster, and a historical action vector for guiding bandwidth resource allocation;
[0050] In a specific embodiment, collect the metric data characterizing the network operating state in the network cluster, and integrate the collected data into a state vector. Thus, the current network motion state in the network cluster can be reflected by this state vector. Among them, the metric data can be at least one type of data for characterizing the current network operating state. In an optional embodiment, the metric data may include, but is not limited to, the switch buffer occupancy rate, the switch port bandwidth utilization rate, the link packet loss rate, and the transmission delay.
[0051] In addition, it is also necessary to obtain a historical action vector for guiding bandwidth resource allocation. In an optional embodiment, the historical action vector can be stored in a specified database and can be directly retrieved from the database when needed.
[0052] It should be noted that when obtaining the state vector and the historical action vector, it can be obtained in real time. Of course, it can be understood that obtaining in real time will increase the pressure on computing resources. Therefore, in a preferred embodiment, it can be obtained once every preset period. For example, it can be obtained once every 10 seconds, that is, the dynamic guidance of the bandwidth allocation policy is performed once every preset period.
[0053] Among them, the historical action vector and the target action vector refer to vectors that can be used to guide the bandwidth resource allocation of each end-side network card in the network cluster. Therefore, the dimension of the action vector is the same as the number of end-side network cards, and the value of each dimension represents the bandwidth allocation percentage of the corresponding end-side network card.
[0054] Figure 2 FIG. is a schematic structural diagram of a network cluster architecture provided by an embodiment of the present application. For ease of understanding, the following will be combined with Figure 2 for illustration. As Figure 2 shown, in the current network cluster architecture, it includes multiple servers, a first switch, and a second switch, and each server includes 4 end-side network cards.
[0055] For example, as Figure 2 shown, one of the servers includes 4 end-side network cards (NICs), specifically including NIC 0, NIC 1, NIC 2, and NIC 3, and the total bandwidth allocated to each end-side network card based on the target action vector is: NIC 0 is 30%, NIC 1 is 35%, NIC 2 is 20%, and NIC 3 is 15%. And the dimension of the target action vector is the same as the number of end-side network cards, that is, it includes 3 dimensions. Thus, the target action vector can be expressed as [30%, 35%, 20%, 15%].
[0056] It should be noted that the historical action vector can be the action vector issued last time, or the average value of historical multiple action vectors. The present application does not make any limitations in this regard. In addition, it should also be noted that the present application does not limit the number of servers, end-side network cards, and switches in the network cluster. In an alternative embodiment, Figure 2 the first switch in
[0057] S11: Input the state vector and the historical action vector into the target model based on semantic understanding and reinforcement learning to obtain the target action vector;
[0058] Furthermore, after obtaining the current state vector and historical action vector of the network cluster, the state vector and the historical action vector are used as inputs to the target model for calculation, so as to obtain the target action vector at the current moment. It can be understood that the target model is used to optimize the action vector according to the historical action vector and the state vector representing the current network operation state, so as to realize the dynamic adjustment of the network cluster traffic.
[0059] Among them, the target model refers to a model that combines a large semantic understanding model and a reinforcement learning module. The introduction of the large semantic understanding model can provide the reinforcement learning model with richer context information and more powerful representation capabilities, thereby guiding the reinforcement learning model to learn the optimal bandwidth allocation strategy. Thus, the best action guidance strategy in the current network state can be obtained, that is, the target action vector.
[0060] It can be understood that in specific embodiments, the network cluster environment is complex and changeable, and traditional learning reinforcement algorithms are prone to overfitting and difficult to adapt to various scenarios. In the embodiments of the present application, the large semantic understanding model is fused with the reinforcement learning model. The introduction of the large semantic understanding model provides rich context information for the reinforcement learning model, enabling it to deeply understand and flexibly respond to complex network environments, ensuring stable performance in diverse network scenarios. Thus, through joint modeling and collaborative optimization, the performance and stability of the congestion control algorithm are improved, providing a new solution for the efficient operation of the network cluster.
[0061] S12: Send the target action vector to each end-side network card in the network cluster to control the end-side network card to perform data transmission according to the bandwidth allocation strategy corresponding to the target action vector.
[0062] Furthermore, send the target action vector output by the target model to each end-side network card in the network cluster to guide the end-side network card to send data packets according to the specified bandwidth percentage, thereby effectively controlling network congestion. That is, based on the target action vector, reasonable allocation of bandwidth resources is achieved.
[0063] In an alternative embodiment, the target action vector can be sent through a network management protocol or a dedicated control interface to ensure accurate transmission and execution of the instructions.
[0064] Thus, the network congestion control method provided by the embodiments of the present application dynamically adjusts the traffic scheduling strategy according to different network states and historical action decisions, improving resource utilization while increasing network throughput and response speed. In addition, through the fusion of semantic understanding and reinforcement learning, semantic information in the state vector and historical action vector is extracted by the semantic understanding model, and the extracted information is used to guide the intelligent decision-making process of the reinforcement learning model, that is, the target action vector generated by the target model guides the allocation of bandwidth resources to achieve active prediction and avoidance of network congestion risks.
[0065] Figure 3 It is a schematic flowchart of a network congestion control method provided by another embodiment of the present application. In an alternative embodiment, as Figure 3 shown, input the state vector and the historical action vector into a target model based on semantic understanding and reinforcement learning to obtain the target action vector, including:
[0066] S30: Obtain the current timestamp and the historical feedback reward value of the network cluster;
[0067] S31: Combine the current timestamp, the historical feedback reward value, the state vector, and the historical action vector into a sample quadruple;
[0068] S32: Use the sample quadruple as the input of the target model for calculation to obtain the target action vector.
[0069] In a specific embodiment, when calculating the target action vector through the target model, in addition to the state vector and the historical action vector, it is also necessary to obtain the current timestamp at the current moment and the historical feedback reward value of the network cluster.
[0070] Among them, the feedback reward value refers to the index data used to characterize the current network performance. For example, latency, packet loss rate, throughput, resource utilization rate, and fairness index data, etc. Through the historical feedback reward value, the impact of the historical action vector on the network state of the network cluster can be evaluated, and the strategy can be adjusted accordingly to optimize the bandwidth resource allocation in the network cluster.
[0071] Furthermore, combine the current timestamp, the historical feedback reward value, the state vector, and the historical action vector to obtain the input sample of the target model, that is, the sample quadruple. In an optional embodiment, the sample quadruple can be expressed as: , where is the current timestamp, is the historical action vector, is the state vector obtained at the current moment, is the historical feedback reward value. Among them, it should be noted that when initially controlling network congestion, a historical action vector can be randomly generated as the input of the target model.
[0072] Furthermore, input the sample quadruple into the target model for calculation to obtain the target action vector that can guide the bandwidth allocation in the network cluster at the next moment.
[0073] Based on the above embodiments, as Figure 3 shown, using the sample quadruple as the input of the target model for calculation to obtain the target action vector includes:
[0074] S320: Perform an embedding operation on the sample quadruple, and perform a positional embedding operation on the historical feedback reward value, the state vector, and the historical action vector to obtain an embedding vector;
[0075] Figure 4 is a schematic diagram of the architecture of a target model provided by an embodiment of the present application. In a specific embodiment, the sample quadruple Input into the target model architecture as shown below for calculation. Figure 4 Specifically, as shown in Figure 4 , the target model separately performs embedding operations on the current timestamp , the historical action vector , the state vector , and the historical feedback reward value . At the same time, position embedding operations are performed on the historical action vector , the state vector , and the historical feedback reward value to obtain the embedding vectors.
[0076] Specifically, as Figure 4 shown, the target model separately performs embedding operations on the current timestamp , the historical action vector , the state vector , and the historical feedback reward value , and at the same time, position embedding operations are performed on the historical action vector , the state vector , and the historical feedback reward value to obtain the embedding vectors.
[0077] Among them, Embedding refers to the technology used to convert high-dimensional data (such as text, images, speech, etc.) into low-dimensional, dense vector representations. And Position Embedding is a special embedding used to capture the position information of each element in the sequence data. When processing sequence data (such as text, time series, etc.), the position information is very important because the same content may have different semantics at different positions.
[0078] S321: After concatenating the embedding vectors, perform Transformer processing and fully connected layer processing in sequence to obtain the target action vector.
[0079] Furthermore, the embedding vectors obtained after the embedding operation are concatenated and used as the input of the Transformer processing as shown in Figure 4 . The output after Transformer processing is used as the output of the fully connected layer processing. After passing through the fully connected layer processing, the target action vector is obtained. Figure 4 Among them, as shown in Figure 4 , the Transformer processing includes the multi-head attention mechanism (Multi-head Attention), addition and normalization (Add&Norm), and the feed forward neural network (Feed Forward Neural Network). . Figure 4 Among them, as shown in Figure 4 , the Transformer processing includes the multi-head attention mechanism (Multi-head Attention), addition and normalization (Add&Norm), and the feed forward neural network (Feed Forward Neural Network).
[0080] Thus, by constructing a target model framework based on semantic understanding and reinforcement learning, the state vector and historical action vector representing the network state are input into the target model to generate an accurate network state representation and action sequence prediction, realizing efficient congestion control of the network.
[0081] As an alternative embodiment, in order to further improve the reliability of congestion control, after the control-side network card transmits data according to the bandwidth allocation policy corresponding to the target action vector, the following steps are further included:
[0082] Determine whether the target model reaches a preset convergence condition;
[0083] If the preset convergence condition is not reached, then perform the following steps:
[0084] Obtain the bandwidth utilization percentage of the end-side network card and the queue length ratio of the switch port;
[0085] Determine the current feedback reward value of the network cluster according to the bandwidth utilization percentage and the queue length ratio;
[0086] Optimize the parameters of the target model according to the current feedback reward value.
[0087] In a specific embodiment, the target model can be continuously optimized to improve the network congestion control performance. Specifically, after the control-side network card transmits data according to the bandwidth allocation policy corresponding to the target action vector, first determine whether the current target model reaches the preset convergence condition, that is, determine whether the target model has reached the optimal state.
[0088] If the convergence condition is not reached, obtain the current feedback reward value of the network cluster to optimize the parameters of the target model through the current feedback reward value. Specifically, obtain the bandwidth utilization percentage of the end-side network card and the queue length ratio of the switch port, and calculate the current feedback reward value according to the bandwidth utilization percentage and the queue length ratio. The specific calculation formula is formula (1):
[0089] (1)
[0090] Where, is the number of end-side network cards, is the bandwidth allocated to the th end-side network card, is the maximum bandwidth of the end-side network card, is the penalty weight, which is used to adjust the penalty for the congestion degree in the reward formula, is the queue queuing length of the switch port corresponding to the th end-side, is the maximum queue length of the switch port, is the bandwidth utilization percentage, is the queue length ratio.
[0091] In a specific embodiment, the calculation of the current feedback reward value aims to encourage the improvement of bandwidth utilization and the reduction of queue length. When the percentage of bandwidth utilization is higher, the current feedback reward value is larger. The longer the queue, the more it indicates that network congestion is occurring currently. According to the change of the current feedback reward value, the parameters of the target model can be iteratively optimized. Through continuous iterative training, the target model can gradually improve its congestion control performance in different network scenarios.
[0092] It should be noted that the preset convergence condition can be that the change of the loss curve of the target model tends to be stable. In another alternative embodiment, it can also be that the loss values calculated any two times do not exceed a threshold. When the target model reaches the preset convergence condition, the target model is not updated iteratively. In addition, it should also be noted that when the target model is updated iteratively, in an alternative embodiment, the target model can be iteratively updated through the temporal difference algorithm.
[0093] In an alternative embodiment, the switch information and link information in the network cluster are obtained; the switch information at least includes the buffer occupancy rate, port bandwidth utilization rate, and port queue length; the link information at least includes the packet loss rate, transmission delay, and throughput;
[0094] The switch information and link information are integrated to obtain a state vector.
[0095] In an alternative embodiment, it is possible to collect the switch information and link information through a network monitoring tool or a device management interface to provide basic information support for subsequent congestion control. Among them, the switch information includes, but is not limited to, the buffer occupancy rate, port bandwidth utilization rate, and port queue length of the switch. The link information includes, but is not limited to, key indicators such as the packet loss rate, transmission delay, and throughput of the link.
[0096] Among them, the buffer occupancy rate of the switch refers to the ratio of the number of data packets stored in the switch queue to the total capacity of the queue. The port bandwidth utilization rate refers to the bandwidth usage of the port at the current moment, which can be expressed as a percentage.
[0097] Furthermore, in an alternative embodiment, after obtaining the switch information and link information initially collected by the network monitoring tool or the device management interface, data cleaning is performed on them to eliminate noise and outliers in the data. In addition, standardization processing is also performed on them to convert the data into a data format suitable for processing by the target model. Specifically, it can include, but is not limited to, converting percentages to decimal forms and normalizing discrete values such as queue lengths. Thus, the quality and consistency of the obtained information are ensured to facilitate the subsequent construction of the state vector.
[0098] For example, in Figure 2In the network cluster shown, the switch information and link information include key information that can characterize the network operation status, such as the number of the first switch and the second switch, port connection conditions, the bandwidth of the uplink and downlink, and the length of the switch buffer queue.
[0099] Furthermore, the switch information and link information after preprocessing (including but not limited to data cleaning and standardization processing) are integrated into a state vector, which comprehensively reflects the current network operation status. The construction of the state vector considers multiple dimensions such as the network congestion degree, bandwidth usage, and data transmission efficiency, providing accurate input for the target model.
[0100] In the above embodiments, the network congestion control method is described in detail. The present application also provides an embodiment corresponding to a network congestion control system.
[0101] Figure 5 It is a schematic structural diagram of a network congestion control system provided by an embodiment of the present application. As Figure 5 shown, the system includes an agent 50 that can implement the network congestion control method of any of the above embodiments, and a simulation platform 51 for building a network cluster.
[0102] In an alternative embodiment, the simulation platform 51 can be an OpenAI Gym simulation platform. The construction of the simulation platform 51 is used to build a simulation environment similar to the actual intelligent computing network cluster to realize the training and verification of the agent 50. Various network states and traffic scenarios are simulated in the simulation environment, enabling the target model in the agent 50 to learn and optimize under safe and controllable conditions. Through comparative tests and result analysis, the effectiveness and stability of the target model are evaluated, and the target model is adjusted and optimized according to the feedback to ensure that it can stably and efficiently control network congestion in actual applications.
[0103] In an alternative embodiment, building a simulation environment similar to the actual intelligent computing network cluster includes building virtual network components such as switches, links, and network cards. The simulation environment should be able to simulate real network traffic, congestion situations, and bandwidth allocation behaviors, providing a reliable platform for the training and verification of the target model.
[0104] In a simulation environment, the target model in the agent 50 is trained. By simulating different network states and traffic scenarios, the target model learns how to make optimal bandwidth allocation decisions in various situations. During the training process, the performance metrics of the target model are monitored in real time. For example, the changes in the feedback reward value, the improvement in bandwidth utilization, the reduction in queue length, etc. are monitored to evaluate the effectiveness and stability of the model. After the training is completed, the target model is verified. By comparing and testing it with traditional congestion control algorithms or other reinforcement learning models, the advantages and performance of the model in practical applications are verified.
[0105] In an optional embodiment, a detailed analysis of the simulation results is carried out to identify the advantages and disadvantages of the target model. For example, poor performance in certain specific scenarios, sensitivity to certain network metrics, etc. According to the analysis results, the target model is further optimized and improved. For example, the structure of the target model, parameter settings, or the calculation function of the feedback reward value are adjusted to improve the overall performance and adaptability of the model.
[0106] Thus, the network congestion control system provided by the embodiments of the present application provides strong support for the verification and optimization of the target model based on semantic understanding and reinforcement learning. Specifically, the simulation platform 51 can simulate a real data center network environment, making up for the lack of a unified evaluation platform in the past. Thus, the model can be fully tested and adjusted, and its performance metrics can be evaluated, providing data support for the improvement and optimization of the model.
[0107] In an optional embodiment, as Figure 5 shown, the network congestion control system further includes: a network topology management module 52, a traffic generator 53, and a network monitoring module 54.
[0108] The network topology management module 52 is used to store the network topology structure information of multiple network clusters. Specifically, the network topology management module 52 is responsible for managing and maintaining the topology structure information of the network clusters. The network topology management module 52 can identify and record the nodes, links, and their connection relationships in the network, providing accurate network topology data for other modules.
[0109] During the simulation process, the network topology management module 52 can flexibly adjust and switch different network topology structures according to actual business needs. For example, tree topology, ring topology, mesh topology, etc. That is, the network congestion control provided by the present application can support multiple network topology structures and traffic patterns. It can flexibly handle various network environments. Thus, the network congestion control system provided by the present application can be widely applied to various data center network scenarios, providing effective congestion control for network clusters of different scales and types.
[0110] The traffic generator 53 is used to generate the data traffic of the network cluster. Specifically, the traffic generator 53 can simulate the data traffic in the network cluster. According to the set traffic patterns and parameters, it generates various types of network traffic, such as bulk data transfer, real-time communication traffic, random traffic, etc.
[0111] In an alternative embodiment, the traffic generator 53 supports multiple traffic distribution models, such as Poisson distribution, normal distribution, etc., to meet the requirements of different application scenarios. The generated traffic data is injected into the simulation platform 51 to provide a real traffic environment for the training and verification of congestion control algorithms.
[0112] Therefore, in a specific embodiment, by reasonably configuring the traffic generator 53, various complex network traffic scenarios can be simulated to fully test and evaluate the performance of congestion control algorithms.
[0113] The network monitoring module 54 is used to collect the switch information and link information in the network cluster in real time, and transmit the collected switch information and link information to the agent 50. Specifically, the network monitoring module 54 is an important component in the network congestion control system responsible for real-time monitoring of network status and traffic information. Through the interface connection with network devices such as switches and routers, it collects key metric data in the network, such as the buffer occupancy rate of the switch, port bandwidth utilization rate, port queue length, and important data such as the packet loss rate, transmission delay, and throughput of the link.
[0114] In a specific embodiment, the agent 50 is the key part to implement congestion control decisions, that is, it can implement the network congestion control method in any of the above embodiments. Based on the construction of the target model of semantic understanding and reinforcement learning, the agent 50 takes the real-time network status information and traffic data collected by the network monitoring module 54 as input, generates a global network status representation, combines the historical action vector, calculates the optimal target action vector, and then realizes the efficient scheduling and congestion control of network traffic based on the target action vector.
[0115] In the above embodiments, the network congestion control method has been described in detail. The present application also provides an embodiment corresponding to a network congestion control device.
[0116] Figure 6 As shown in the structural schematic diagram of a network congestion control device provided by an embodiment of the present application, Figure 6 shown, the device includes:
[0117] A vector acquisition module 60, configured to acquire a state vector for characterizing the current network operation state of the network cluster, and a historical action vector for guiding bandwidth resource allocation;
[0118] A target vector determination module 61, configured to input a state vector and a historical action vector into a target model based on semantic understanding and reinforcement learning to obtain a target action vector;
[0119] A bandwidth allocation module 62, configured to send the target action vector to each edge network card in the network cluster to control the edge network card to perform data transmission according to the bandwidth allocation policy corresponding to the target action vector.
[0120] In addition, the network congestion control device provided by the embodiment of the present application further includes:
[0121] A first information acquisition module, configured to acquire a current timestamp and a historical feedback reward value of the network cluster;
[0122] A sample quadruple composition module, configured to combine the current timestamp, the historical feedback reward value, the state vector, and the historical action vector into a sample quadruple;
[0123] A calculation module, configured to use the sample quadruple as an input of the target model for calculation to obtain a target action vector.
[0124] An embedding operation module, configured to perform an embedding operation on the sample quadruple, and perform a positional embedding operation on the historical feedback reward value, the state vector, and the historical action vector to obtain an embedding vector;
[0125] An embedding vector processing module, configured to splice the embedding vectors, and then perform a transformer processing and a fully connected layer processing in sequence to obtain a target action vector.
[0126] A convergence determination module, configured to determine whether the target model reaches a preset convergence condition; if the preset convergence condition is not reached, then call a second information acquisition module, a current feedback reward value determination module, and an optimization module;
[0127] A second information acquisition module, configured to acquire the bandwidth utilization percentage of the edge network card and the queue length occupancy ratio of the switch port;
[0128] A current feedback reward value determination module, configured to determine the current feedback reward value of the network cluster according to the bandwidth utilization percentage and the queue length occupancy ratio;
[0129] An optimization module, configured to optimize the parameters of the target model according to the current feedback reward value.
[0130] A third information acquisition module, configured to acquire switch information and link information in the network cluster; the switch information at least includes a buffer occupancy rate, a port bandwidth utilization rate, and a port queue length; the link information at least includes a packet loss rate, a transmission delay, and a throughput;
[0131] An integration module, configured to integrate the switch information and the link information to obtain a state vector.
[0132] Figure 7 The following is a schematic structural diagram of a network congestion control device provided by another embodiment of this application. As Figure 7 shown, the network congestion control device includes: a memory 70 for storing a computer program;
[0133] a processor 71 for implementing the steps of the network congestion control method mentioned in the above embodiment when executing the computer program.
[0134] The network congestion control device provided in this embodiment may include, but is not limited to, a laptop computer or a desktop computer, etc.
[0135] Among them, the processor 71 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 71 may be implemented in at least one hardware form of a digital signal processor (DSP for short), a field-programmable gate array (FPGA for short), or a programmable logic array (PLA for short). The processor 71 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 71 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 71 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.
[0136] The memory 70 may include one or more computer-readable storage media, which may be non-transitory. The memory 70 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 70 is at least used to store the following computer program 701. After the computer program is loaded and executed by the processor 71, the relevant steps of the network congestion control method disclosed in any of the foregoing embodiments can be implemented. In addition, the resources stored in the memory 70 may also include an operating system 702 and data 703, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 702 may include Windows, Unix, Linux, etc. The data 703 may include, but is not limited to, the relevant data involved in the network congestion control method.
[0137] In some embodiments, the network congestion control device may further include a display screen 72, an input / output interface 73, a communication interface 74, a power supply 75, and a communication bus 76.
[0138] Those skilled in the art can understand that Figure 7 the structure shown in does not constitute a limitation on the network congestion control device, and may include more or fewer components than shown in the figure.
[0139] The network congestion control device provided by the embodiment of the present application includes a memory and a processor. When the processor executes the program stored in the memory, the network congestion control method in the above embodiment can be implemented.
[0140] It should be noted that although the operations are depicted in a specific order in the drawings, this should not be construed as requiring the operations to be performed in the specific order shown or sequentially, or requiring all of the illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. In addition, the separation of the various system modules and components in the above embodiments should not be understood as required in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
Claims
1. A network congestion control method, characterized in that: The method comprises: Obtaining a state vector for characterizing the current network operation state of the network cluster and a historical action vector for guiding bandwidth resource allocation; Inputting the state vector and the historical action vector into a target model based on semantic understanding and reinforcement learning to obtain a target action vector; Sending the target action vector to each end-side network card in the network cluster to control the end-side network card to perform data transmission according to the bandwidth allocation strategy corresponding to the target action vector; The step of inputting the state vector and the historical action vector into a target model based on semantic understanding and reinforcement learning to obtain a target action vector includes: Get the current timestamp and the historical feedback reward value of the network cluster; Combining the current timestamp, the historical feedback reward value, the state vector, and the historical action vector into a sample quadruple; The sample quadruple is used as an input of the target model to perform calculations to obtain the target action vector.
2. The network congestion control method according to claim 1, characterized in that: The step of calculating the target action vector by using the sample quadruple as an input of the target model includes: Performing an embedding operation on the sample quadruple, and performing a position embedding operation on the historical feedback reward value, the state vector, and the historical action vector to obtain an embedding vector; After concatenating the embedding vectors, transformer processing and fully connected layer processing are performed in sequence to obtain the target action vector.
3. The network congestion control method according to claim 1, characterized in that: After controlling the end-side network card to perform data transmission according to the bandwidth allocation strategy corresponding to the target action vector, the method further includes: Determining whether the target model reaches a preset convergence condition; If the preset convergence condition is not met, perform the following steps: Obtaining the bandwidth utilization percentage of the end-side network card and the queue length ratio of the switch port; Determining a current feedback reward value of the network cluster according to the bandwidth utilization percentage and the queue length ratio; Optimizing the parameters of the target model according to the current feedback reward value.
4. The network congestion control method according to claim 1, characterized in that: Get the state vector used to characterize the current network operation status of the network cluster, including: Obtaining switch information and link information in the network cluster; the switch information at least includes buffer occupancy, port bandwidth utilization and port queue length; the link information at least includes packet loss rate, transmission delay and throughput; The switch information and the link information are integrated to obtain the state vector.
5. A network congestion control system, characterized in that: The system includes an intelligent agent that implements the network congestion control method described in any one of claims 1 to 4, and a simulation platform for building the network cluster.
6. The network congestion control system according to claim 5, characterized in that: The system also includes: a network topology management module, a traffic generator and a network monitoring module; The network topology management module is used to store network topology information of a plurality of network clusters; The traffic generator is used to generate data traffic of the network cluster; The network monitoring module is used to collect the switch information and link information in the network cluster in real time; and transmit the collected switch information and link information to the intelligent agent.
7. A network congestion control device, characterized in that: The device comprises: A vector acquisition module is used to acquire a state vector for representing the current network operation state of the network cluster, and a historical action vector for guiding bandwidth resource allocation; A target vector determination module, used for inputting the state vector and the historical action vector into a target model based on semantic understanding and reinforcement learning to obtain a target action vector; A bandwidth allocation module, used to send the target action vector to each end-side network card in the network cluster, so as to control the end-side network card to perform data transmission according to the bandwidth allocation strategy corresponding to the target action vector; The first information acquisition module is used to obtain the current timestamp and the historical feedback reward value of the network cluster; The sample quadruple composition module is used to combine the current timestamp, historical feedback reward value, state vector and historical action vector into a sample quadruple; The calculation module is used to calculate the sample quadruple as the input of the target model to obtain the target action vector.
8. A network congestion control device, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps of the network congestion control method according to any one of claims 1 to 4 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the network congestion control method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Cloud resource scheduling performance bottleneck prediction method based on reinforcement learning
CN112422651A
Data transmission method and system of industrial edge gateway
CN119316424A