Traffic scheduling method and electronic equipment
By generating candidate port values and probe packets at the data sending end, and combining them with a reinforcement learning model to select the optimal source port value, the problem of unbalanced network load in multi-path networks is solved, achieving dynamic load balancing and efficient network transmission.
Patent Information
- Application Number
- CN202511613565.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-06
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-11-06
AI Technical Summary
In existing technologies, the problem of unbalanced multi-path network load leads to congestion on some links, decreased bandwidth utilization, increased network transmission latency, and packet loss. Existing solutions are inflexible in deployment, costly, and unable to adapt to dynamic traffic patterns in real time.
Multiple candidate port values are generated at the data sending end. Performance metrics of candidate network paths are obtained by probing data packets. The optimal source port value is evaluated and selected based on a reinforcement learning model. The source port field of the data packets is dynamically adjusted to optimize path selection.
It achieves dynamic and adaptive load balancing, avoids the negative impact of hash collisions, improves network bandwidth utilization and transmission efficiency, reduces deployment costs, and is compatible with existing network devices and applications.
Smart Images

Figure CN121098802A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of traffic scheduling, and in particular to a traffic scheduling method and an electronic device. BACKGROUND
[0002] Network devices rely on hash algorithms to distribute traffic over multiple paths, while hash collisions are inevitable, leading to uneven distribution of network load over multiple paths. This imbalance can cause problems such as congestion on some links, decreased bandwidth utilization, increased network transmission latency, and packet loss, which seriously affect the overall forwarding efficiency and performance of the network.
[0003] To address the problem of uneven load distribution in multi-path networks, related solutions mostly focus on optimizing network core devices (such as switches and routers), such as improving their built-in hash algorithms or implementing weighted cost multi-path routing (WCMP). However, such solutions have inherent limitations, making it difficult to effectively solve the above-mentioned load imbalance problem fundamentally: first, they are not flexible and costly to deploy, requiring upgrading or reconfiguring existing network infrastructure; second, they have poor compatibility and are difficult to deploy uniformly in heterogeneous network environments; most importantly, these solutions are still essentially static or semi-static optimizations, which are difficult to perceive the quality of end-to-end paths in real time, and cannot adapt to dynamically changing network traffic patterns, thus cannot achieve true dynamic load balancing. SUMMARY
[0004] The present application provides a traffic scheduling method and an electronic device to at least solve the technical problem of uneven load distribution in multi-path networks in related technologies.
[0005] The present application provides a traffic scheduling method applied to a data sending end, the traffic scheduling method comprising: In response to a to-be-sent data stream conforming to a transport layer protocol and the traffic size of the to-be-sent data stream being greater than a preset threshold, generating a plurality of candidate port values for identifying the to-be-sent data stream at the data sending end, each candidate port value corresponding to a candidate network path; creating a plurality of probe data packets, wherein the source port field of each probe data packet is set to a different candidate port value; sending the probe data packets to a data receiving end, and obtaining performance indicators of the candidate network paths corresponding to the plurality of probe data packets based on the response of the data receiving end; evaluating the plurality of candidate port values based on the performance indicators, and determining a target source port value from the plurality of candidate port values according to the evaluation result; modifying the source port field of the data packets corresponding to the to-be-sent data stream to the target source port value, and sending the modified data packets corresponding to the to-be-sent data stream to the data receiving end.
[0006] The application further provides an electronic device, comprising a memory for storing a computer program, and a processor for executing the computer program to implement the steps of the traffic scheduling method in the embodiments.
[0007] In response to the to-be-sent data stream conforming to a transport layer protocol and the traffic size of the to-be-sent data stream being greater than a preset threshold, a plurality of candidate port values for identifying the to-be-sent data stream at a data sending end are generated, each candidate port value corresponding to a candidate network path; a plurality of probe data packets are created, wherein the source port fields of the probe data packets are set to different candidate port values; the probe data packets are sent to a data receiving end, and the performance indicators of the candidate network paths corresponding to the plurality of probe data packets are obtained based on the response of the data receiving end; the plurality of candidate port values are evaluated based on the performance indicators, and a target source port value is determined from the plurality of candidate port values according to the evaluation result; the source port field of the data packet corresponding to the to-be-sent data stream is modified to the target source port value, and the modified data packet corresponding to the to-be-sent data stream is sent to the data receiving end.
[0008] The traffic scheduling method provided by the application comprises the following steps: in response to the to-be-sent data stream conforming to a transport layer protocol and the traffic size of the to-be-sent data stream being greater than a preset threshold, a plurality of candidate port values for identifying the to-be-sent data stream at a data sending end are generated, each candidate port value corresponding to a candidate network path; a plurality of probe data packets are created, wherein the source port fields of the probe data packets are set to different candidate port values; the probe data packets are sent to a data receiving end, and the performance indicators of the candidate network paths corresponding to the plurality of probe data packets are obtained based on the response of the data receiving end; the plurality of candidate port values are evaluated based on the performance indicators, and a target source port value is determined from the plurality of candidate port values according to the evaluation result; the source port field of the data packet corresponding to the to-be-sent data stream is modified to the target source port value, and the modified data packet corresponding to the to-be-sent data stream is sent to the data receiving end.
[0009] Thus, by setting the to-be-sent data stream to satisfy the double judgment conditions of the transmission layer protocol and the traffic size, it is ensured to detect the data stream of the large data stream which supports the modification of the source port field, to prevent waste of system resources, and compared with the passive acceptance of the default hash result in the traditional way, the application actively creates a detection data packet, sets the source port field of each detection data packet to a different candidate port value, simulates data sending for each candidate port value, logically establishes the mapping relationship between the source port value and the network path, measures the performance indicators of the simulated sending, and the system can accurately perceive the health status and congestion degree of each candidate network path, actively avoids the congested path with high delay and high packet loss, selects the path with the optimal quality, avoids the negative effects caused by the hash conflict from the root, realizes dynamic and adaptive load balancing, and realizes the traffic scheduling target by modifying only the source port field, while ensuring complete compatibility with related network devices, protocols and applications, without modifying any network infrastructure or application program, and the deployment cost is low. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0011] Figure 1 The flowchart of the traffic scheduling method provided by an embodiment of the present application; Figure 2 The structural diagram of the traffic scheduling system provided by an embodiment of the present application; Figure 3 The internal structure diagram of the electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0012] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.
[0013] It should be noted that in the description of the present application, the terms "comprising", "containing" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or network device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or network device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0014] In order to enable those skilled in the art to better understand the present application, the present application will be further described in detail below in combination with the drawings and specific embodiments.
[0015] In modern data center, cloud computing and high-speed enterprise network environment, in order to fully utilize the switching capability of network equipment, multi-path routing technology such as equal-cost multi-path routing (ECMP) or link aggregation group (LAG) is widely used. In these technologies, network switches usually disperse data streams to multiple physical links by performing hash operation on specific fields in the packet header of data packets (such as source IP address, destination IP address, protocol type, source port, destination port, etc.) to achieve load balancing and improve overall bandwidth.
[0016] However, this hash-based load balancing mechanism has an inherent defect: hash collision. That is, different data streams may be calculated as the same hash value, and thus mapped to the same physical link. In actual network traffic, due to the asymmetry of traffic (such as the existence of "elephant flow") or the uneven distribution of hash fields (for example, a large amount of traffic may be directed to a few destination ports), the occurrence of hash collision will cause serious imbalance of load on multiple logical links. Some links may be congested, resulting in queue delay or even packet loss, while other links are in light load or idle state. This undoubtedly causes waste of network bandwidth resources and significantly reduces the overall forwarding efficiency and performance of the network. Currently, the industry's solutions to this problem mainly focus on the network equipment side. For example, optimizing the hash algorithm of the switch (such as using CRC32 or other algorithms with more uniform distribution), using weighted cost multi-path routing (WCMP) or dynamically adjusting the path weight through a centralized controller. However, these solutions require the support of network core equipment, have poor deployment flexibility, and cannot adapt to all dynamically changing traffic patterns, making it difficult to fundamentally solve the load imbalance problem caused by hash collision.
[0017] In view of the above technical problems, as shown in Figure 1 The embodiment of the present application provides a traffic scheduling method, which specifically comprises the following steps: Step 101: in response to the to-be-sent data stream conforming to a transport layer protocol and a traffic size of the to-be-sent data stream being greater than a preset threshold, generating a plurality of candidate port values for identifying the to-be-sent data stream at a data sending end, each candidate port value corresponding to a candidate network path.
[0018] In actual service traffic, the source IP (source network address) and the destination IP (destination network address) are determined by both parties of communication and are fixed. The protocol type is determined by an application program and is also fixed. The destination port is determined by a server application program listening port (such as Web-80 / 443 and SSH-22 port) and is fixed. Therefore, the only variable left is the source port (Source Port).
[0019] The traffic scheduling method provided in the application is applied to a data sending end, and a specific implementation carrier thereof can be a network card driver program running in an operating system kernel or an intelligent network card or a data processing unit, so as to realize hardware offloading with zero occupation of host resources.
[0020] The execution subject of the traffic scheduling method is transferred from a traditional network core device (such as a switch) to the data sending end by setting. This fundamental perspective conversion breaks the dependence on complex modification of network infrastructure and realizes a distributed and easy-to-deploy optimization mechanism by implementing preprocessing before a data packet enters a network.
[0021] The to-be-sent data stream refers to a set of data packets with the same five-tuple (source IP, destination IP, protocol type, source port, and destination port) that need to be sent through a network.
[0022] The candidate port value is a set of mutually non-repeated source port values generated by the system in a temporary port range of an operating system kernel. Each candidate port value corresponds to a candidate network path potentially after hash calculation via a network device.
[0023] Specifically, it is determined whether the data packet corresponding to the to-be-sent data stream includes a modifiable source port field and whether the traffic size of the to-be-sent data stream is greater than a preset threshold. In response to the data packet corresponding to the to-be-sent data stream including the modifiable source port field and the traffic size of the to-be-sent data stream being greater than the preset threshold, a plurality of non-repeated port values are generated in a temporary port range of an operating system kernel. The port values are bound to to-be-sent data stream key information to generate a plurality of candidate port values for identifying the to-be-sent data stream at a data sending end. The to-be-sent data stream key information includes a source network address, a destination network address, a source port, a destination port, and a transport layer protocol number.
[0024] By judging whether the to-be-sent data stream includes a modifiable source port field, judging whether the to-be-sent data stream conforms to a transport layer protocol (TCP or UDP protocol), and estimating the total byte number of the data stream and comparing the total byte number with a preset threshold, when the total byte number exceeds the threshold, it is judged that the traffic size of the to-be-sent data stream is greater than the preset threshold, that is, the data stream is identified as an elephant flow. Only when the above two conditions are met at the same time, the subsequent traffic scheduling method is executed; for the regular data stream (such as ICMP packet or short flow) that does not meet the condition, the system will keep the original source port unchanged and send it according to the default rule, so as to ensure the efficiency and economy of the system.
[0025] The system will generate a plurality of candidate source port values in the temporary port range preset in the operating system kernel. In order to ensure that these candidate port values can be correctly applied to the target data stream and tracked, the system will bind each candidate port value with the five-tuple key information of the to-be-sent data stream. The five-tuple includes source IP address, destination IP address, source port, destination port and transport layer protocol. This binding operation generates a plurality of candidate port value sets for uniquely identifying the to-be-sent data stream at the data sending end. Each candidate port value, after entering the network, will logically correspond to a potential candidate network path through the hash calculation of the switch, laying a solid foundation for subsequent path detection and optimization. In this way, by abandoning the strategy of passively accepting the default hash result and turning to actively generating a plurality of candidate source port values, the network hash result can be perceived in advance through subsequent detection of the corresponding path state. This mechanism changes from blind sending to pre-detection and post-sending, avoiding major performance loss caused by hash conflict with minimal detection overhead, and embodies intelligent predictability.
[0026] In an embodiment, a plurality of candidate source ports can also be generated by a candidate port generator, which can be a functional module located at the data sending end and responsible for the task of generating a candidate source port set. Its generation strategy aims to ensure the diversity and dispersion of the port values, for example, using a random or pseudo-random algorithm to ensure that these ports can be mapped to as many different physical paths as possible after being hashed by the network device.
[0027] Step 102: creating a plurality of probe data packets, wherein the source port field of each probe data packet is set to a different candidate port value.
[0028] Step 103: sending the probe data packets to the data receiving end and obtaining the performance indicators of the candidate network paths corresponding to the plurality of probe data packets based on the responses of the data receiving end.
[0029] The probe data packet refers to a data packet specially constructed for performing path performance probing, the fields of destination IP, destination port, protocol type, etc. of which are consistent with the real data stream to be sent, but the source port field is set to a candidate port value, thereby simulating the network path when the data stream uses the port.
[0030] The source port field refers to a specific field in the protocol header of a network data packet (such as a TCP packet or a UDP packet), which is used to identify the sending application port. The present application actively modifies this field as a control variable for controlling the data stream to select different network paths.
[0031] The data receiving end refers to the opposite host communicating with the data sending end, which is responsible for receiving the probe data packet and returning a response, thereby forming a complete end-to-end communication to measure the path performance.
[0032] The core feature of the probe data packet lies in the simulation of its five-tuple: except for the source port field, the remaining fields (including the destination IP address, the destination port, and the protocol type) are completely consistent with the real data stream to be scheduled, so as to ensure the consistency of the probe path and the real data transmission path. Each candidate port value is sequentially filled into the source port field of the probe data packet. Thus, each probe packet loaded with a different candidate source port becomes a probe for a specific potential path. In this way, the key physical operation of mapping the port to the path is realized.
[0033] The probe data packets created above are sequentially or in parallel sent to the data receiving end. When the data receiving end receives these probe packets, it will return a response packet in accordance with the standard network protocol (such as the ACK mechanism of TCP or the customized response of UDP). In the data sending end, the system captures the time delay of this interaction through a high-precision mechanism.
[0034] In the embodiment executed by the intelligent network card or the data processing unit, this mechanism is the hardware timestamp, which can provide nanosecond-level precision, thereby accurately measuring the round-trip time (RTT).
[0035] In the software embodiment executed by the operating system driver, the feedback of the kernel protocol stack and the system timestamp are calculated. By recording the time from sending each probe packet to receiving the response, the system calculates one or more performance indicators for each candidate source port value (i.e. each candidate network path), and the most core one is the RTT. Further, the system can also calculate the jitter of the RTT through multiple probes to evaluate the stability of the path.
[0036] Step 104: evaluating the multiple candidate port values based on the performance indicators, and determining the target source port value from the multiple candidate port values according to the evaluation results.
[0037] Specifically, based on the performance metrics of candidate network paths, a state vector representing each candidate network path is constructed; the state vector is used as the input to the reinforcement learning model, and the target source port value is output through the reinforcement learning model.
[0038] The state vector of each candidate network path is represented as follows: ; Among them, S t (i) Represents the state vector of network path i at time t, RTT t (i) Jitter represents the smoothed round-trip time of network path i. t (i) LossRate represents the latency jitter of network path i. t (i) Q represents the packet loss rate of network path i. t (i) B represents the normalized queue depth corresponding to network path i. t (i) T represents the estimated available bandwidth of the normalized network path i. t This represents normalized time.
[0039] In one implementation, the target source port value is determined as the output of the reinforcement learning model based on the reward function of the model. The reward function is expressed as: ; Where, r t RTT represents the reward value obtained by the data sender after selecting candidate port value A at time t. t (i) The smoothed round-trip time of network path i is represented by β, which is the latency sensitivity parameter, and LossRate is the loss rate. t (i) Let λ represent the packet loss rate of network path i, λl represent the reliability sensitivity parameter, and B represent the packet loss rate of network path i. t (i) B represents the estimated available bandwidth of the normalized network path i. max λ represents the physical bandwidth limit of the network link, λt represents the throughput sensitivity index, and λs represents the handover penalty coefficient, used to balance the cost of path handover. t A represents the candidate port value selected at time t. t-1 This represents the candidate port value selected at time t-1, 1 {At≠At-1} As an indicator function, when A t With A t-1 If the values are different, its value is 1; otherwise, it is 0.
[0040] The evaluation mechanism of the present application is driven by a reinforcement learning model, the core of which is a multi-objective reward function, which aims to select a path that simultaneously achieves low latency, high reliability and high throughput, and maintains a certain stability to avoid the overhead caused by frequent switching. The reward function integrates these several competing objectives into a scalar reward value. Among them, the delay utility term ensures strict punishment for high-latency paths; the reliability term makes any packet loss significantly reduce the reward; the throughput term encourages the use of high-bandwidth paths, while the stability penalty term applies a fixed penalty to the current selection of a different path behavior from the last time to prevent jitter.
[0041] Among them, the reward function based on the reinforcement learning model determines the target source port value as the output of the reinforcement learning model, including: obtaining a plurality of candidate port values corresponding Gaussian distribution; the reward value of the candidate port value is subject to the Gaussian distribution corresponding to the candidate port value; in response to the need to determine the target source port value, sample values corresponding to each candidate port value are obtained from the Gaussian distribution corresponding to each candidate port value; the candidate port value with the largest sample value is taken as the target source port value.
[0042] The Gaussian distribution is represented as: ; Among them, Q t (A) represents the maximum sample value, N(μ A (t), σ A 2 (t)) represents the Gaussian distribution, μ A (t) represents the reward mean at time t, and σ A 2 (t) represents the reward variance at time t.
[0043] When making a specific decision, the system uses the Thompson Sampling algorithm, which maintains a Gaussian distribution of rewards, where the reward mean represents the estimation of the quality of the port based on historical experience, and the reward variance represents the uncertainty of this estimation. When a decision needs to be made, the system samples a temporary reward value (Q t (A)) from the Gaussian distribution corresponding to each candidate port value. Then, the system selects the candidate port with the largest sample value as the target source port value for this time. This sampling and optimization mechanism cleverly balances the selection of the candidate port value with the best current estimate and the candidate port value with high uncertainty, enabling the system to adaptively discover potentially better paths.
[0044] The application also sets the candidate port value with the maximum sampling value as the target source port value, and then includes obtaining the reward value corresponding to the target source port value based on the target source port value sending the data packet corresponding to the to-be-sent data stream; and updating the Gaussian distribution parameter corresponding to the target source port value based on the reward value corresponding to the target source port value.
[0045] Specifically, the Gaussian distribution parameter corresponding to the target source port value is updated based on a Bayesian function, and the Bayesian function is represented as: ; ; ; Wherein, n A (t+1) represents the total number of times that the candidate port value is selected at time t+1, n A (t) represents the total number of times that the candidate port value is selected at time t, μ A (t) represents the reward mean value at time t, μ A (t+1) represents the reward mean value at time t+1, σ A 2 (t+1) represents the reward variance at time t+1, σ0 2 represents the initial prior variance of the Gaussian distribution, σ r 2 represents the noise variance of the observed reward, which is used to measure the degree of unreliability of a single reward observation value.
[0046] When the real data stream is sent using the target source port, the performance of the path is continuously measured, and the reward function is substituted again to calculate the actual reward value corresponding to the target source port value of this decision. Subsequently, the system starts the Bayesian updating process, and the actual reward value is used to update the Gaussian distribution parameter corresponding to the target source port value. As shown in the Bayesian updating formula, the reward mean value moves in the direction of the actual reward value of the newly observed target source port value, and the moving step decreases with the increase of the selection times; at the same time, the reward variance decreases, indicating that the estimation confidence of the port reward value increases. Through this continuous “decision-observation-update” cycle, the reinforcement learning model can track the changes of the network state in real time, make the decision agent increasingly accurate, and finally realize the dynamic optimal balance of the global load.
[0047] Step 105: modifying the source port field of the data packet corresponding to the to-be-sent data stream to the target source port value, and sending the modified data packet corresponding to the to-be-sent data stream to the data receiving end.
[0048] The target source port value refers to the candidate source port value that is considered optimal under the current network state and is ultimately determined through the aforementioned intelligent assessment and decision-making process (such as the output of the reinforcement learning model). It is the final output of the entire optimization process.
[0049] The source port field modification refers to the technical operation of replacing the source port field value in the protocol header (such as the TCP header or UDP header) of the original data packet to be sent with the target source port value at the data sending end. This operation is completed before the data packet enters the network.
[0050] The modified data stream to be sent refers to the original data stream whose source port field of all data packets has been uniformly modified to the target source port value. This data stream will be continuously sent through the selected optimal path.
[0051] The system will intercept the original data stream that should be sent directly and modify the source port field of all data packets from the original port used by the application program to the target source port value obtained through optimization in bulk and consistently. Subsequently, the batch of messages with uniform source ports is delivered to the network interface and sent to the data receiving end. Since network switching devices use the source port as a key input for the hash algorithm when performing multi-path routing, the modified data stream will be consistently mapped to the previously detected physical path with optimal performance, ensuring that the entire data stream enjoys low latency and high throughput transmission experience.
[0052] The source port is a completely legal and expected variable field in the TCP / UDP protocol. Modifying it does not violate any network protocol specifications and will not be discarded by network devices as an exception. In this application, the source port, a standard protocol field, is creatively selected as the core for controlling data flow paths, achieving complete transparency to the upper-layer application. Specifically, from traffic identification, intelligent detection to the final port modification, all operations are completed in the operating system kernel or intelligent network card. For the application program that generates the data stream, it still performs normal network communication on its own socket and is completely unaware of the complex path optimization process behind it. No code modification is required for existing application programs, greatly reducing the technical threshold and cost of deployment.
[0053] In an embodiment, for scenarios where there are a large number of data sending ends that need to cooperate for traffic optimization in large-scale data centers or cloud computing environments, the application sets up a distributed federated learning mechanism, allowing multiple data sending ends to share learning experience and achieve group intelligence.
[0054] The data sending end is multiple in the application, and the flow scheduling method further includes: establishing communication between the multiple data sending ends; selecting any data sending end as a coordination node from the multiple data sending ends; the coordination node obtains the reinforcement learning model parameters of all data sending ends, performs weighted average on the reinforcement learning model parameters of all data sending ends, obtains coordination model parameters, and distributes the coordination model parameters to the multiple data sending ends; for any data sending end, the corresponding reinforcement learning model thereof is updated based on the coordination model parameters.
[0055] It needs to be explained that the coordination node is a logical central node dynamically elected or designated in the multiple data sending ends, and the function thereof is to collect the model parameters of each end, perform a federal average algorithm, and distribute the obtained coordination model parameters after processing, and the coordination node itself is also a data sending end.
[0056] The reinforcement model parameters herein refer to internal state data of a local reinforcement learning intelligent agent (such as a Thompson Sampling model) of the data sending end. Specifically, the reinforcement model parameters can include reward mean, reward variance, and selection times for each candidate port value.
[0057] The coordination model parameters include reward mean, reward variance, and selection times, the weighted average of the reinforcement learning model parameters of all data sending ends is performed to obtain the coordination model parameters, and the corresponding reinforcement learning model of any data sending end is updated based on the coordination model parameters, including: calculating the reward mean, reward variance, and selection times corresponding to each candidate port value based on a weighted average calculation formula of the reinforcement learning model parameters; each data sending end updates the corresponding reinforcement learning model based on the reward mean, reward variance, and selection times.
[0058] The weighted average calculation formula of the reinforcement learning model parameters is represented as: ; ; ; K represents the number of data sending ends participating in coordination, k represents the reinforcement learning model parameters of the kth data sending end, μ A represents the reward mean corresponding to the candidate port value A, σ A represents the reward variance corresponding to the candidate port value A, σ0 r represents the noise variance of the observed reward, N A represents the total selection times of the candidate port value A, n A represents the reset selection times generated according to the total selection times.
[0059] In the present application, a plurality of communication links between data sending ends are established to form a coordination group. Subsequently, a data sending end is elected as a temporary coordination node from the group through a lightweight distributed negotiation algorithm (such as the bully algorithm) or designated by a central controller.
[0060] When a predetermined coordination period arrives or a specific condition is triggered, the coordination node initiates a request to all other data sending ends in the group to obtain their respective locally maintained reinforcement learning model parameters. After receiving all the parameters, the coordination node does not collect any raw traffic data, but performs a weighted average of these model parameters through a federated averaging algorithm. This weighting ensures that data sending ends with more experience (more attempts) contribute more to the global model. After calculating the coordination model parameters, the coordination node distributes them to all data sending ends participating in the coordination. After receiving the coordination model parameters, each data sending end does not completely discard the local model, but uses it to update or initialize the local reinforcement learning model. For example, the parameters of the local model can be replaced with the coordination model parameters, or the two can be fused.
[0061] In this way, through the sharing of reinforcement learning parameters among multiple data sending ends, a newly joined sending end can quickly obtain historical experience, avoiding the exploration overhead of starting from zero and accelerating the convergence speed of the entire system. At the same time, it enables all sending ends to collaboratively discover and avoid global network congestion points, achieving a leap from local optimization to global optimization, thereby achieving true dynamic load balancing at the system level. In addition, since only model parameters are exchanged and not raw data, the mechanism also fully protects the local data privacy of each data sending end.
[0062] In a specific embodiment, to adapt to general servers, virtualization environments, and other scenarios that require rapid deployment and are cost-sensitive, the traffic scheduling method of the present application is implemented through a high-level network card driver in the server operating system kernel. The core operation process is as follows: The driver first determines whether the data packet to be sent is TCP or UDP protocol based on the protocol stack information, and further estimates the traffic size. When it is identified as large data traffic, the optimization process is triggered (i.e., the process of generating a plurality of candidate port values for identifying the data flow to be sent at the data sending end, and the subsequent process). Subsequently, the driver generates a set of diversified candidate source ports in the kernel state and schedules a series of probe packets to be sent. These probe packets differ only in source port, and the rest of the content is consistent with the actual data. Based on the ACK confirmation mechanism of the kernel protocol stack and the system timestamp, the driver accurately calculates the round-trip time (RTT) and its fluctuation degree (Jitter) of the path corresponding to each candidate port. Finally, the driver selects the candidate port with the lowest and most stable average RTT as the final source port, and modifies the source port field of the real data flow before sending it.
[0063] This embodiment realizes pure software, without any special hardware support, and can be deployed by upgrading the driver only, with extremely low deployment cost and extremely wide compatibility. In addition, the entire optimization process is completely transparent to the upper layer application, and is fully compatible with existing network devices.
[0064] In another specific embodiment, to meet the high-performance computing, financial transactions and other delay and host resource consumption extremely sensitive scenarios, the traffic scheduling method of the application is executed by the intelligent network card or data processing unit (DPU) hardware offloading. The core operation process is as follows: The host kernel offloads the large flow sending task and its context information to the intelligent network card. The programmable engine in the network card automatically generates a candidate source port set and sends a probe packet using hardware timestamp to independently measure the RTT of each path with nanosecond precision. The probe and decision process is completely completed on the network card, and the final source port of the real data stream is modified to the optimal selection by the network card hardware, and the checksum is updated and sent, without the intervention of the host CPU.
[0065] This hardware offloading embodiment can achieve extreme performance: first, it realizes zero CPU occupation of the host, completely isolates the computing overhead, and guarantees the performance of the host service. Second, hardware timestamp and line speed processing bring ultra-low and stable processing delay, meeting the performance requirements of the most demanding applications. In addition, thanks to the programmability of the intelligent network card, the destination port can be further modified or the protocol type can be added to modify the encapsulation header, realizing more flexible traffic control capability.
[0066] In yet another specific embodiment, for large cloud data centers, super-large clusters and other environments with extreme requirements for network performance, the application cooperates with the network switch through the intelligent network card supporting in-band network telemetry (INT) to realize global optimization. The core operation process is as follows: the probe packet generated by the sending end intelligent network card is embedded with INT metadata instructions. When this probe packet traverses the network, each INT supporting switch on the path will automatically write its local queue depth, timestamp, congestion mark, ingress / egress port and other detailed information into the packet. The receiving end encapsulates the INT metadata containing the complete path information in the response packet and returns it. After the sending end network card parses the response, it not only knows the end-to-end delay, but also clearly understands the real-time health status of each hop on the entire path, so as to make more fine and forward-looking decisions than simply based on RTT, such as actively avoiding a specific switch whose internal queue has been accumulated.
[0067] This INT telemetry embodiment can upgrade the decision basis from end-to-end delay to path-wide health. This enables the system to make proactive optimization, avoiding congestion before it actually affects end-to-end performance, thus achieving true global load optimization in a complex data center network environment, and significantly improving the efficiency of large-scale data-intensive applications.
[0068] The beneficial effects of the present application mainly manifest in the following aspects: 1. Intelligent triggering and resource optimization: By setting up an intelligent triggering mechanism that only starts optimization for elephant flows at the transport layer protocol, it ensures that the optimization benefit is much greater than the cost, thereby significantly improving system efficiency and optimizing the use of computing and network resources, and fundamentally reducing the overall total cost of ownership.
[0069] 2. Active probing and dynamic optimization: By actively sending probe packets to measure path performance, the system can real-time perceive network state changes and accurately predict the forwarding quality of different source ports. Subsequently, based on the probe results, the current optimal source port is selected, thereby effectively avoiding the inherent conflict problem of traditional hash mechanism, greatly improving network bandwidth utilization, and significantly reducing transmission delay and packet loss rate.
[0070] 3. Batch application and seamless compatibility: After determining the optimal source port, it is applied to all packets of the entire data flow in batches, ensuring the stability and consistency of flow transmission. Most importantly, the entire scheme is implemented by modifying only the standard protocol field of the source port, ensuring seamless compatibility with existing network devices, protocols, and upper-layer applications, achieving zero-reform deployment, and greatly reducing deployment costs and landing risks.
[0071] 4. Global optimization and business guarantee: The above key points work together to enable the scheme to intelligently and adaptively adjust traffic distribution and achieve global load balancing. This not only relieves the pressure on network core devices, but also effectively avoids interference with delay-sensitive key services by intelligently scheduling large flows, thereby guaranteeing the service quality of the overall business.
[0072] In a feasible implementation, in order to enable the traffic scheduling to intelligently adapt to the performance requirements of different applications, the business type identifier of the data flow to be sent can be obtained, and according to the business type identifier, the strategy parameters relied on in the evaluation process are dynamically adjusted, wherein the business type at least includes delay-sensitive business and throughput-sensitive business. Specifically, the strategy parameters relied on in the evaluation process are dynamically adjusted, including: when the business type is delay-sensitive, the weights of the delay sensitivity parameter β and the reliability sensitivity parameter λl in the reward function are adjusted to be higher; when the business type is throughput-sensitive, the weight of the throughput sensitivity index λt in the reward function is adjusted to be higher.
[0073] Specifically, in the reinforcement learning-based decision model, the system dynamically configures the sensitivity parameters of the reward function. For example, when identifying that the service type is delay-sensitive (such as video conferencing, online gaming), the system increases the weights of the delay sensitivity parameter β and the reliability sensitivity parameter λl. This makes the reward function more sensitive to the delay growth and packet loss behavior of the path, thereby driving the system to choose the path with the lowest delay and the most stable path at all costs to ensure the user experience of interactive services. When identifying that the service type is throughput-sensitive (such as large data backup, scientific computing), the system increases the weight of the throughput sensitivity index λt. This will guide the system to preferentially select the path with the highest available bandwidth to maximize data transmission efficiency and shorten job completion time. In this way, by customizing the optimization goal for different service types, the mismatch of resource allocation is effectively avoided. For example, it prevents scheduling massive backup traffic to a path with low delay but limited bandwidth, or scheduling real-time voice traffic to a path with high bandwidth but congestion jitter, thereby achieving the optimal overall service performance at the system level and adapting to complex and heterogeneous network environments to meet the differentiated needs of different services from real-time interaction to bulk transmission.
[0074] As shown in Figure 2 The embodiments of the present application also provide a traffic scheduling system deployed at a data sending end, which specifically includes a traffic monitor, a candidate port generator, a probe scheduler, a packet modifier, and a core decision maker.
[0075] The traffic monitor, as the triggering starting point of the system, continuously monitors the data flow to be sent at the data sending end, and in response to identifying a data flow that meets the transmission layer protocol and has a traffic size exceeding a preset threshold, triggers the subsequent optimization process.
[0076] The candidate port generator, in response to the triggering of the traffic monitor, generates a plurality of candidate source port values for identifying the data flow within the temporary port range of the operating system kernel.
[0077] The probe scheduler, in response to the generation of the candidate port set, is responsible for creating probe data packets and setting their source port fields to different candidate port values. Subsequently, it schedules the sending of these probe packets and synchronously notifies the core decision maker to start performance monitoring.
[0078] The core decision maker, as the intelligent hub of the system, in response to the notification of the probe scheduler, starts receiving and processing probe responses from the data receiving end. It calculates the performance indicators of each candidate path based on these responses and executes an intelligent evaluation algorithm (such as a reinforcement learning model), thereby deciding the optimal target source port value.
[0079] The packet modifier performs two key operations in response to the target source port value received from the core decision maker: one is to send the probe packets generated by the probe scheduler to the network in the probe stage; the other is to modify the source port field of all data packets of the original data stream to be sent to the target source port value after the decision is completed, and perform sending.
[0080] The traffic scheduling system deployed at the data sending end provided by the application constructs an efficient and automated closed-loop control process, seamlessly connects intelligent triggering, active probing, dynamic decision making and lossless execution, and finally realizes intelligent and adaptive scheduling of network traffic at the data sending end.
[0081] The embodiments of the application also provide an electronic device, as shown in the accompanying drawings, comprising a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the above traffic scheduling method embodiments. Figure 3
[0082] The embodiments of the application also provide a computer readable storage medium, which stores a computer program, wherein the computer program is configured to perform the steps in any of the above traffic scheduling method embodiments when running.
[0083] In an exemplary embodiment, the above computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile network device, a magnetic disk or an optical disk, and various media that can store computer programs.
[0084] The skilled person can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the application.
[0085] The above describes in detail the flow scheduling method provided by the present application. The principles and implementation manners of the present application are described by using specific examples, and the above description of the examples is only used to help understand the method of the present application and the core idea thereof. It should be pointed out that, for those skilled in the art, some improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the protection scope of the present application.
Claims
1. A traffic scheduling method, characterized in that, Applied to the data sending end, the traffic scheduling method includes: In response to the data stream to be sent conforming to the transport layer protocol and the traffic size of the data stream to be sent being greater than a preset threshold, multiple candidate port values are generated to identify the data stream to be sent at the data sending end, and each candidate port value corresponds to a candidate network path. Create multiple probe packets, where the source port field of each probe packet is set to a different candidate port value; The probe data packets are sent to the data receiving end, and the performance indicators of the candidate network paths corresponding to the multiple probe data packets are obtained based on the response of the data receiving end. The performance metrics are used to evaluate multiple candidate port values, and the target source port value is determined from the multiple candidate port values based on the evaluation results. Modify the source port field of the data packet corresponding to the data stream to be sent to the target source port value, and send the modified data packet corresponding to the data stream to be sent to the data receiving end.
2. The traffic scheduling method according to claim 1, characterized in that, The step of generating multiple candidate port values for identifying the data stream to be sent at the data sending end in response to the data stream conforming to the transport layer protocol and the data stream's size being greater than a preset threshold includes: Determine whether the data packet corresponding to the data stream to be sent includes a modifiable source port field, and determine whether the traffic size of the data stream to be sent exceeds a preset threshold; In response to the data packet corresponding to the data stream to be sent including a modifiable source port field and the traffic size of the data stream to be sent being greater than a preset threshold, multiple non-repeating port values are generated within the temporary port range of the operating system kernel. The port value is bound to key information of the data stream to be sent to generate multiple candidate port values for identifying the data stream to be sent at the data sending end. The key information of the data stream to be sent includes the source network address, destination network address, source port, destination port, and transport layer protocol number.
3. The traffic scheduling method according to claim 1, characterized in that, The process of evaluating multiple candidate port values based on the performance metrics and determining the target source port value from the multiple candidate port values based on the evaluation results includes: Based on the performance metrics of candidate network paths, a state vector representing each candidate network path is constructed. The state vector is used as the input to the reinforcement learning model, and the target source port value is output through the reinforcement learning model. The state vector is represented as: ; Among them, S t (i) Represents the state vector of network path i at time t, RTT t (i) Jitter represents the smoothed round-trip time of network path i. t (i) LossRate represents the latency jitter of network path i. t (i) Q represents the packet loss rate of network path i. t (i) B represents the normalized queue depth corresponding to network path i. t (i) T represents the estimated available bandwidth of the normalized network path i. t This represents normalized time.
4. The traffic scheduling method according to claim 3, characterized in that, The output of the target source port value through the reinforcement learning model includes: The target source port value determined by the reward function of the reinforcement learning model is used as the output of the reinforcement learning model; The reward function is expressed as follows: ; Where, r t RTT represents the reward value obtained by the data sender after selecting candidate port value A at time t. t (i) The smoothed round-trip time of network path i is represented by β, which represents the latency sensitivity parameter, and LossRate is the loss rate. t (i) Let λ represent the packet loss rate of network path i, λl represent the reliability sensitivity parameter, and B represent the packet loss rate of network path i. t (i) B represents the estimated available bandwidth of the normalized network path i. max λ represents the physical bandwidth limit of the network link, λt represents the throughput sensitivity index, and λs represents the handover penalty coefficient, used to balance the cost of path handover. t A represents the candidate port value selected at time t. t-1 This represents the candidate port value selected at time t-1, 1 {At≠At-1} As an indicator function, when A t With A t-1 If the values are different, its value is 1; otherwise, it is 0.
5. The traffic scheduling method according to claim 4, characterized in that, The target source port value determined by the reward function based on the reinforcement learning model, as the output of the reinforcement learning model, includes: Obtain a Gaussian distribution corresponding to multiple candidate port values, wherein the reward value of each candidate port value follows a Gaussian distribution corresponding to the candidate port value. In response to the need to determine the target source port value, samples are taken from the Gaussian distribution corresponding to each candidate port value to obtain the sampled value corresponding to each candidate port value; The candidate port value with the largest sampled value is used as the target source port value; The Gaussian distribution is represented as: ; Among them, Q t (A) represents the maximum sampled value, N(μ) A (t), σ A 2 (t) represents a Gaussian distribution, μ A (t) represents the average reward at time t, σ A 2 (t) represents the reward variance at time t.
6. The traffic scheduling method according to claim 5, characterized in that, The step of selecting the candidate port value with the largest sampled value as the target source port value includes: After sending the data packet corresponding to the data stream to be sent based on the target source port value, obtain the reward value corresponding to the target source port value; Based on the reward value corresponding to the target source port value, update the Gaussian distribution parameters corresponding to the target source port value.
7. The traffic scheduling method according to claim 6, characterized in that, The Gaussian distribution parameters corresponding to the updated target source port value include: Update the Gaussian distribution parameters corresponding to the target source port value based on the Bayesian function; The Bayesian function is expressed as: ; ; ; Where, n A (t+1) represents the total number of times the candidate port value was selected at time t+1, where n is the number of times. A (t) represents the total number of times a candidate port value is selected at time t, μ A (t) represents the mean reward at time t, μ A (t+1) represents the average reward at time t+1, σ A 2 (t+1) represents the reward variance at time t+1, σ0 2 Let σ represent the initial prior variance of the Gaussian distribution. r 2 This represents the noise variance of the observation reward, used to measure the unreliability of a single reward observation.
8. The traffic scheduling method according to claim 1, characterized in that, The data sending ends are multiple, and the traffic scheduling method further includes: Establish communication between multiple data senders; Select any one of the multiple data senders as the cooperating node; The collaborative node obtains the reinforcement learning model parameters of all data senders, performs a weighted average of the reinforcement learning model parameters of all data senders to obtain the collaborative model parameters, and distributes the collaborative model parameters to multiple data senders. For any data sender, update its corresponding reinforcement learning model based on the collaborative model parameters.
9. The traffic scheduling method according to claim 8, characterized in that, The collaborative model parameters include the mean reward, the variance of the reward, and the number of times the model is selected. The collaborative model parameters are obtained by weighted averaging of the reinforcement learning model parameters of all data senders. For any data sender, updating its corresponding reinforcement learning model based on the collaborative model parameters includes: The weighted average calculation formula based on the parameters of the reinforcement learning model is used to calculate the mean reward, variance of the reward, and number of times each candidate port value is selected. Each data sender updates its corresponding reinforcement learning model based on the mean reward, the variance of the reward, and the number of times it is selected; The weighted average calculation formula for the parameters of the reinforcement learning model is expressed as follows: ; ; ; Where K represents the number of data senders participating in the collaboration, k represents the reinforcement learning model parameters of the k-th data sender, and μ A ' represents the mean reward corresponding to candidate port value A, σ A ²' represents the reward variance corresponding to candidate port value A, σ0² represents the initial prior variance, and σ r ² represents the noise variance of the observed reward, N A n represents the total number of times candidate port value A was selected. A This indicates the reset selection count generated based on the total number of selections.
10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the traffic scheduling method as described in any one of claims 1 to 9 when executing the computer program.
Citation Information
Patent Citations
TCP (Transmission Control Protocol) data enhancement method and system for encrypted traffic classification
CN120128546A
Service network message cross-domain transmission method and system based on SRv6 SID function extension
CN120200953A
Sublimation computing power-driven dynamic link joint optimization method and system, terminal and storage medium
CN120321686A
Network flow control method, system, engine, node and related equipment
CN120567772A
Service transmission method and device of optical transmission network, electronic equipment and storage medium
CN120583340A
Cited By
Flow processing system and method, data processing method, equipment, medium and product
CN121357103A
Traffic processing system and method, data processing method, device, medium and product
CN121357103B