Data flow control methods, devices, electronic equipment, media, and software products

By predicting the rate of change of storage node buffers and dynamically selecting the target sender to pause data flow, the problem of inaccurate network congestion handling in distributed storage systems is solved, achieving precise flow control, reducing tail latency, and improving system performance and reliability.

CN120639704BActive Publication Date: 2026-01-30JINAN INSPUR DATA TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511130813.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-13
Publication Date
2026-01-30
Estimated Expiration
2045-08-13

AI Technical Summary

Technical Problem

In existing technologies, the network congestion handling of distributed storage systems is inaccurate, resulting in high tail latency and high peak buffer ratio in Incast scenarios, making it impossible to achieve the coordinated requirements of lossless transmission and low latency.

Method used

By predicting the rate of change of the storage node buffer, calculating the remaining time and pause time, and dynamically selecting the target sending end to pause data traffic, combined with CRUSH mapping and traffic proportion analysis, precise traffic control is achieved.

Benefits of technology

It reduces tail latency, improves data transmission efficiency and network resource utilization, ensures traffic fairness and network transmission accuracy, and enhances the overall performance and reliability of the distributed storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639704B_ABST
    Figure CN120639704B_ABST
Patent Text Reader

Abstract

This application discloses a data traffic control method, apparatus, electronic device, medium, and program product, relating to the fields of computer networks and distributed storage technology. The method includes: predicting the remaining time and pause time corresponding to the buffers of a first storage node based on the rate of change of the buffers of that first storage node; when the remaining time is less than a time threshold, selecting a target sender from multiple senders corresponding to the source distribution of data traffic to the first storage node, and instructing the target sender to pause sending data traffic to the first storage node based on the pause time. This solves the problem of inaccurate handling of network congestion in distributed storage systems in related technologies, achieving the effect of predicting and accurately handling network congestion in distributed storage systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer networks and distributed storage technology, and in particular to a data traffic control method, apparatus, electronic device, medium, and program product. Background Technology

[0002] In distributed storage systems, Remote Direct Memory Access (RDMA) technology achieves low-latency, high-throughput data transmission through RDMA over Converged Ethernet (RoCE) networks. RoCE networks achieve lossless transmission through Layer 2 Priority Flow Control (PFC) and Layer 3 Data Center Quantized Congestion Notification (DCQCN), but existing mechanisms have significant limitations.

[0003] PFC, as a Layer 2 congestion control mechanism, relies on a fixed threshold to passively trigger pauses, making it unable to predict congestion in advance. This results in tail latency as high as 820μs in 16:1 many-to-one injection (Incast) scenarios, and also suffers from head blocking and traffic unfairness issues. DCQCN relies on Explicit Congestion Notification (ECN) marking and explicit rate control, but feedback delays lead to excessively long control loops, making it unable to suppress sudden congestion in a timely manner. Insufficient coordination between these two mechanisms results in peak buffer usage reaching 100% in distributed storage networks during Incast scenarios, with tail latency fluctuations exceeding ±50%, urgently requiring a more precise congestion control mechanism.

[0004] The problem of inaccurate handling of network congestion in distributed storage systems in related technologies has not yet been effectively solved. Summary of the Invention

[0005] This application provides a data traffic control method, apparatus, electronic device, medium, and program product to at least solve the problem of inaccurate handling of network congestion in distributed storage systems in related technologies.

[0006] This application provides a data traffic control method, comprising: predicting the remaining time and pause time corresponding to the buffers of the first storage node based on the buffer change rate of the first storage node, wherein the remaining time is used to indicate the time interval from the current time to the time when the buffer is filled, and the pause time is used to indicate the time during which the first storage node needs to pause receiving data traffic when the remaining time is less than a time threshold; when the remaining time is less than the time threshold, selecting a target sender from multiple senders corresponding to the source distribution of the data traffic of the first storage node, and instructing the target sender to pause sending data traffic to the first storage node based on the pause time.

[0007] This application also provides a data traffic control device, comprising: a prediction module, configured to predict the remaining time and pause time corresponding to the buffers of the first storage node based on the buffer change rate of the first storage node, wherein the remaining time is used to indicate the time interval from the current time to the time when the buffer is filled, and the pause time is used to indicate the time during which the first storage node needs to pause receiving data traffic when the remaining time is less than a time threshold; and an indication module, configured to select a target sender from a plurality of senders corresponding to the source distribution of the data traffic of the first storage node when the remaining time is less than the time threshold, and instruct the target sender to pause sending data traffic to the first storage node based on the pause time.

[0008] This application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-described data flow control methods when executing the computer program.

[0009] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described data flow control methods.

[0010] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above-described data flow control methods.

[0011] This application predicts the remaining time and pause time corresponding to the buffer of a first storage node based on the buffer change rate of the first storage node. The remaining time indicates the time interval from the current moment to the moment the buffer is filled, and the pause time indicates the time during which the first storage node needs to pause receiving data traffic when the remaining time is less than a time threshold. When the remaining time is less than the time threshold, a target sender is selected from multiple senders corresponding to the source distribution of the data traffic to the first storage node, and the target sender is instructed to pause sending data traffic to the first storage node based on the pause time. Therefore, this solves the problem of inaccurate handling of network congestion in distributed storage systems in related technologies, achieving the technical effect of predicting and accurately handling network congestion in distributed storage systems. Attached Figure Description

[0012] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a hardware structure block diagram of a computer terminal for a data traffic control method according to an embodiment of this application;

[0014] Figure 2 This is a flowchart of a data traffic control method according to an embodiment of this application;

[0015] Figure 3 This is an architecture diagram of a distributed storage system according to an embodiment of this application;

[0016] Figure 4 This is a schematic diagram of module interaction according to an embodiment of this application;

[0017] Figure 5 This is a structural block diagram of a data flow control device according to an embodiment of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0019] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0020] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] The specific application environment architecture or specific hardware architecture on which the execution of the data traffic control method depends is described here.

[0022] The methods and embodiments provided in this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure block diagram of a computer terminal for a data traffic control method according to an embodiment of this application. Figure 1 As shown, a computer terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the computer terminal described above. For example, the computer terminal may also include components that are more complex than those described above. Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0023] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the data flow control method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the aforementioned method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to a computer terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0024] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the computer terminal. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0025] Figure 2 This is a flowchart of a data traffic control method according to an embodiment of this application, applied to storage nodes in a distributed storage system. Figure 2 As shown, the process includes the following steps:

[0026] Step S202: Predict the remaining time and pause time corresponding to the buffer of the first storage node by the buffer change rate of the first storage node, wherein the remaining time is used to indicate the time interval between the current time and the time when the buffer is filled, and the pause time is used to indicate the time when the first storage node needs to pause receiving data traffic when the remaining time is less than a time threshold.

[0027] Step S204: If the remaining time is less than the time threshold, select a target sender from multiple senders corresponding to the source distribution of the data traffic of the first storage node, and instruct the target sender to pause sending data traffic to the first storage node based on the pause time.

[0028] The first storage node is an object storage device (OSD).

[0029] For example, if the remaining time is less than the time threshold of 5ms, it is determined that the buffer of the first storage node has a potential congestion risk and intervention is required. The intervention method is to instruct the target sending end to suspend sending data traffic to the first storage node based on the pause time.

[0030] Through the above steps, the remaining time and pause time corresponding to the buffer of the first storage node are predicted based on the buffer change rate of the first storage node. The remaining time indicates the time interval from the current moment to the moment the buffer is filled, and the pause time indicates the time during which the first storage node needs to pause receiving data traffic if the remaining time is less than a time threshold. If the remaining time is less than the time threshold, a target sender is selected from multiple senders corresponding to the source distribution of the data traffic to the first storage node, and the target sender is instructed to pause sending data traffic to the first storage node based on the pause time. Therefore, this solves the problem of inaccurate handling of network congestion in distributed storage systems in related technologies, achieving the technical effect of predicting and accurately handling network congestion in distributed storage systems.

[0031] The embodiments of this application provide a data traffic control method, and the method is described in detail below in conjunction with the execution flow of the data traffic control method.

[0032] In an exemplary embodiment, before predicting the remaining time and pause time corresponding to the buffer of the first storage node by the buffer change rate of the first storage node, the method further includes: continuously sampling the occupied size of the buffer according to the sampling time interval to obtain multiple sampling data pairs, wherein each sampling data pair includes: the occupied size of the buffer at two times before and after the sampling time interval; and determining the buffer change rate by the multiple sampling data pairs.

[0033] Furthermore, determining the buffer change rate through the multiple sampled data pairs includes: calculating the target buffer change rate corresponding to each sampled data pair using each sampled data pair and the sampling time interval; and inputting multiple consecutive target buffer change rates into a sliding window to obtain the buffer change rate.

[0034] Target buffer change rate The calculation formula is:

[0035] .

[0036] in, The sampling time interval can be 10ms; Let t be the size of the buffer at time t. This represents the size of the buffer after the sampling time interval at time t. It is generally used for continuous sampling, for example, sampling sequentially at t=10ms, t=20ms, and t=30ms.

[0037] Therefore, the rate of change of the buffer zone The calculation formula is:

[0038] .

[0039] in, This indicates the number of sampled data pairs (i.e., the number of rate of change of the input target buffer), for example, This represents a sliding window with three sampling time points.

[0040] Through the above embodiments, by periodically sampling the buffer occupancy size and calculating the rate of change, the growth trend of buffer usage over time can be captured in a timely manner. Using a sliding window algorithm to smooth the sampled data can filter out instantaneous fluctuations, resulting in a more stable and reliable rate of change for the buffer. This helps to identify potential network congestion in advance, allowing preventative measures to be taken before congestion occurs.

[0041] Furthermore, after determining the buffer change rate through the multiple sampled data pairs, the method further includes: determining whether the buffer is in an overflow state after a first target time through the buffer change rate; if the buffer is in an overflow state after the first target time, instructing the multiple sending ends to reduce the speed of sending data traffic at a notification time, wherein the notification time is separated from the current time by a second target time, and the second target time is less than 1 / 2 of the first target time.

[0042] For example, if the calculated buffer change rate is 0.67MB, and this rate of change is maintained, the buffer will increase by 0.67 × 100 = 67MB after the first target time of 100ms. Adding this to the current 45MB will cause an overflow (assuming a total capacity of 100MB). In this case, an action can be triggered in advance (e.g., notifying the sender to slow down at t=40ms) to avoid actual overflow.

[0043] In one exemplary embodiment, predicting the remaining time and pause time corresponding to the buffers of the first storage node based on the buffer change rate of the first storage node includes: predicting the remaining time based on the buffer change rate and the round-trip delay corresponding to the first storage node; and predicting the pause time based on the buffer change rate and a safety factor.

[0044] Furthermore, predicting the remaining time using the buffer change rate and the round-trip delay corresponding to the first storage node includes: determining the difference between the total size of the buffer and the size occupied by the buffer at the current moment, and determining the ratio of the difference to the buffer change rate; and predicting the remaining time using the ratio and the round-trip delay.

[0045] Specifically, the remaining time The calculation formula is:

[0046] .

[0047] in, Indicates the total size of the buffer. This indicates the current occupied size. This indicates the round-trip delay, where the round-trip delay (RTT) is abbreviated as RTT.

[0048] By monitoring the rate of change of the buffer in real time and combining it with the round-trip latency of the storage nodes, the system can calculate in advance the remaining time before the buffer reaches full capacity, thus predicting when congestion will occur.

[0049] Furthermore, predicting the pause time using the buffer change rate and the safety factor includes: determining the difference between the total size of the buffer and the current occupied size of the buffer, and determining the ratio of the difference to the buffer change rate; and predicting the pause time using the ratio and the safety factor.

[0050] Furthermore, predicting the pause time using the ratio and the safety factor includes: determining a historical adjustment factor using the k previous historical adjustment data at the current moment, wherein the i-th historical adjustment data in the k previous historical adjustment data includes: the pause time predicted in the i-th instance and the actual time of the i-th unblocking of the first storage node, where k is a positive integer and i = 1, 2, 3...k; calculating a target pause time using the ratio and the safety factor; and adjusting the target pause time using the historical adjustment factor to obtain the pause time.

[0051] Specifically, the target pause time The calculation formula is:

[0052] .

[0053] in, For safety, the default value is 0.2.

[0054] Optional, taking k=5 as an example, historical adjustment factor The calculation method is as follows:

[0055] .

[0056] in, The pause time obtained from the i-th prediction, The actual time for the i-th time that the first storage node is decongested.

[0057] The target pause time is thus adjusted using the historical adjustment factor to obtain the pause time, including: ;in, Indicates the pause time.

[0058] Based on the buffer change rate and safety factor, the system can dynamically predict the optimal pause time, ensuring that data inflow is reduced in a timely manner before the buffer reaches a critical value, thus avoiding buffer overflow. This predictive mechanism allows the system to adjust network traffic more accurately and promptly, avoiding traditional global passive response methods, reducing the impact on innocent traffic, and improving the accuracy of flow control and the fairness of network transmission. It reduces the possibility of buffer overflow, lowers tail latency, improves data transmission efficiency and network resource utilization, and enhances the overall performance and reliability of the distributed storage system.

[0059] In an exemplary embodiment, selecting a target sender from a plurality of senders corresponding to the source distribution of data traffic of the first storage node includes: determining the total source amount of data traffic of the first storage node, and determining each source amount corresponding to each sender; determining the traffic proportion of each sender among the plurality of senders by the ratio of each source amount to the total source amount, wherein the traffic proportion is used to indicate the source distribution; selecting a target traffic proportion from the plurality of traffic proportions sorted according to a target order, and taking the sender corresponding to the target traffic proportion as the target sender.

[0060] Understandable is the proportion of traffic. The calculation formula is: ;in, This represents the quantity of each source corresponding to each sender. This represents the total amount of data traffic originating from the first storage node. m represents the total number of source OSDs sending traffic to the target OSD. Multiple senders include m source OSDs.

[0061] Sort the m source OSDs in descending order of traffic percentage according to the target order; select source OSDs from high to low until the cumulative selected traffic percentage exceeds 75%, and then use the selected source OSDs as the target senders.

[0062] Specifically, by performing a detailed analysis of the data traffic received by the first storage node (OSD node), the total contribution of all senders to the data traffic of the target node can be determined, and the traffic source of each sender can be identified. By calculating the ratio of the traffic source of each sender to the total source, i.e., the traffic proportion, the distribution of data traffic can be understood, which helps to identify the data streams that have the greatest impact on network congestion.

[0063] Furthermore, based on the proportion of traffic, the sending ends are sorted according to a target order, and the sending ends with the highest traffic proportions are selected as "target sending ends," and traffic control measures, such as suspension or rate limiting, are applied to them, while other sending ends are exempt from control. This method avoids the traditional one-size-fits-all suspension of network traffic, reduces interference with normal data transmission, and ensures the rational allocation and efficient utilization of network resources.

[0064] In an exemplary embodiment, instructing the target sender to pause sending data traffic to the first storage node based on the pause time includes: a sending step: sending a preset data frame to the target sender, wherein the preset data frame is used to instruct the target sender to pause sending data traffic to the first storage node; an updating step: calculating a fairness adjustment factor corresponding to each sender in the target sender, updating the traffic share of each sender among the plurality of senders through the fairness adjustment factor, and updating the target sender through the updated traffic share; cyclically executing the above sending step and updating step until a preset condition is met, determining that the instruction to the target sender to pause sending data traffic to the first storage node based on the pause time has been completed, wherein the preset condition includes one of the following: the pause time ends, or the remaining time is detected to be greater than or equal to the time threshold.

[0065] Furthermore, calculating the fairness adjustment factor for each transmitter in the target transmitter includes: determining the total time each transmitter was suspended in the most recent statistical period; and determining the fairness adjustment factor by the ratio of the total time to the period time corresponding to the statistical period.

[0066] The preset data frame is the PFC pause frame, and the fairness adjustment factor is... The calculation formula is:

[0067] .

[0068] in, The total time each transmitter (e.g., the i-th source OSD) was paused within the most recent statistical period (e.g., the past minute); then the statistical period... .

[0069] Traffic share The adjustment formula is: .

[0070] The above embodiments achieve dynamic, fair, and efficient network flow control, especially in distributed storage systems facing network congestion. By sending a preset data frame to the target sender to indicate a pause, the system can respond instantly to buffer congestion, avoiding packet drop and increased transmission latency caused by buffer overflow. Simultaneously, a fairness adjustment factor is introduced to dynamically adjust the traffic share of each sender, ensuring that, in the long run, each sender will not suffer unfair data transmission restrictions due to frequent pauses, thus avoiding traffic starvation and enhancing the fairness of network flow control. The mechanism of cyclically executing the sending and updating steps until preset conditions are met ensures the real-time performance and effectiveness of network congestion control. Furthermore, by gradually relaxing traffic restrictions, the system allows for rapid resumption of data transmission after the buffer returns to normal, improving the response speed and overall performance of the distributed storage network.

[0071] In an exemplary embodiment, after instructing the target sending end to suspend sending data traffic to the first storage node based on the pause time, the method further includes: when there is data traffic to be stored in a client communicating with a distributed storage system including the first storage node, calculating the buffer occupancy, network latency, and load factor of the first storage node respectively; performing a weighted summation of the buffer occupancy, network latency, and load factor to determine the network congestion factor of the first storage node, and instructing the client to select a second storage node in the distributed storage system corresponding to the data traffic to be stored based on the network congestion factor.

[0072] Among them, network congestion factor The calculation formula is as follows:

[0073] .

[0074] in, Indicates buffer occupancy. Indicates network latency. Indicates the load factor. Optional. , , . The round-trip delay is based on the baseline.

[0075] For example, a storage node with a network congestion factor of less than 0.6 can be selected as the second storage node.

[0076] Furthermore, after determining the network congestion factor of the first storage node by weighted summation of the buffer occupancy, the network latency, and the load factor, the method further includes at least one of the following: instructing a monitoring node in the distributed storage system including the first storage node to perform fragment migration of the first storage node when the network congestion factor is higher than a preset threshold; instructing the monitoring node to calculate the congestion status of multiple storage nodes in the distributed storage system using a congestion prediction model and multiple network congestion factors, and instructing the monitoring node to perform fragment migration of the first storage node according to the congestion status.

[0077] The congestion prediction model is as follows:

[0078] .

[0079] in, This is an indicator function used to determine the congestion factor of a single OSD. Whether it exceeds 0.8 (which can be understood as "the OSD is in a high-risk congestion state"), return 1 if the condition is met, otherwise return 0; The total number of OSD nodes in the cluster, when the congestion probability... In such cases, rebalancing is triggered in advance, and the rebalancing cycle is dynamically adjusted: 5 minutes in normal state, 30 seconds in high congestion warning state, and 10 seconds in congestion state.

[0080] In an exemplary embodiment, before predicting the remaining time and pause time corresponding to the buffers of the first storage node based on the buffer change rate of the first storage node, the method further includes: determining the load state of the first storage node, wherein the load state includes: CPU utilization; adjusting the total size of the buffers based on the load state; and determining the response strategy of the first storage node to data requests from the plurality of sending ends based on the load state.

[0081] In one exemplary embodiment, adjusting the total size of the buffer based on load status includes: determining the total size of the buffer as a first value when the load status is in a first adjustment range; determining the total size of the buffer as a second value when the load status is in a second adjustment range; and determining the total size of the buffer as a third value when the load status is in a third adjustment range; wherein the first value is greater than the second value, the second value is greater than the third value, the maximum value in the first adjustment range is less than the minimum value in the second adjustment range, and the maximum value in the second adjustment range is less than the minimum value in the third adjustment range.

[0082] Among them, the Central Processing Unit (CPU) is the most important component.

[0083] In other words, the first storage node dynamically adjusts the total size of the receive buffer according to its own load status: when the load status in the first adjustment range is low load (CPU utilization <30%), the buffer size is adjusted to 16MB; when the load status in the second adjustment range is medium load (CPU utilization 30%-70%), the total size of the buffer is adjusted to 12MB; when the load status in the third adjustment range is high load (CPU utilization >70%), the total size of the buffer is adjusted to 8MB.

[0084] In an exemplary embodiment, determining the response strategy of the first storage node to the data requests from the plurality of sending ends based on the load state includes: when the load state is in a first response interval, determining that receiving the data requests from the plurality of sending ends is allowed; when the load state is in a second response interval, determining that predicting the remaining time and pause time corresponding to each of the buffers is allowed; and when the load state is in a third response interval, determining that returning a temporary unavailable state to the plurality of sending ends and rejecting receiving the data requests from the plurality of sending ends.

[0085] In other words, a buffer level early warning mechanism is introduced, dividing buffer usage into three response intervals: the first response interval is the safe zone (0%-70%): normal data reception; the second response interval is the early warning zone (70%-90%): the remaining time and pause time corresponding to the buffer are predicted, and the pause time is calculated based on the buffer derivative; the third response interval is the danger zone (90%-100%): new write requests are rejected and a temporary unavailable state is returned.

[0086] To better understand the process of the above data traffic control method, the implementation flow of the above data traffic control method will be described below in conjunction with optional embodiments, but this is not intended to limit the technical solution of the embodiments of this application.

[0087] The technological development of distributed storage systems stems from the explosive growth in demand for massive data storage and processing. With the rise of technologies such as cloud computing, big data, and artificial intelligence (AI), traditional centralized storage, due to its poor scalability and significant performance bottlenecks, struggles to handle exabyte (EB) level data volumes and high-concurrency access. Distributed storage addresses this by distributing data across multiple nodes, utilizing replication redundancy, load balancing, and distributed consensus mechanisms to achieve linear scaling of storage capacity and performance. The widespread adoption of high-speed network technologies such as RDMA has accelerated its application in low-latency scenarios (such as financial transactions and real-time analytics), but distributed storage systems still face challenges such as network congestion caused by incast traffic and maintaining consistency across multiple replicas.

[0088] In distributed storage systems, Remote Direct Memory Access (RDMA) technology achieves low-latency, high-throughput data transmission through RoCE networks. However, the "many-to-one" incast traffic pattern (such as data replication and cluster reconstruction) often leads to switch buffer overflows. PFC, as a Layer 2 congestion control mechanism, relies on a fixed threshold to passively trigger pauses, making it impossible to predict congestion in advance. This results in tail latency as high as 820μs in 16:1 incast scenarios, and also suffers from head-end blocking and traffic unfairness. While end-to-end DCQCN alleviates the shortcomings of PFC, its slow convergence speed still leads to switch queue accumulation during congestion, creating a vicious cycle of "latency-congestion," making it difficult to meet the dual requirements of distributed storage for lossless networks and low latency.

[0089] In other words, RoCE networks achieve lossless transmission through the collaboration of Layer 2 PFC and Layer 3 DCQCN, but the existing mechanisms have significant limitations. PFC passively responds based on switch buffer thresholds, uniformly suspending all traffic on the same input port, easily affecting innocent flows; its coarse-grained control cannot distinguish the main congestion sources, often leading to the blocking of small-volume services in mixed traffic scenarios (such as the coexistence of small metadata input / output (IO) and large data block IO). DCQCN relies on ECN marking and explicit rate control, but the feedback delay results in an excessively long control loop, failing to suppress sudden congestion in a timely manner. Insufficient collaboration between the two leads to a peak buffer ratio of 100% in distributed storage networks under Incast scenarios, with tail latency fluctuations exceeding ±50%, urgently requiring a more precise congestion control mechanism.

[0090] Specifically, to address network congestion between OSDs in a distributed storage system, relevant technologies and solutions related to PFC and DCQCN mainly include those based on traffic scheduling and path optimization, such as:

[0091] 1) Real-time identification mechanism for congested OSD nodes.

[0092] Multi-dimensional monitoring indicators. Buffer pool level monitoring: Real-time collection of receive buffer occupancy rate via the OSD node kernel buffer pool interface. An alert is triggered when the occupancy rate exceeds a threshold (e.g., 70%) for three consecutive sampling periods (1ms / period). Network latency detection: Using active probe packets (sending Internet Control Message Protocol (ICMP) Echo Requests every 50ms), if the RTT between OSDs exceeds 150% of the baseline value (e.g., 2ms) for five consecutive times (i.e., 3ms), it is marked as a congestion-prone node. Traffic anomaly detection: Analyzing traffic patterns between OSDs, if the inbound traffic rate of a certain OSD surges by more than 200% of the historical average within 10ms, accompanied by a PFC pause frame reception frequency ≥ 5 times / second, it is confirmed as a congestion-prone node.

[0093] 2) Dynamic path rerouting based on the controlled replication under scalable hashing (CRUSH) algorithm.

[0094] Topology awareness and traffic reallocation. CRUSH mapping update: After the MON (Monitor node) obtains the list of congested OSDs, it triggers incremental rebalancing (not full reconstruction) using the CRUSH algorithm, remapping replication flows originally destined for congested OSDs to non-congested OSDs. Shard migration strategy: Migration is performed at the data shard granularity, prioritizing the migration of incomplete replication tasks, and synchronizing new path information to all relevant OSDs via the osd map update command; Traffic scheduling priority: Rerouting traffic is prioritized lower than normal service traffic to avoid bandwidth contention, and migration bandwidth is limited to 30% of link capacity.

[0095] 3) Multi-level triggering and execution of load balancing.

[0096] A three-tier triggering mechanism. Primary trigger: When the length of a single OSD's receive queue exceeds 1000 packets (approximately 1MB), local traffic scheduling is initiated, adjusting the traffic distribution ratio to other OSDs within the same rack. Intermediate trigger: When three or more OSDs simultaneously meet the primary condition, cross-rack path rerouting is triggered, distributing traffic to OSDs in other racks via CRUSH. Advanced trigger: When more than 10% of the OSDs across the network experience congestion, global load balancing is initiated, recalculating the storage location of all data shards. Incremental replication mode is enabled when the migration amount exceeds 5% of the total data volume.

[0097] 4) Feedback and verification of congestion mitigation.

[0098] Real-time status feedback: Congested OSDs report buffer occupancy and RTT change curves to the MON node every 100ms. When the indicators recover to within 120% of the baseline value for 10 consecutive cycles, it is marked as congestion relief. Path validity verification: Traffic after rerouting needs to pass a 30-second stability verification. If the RTT fluctuation of the new path exceeds 20% or there are more than 3 PFC pauses, it will fall back to the original path and reselect the target OSD.

[0099] However, the solution based on traffic scheduling and path optimization has the following drawbacks:

[0100] 1) Quantitative impact of scheduling delay: The minimum cycle of CRUSH rebalancing is 100ms, which cannot cope with sudden congestion at the 10ms level (e.g., when the peak duration of Incast traffic is <50ms, congestion has already occurred before scheduling is completed). 2) Specific manifestations of metadata overhead: Each incremental rebalancing requires updating the mapping table of all OSDs in the network. When the cluster size exceeds 500 OSDs, the control plane traffic of a single update can reach 20MB, occupying 10% of the management network bandwidth. 3) Typical scenarios for local optimization: In clusters where cross-rack traffic accounts for more than 40%, path adjustment alone cannot solve the congestion bottleneck of the core switch. For example, in a 16:1 Incast scenario, rerouting may still lead to accumulation of the egress queue of the aggregation switch.

[0101] To address the aforementioned shortcomings, this application's optional embodiments construct a predictive flow control mechanism based on RoCE networks to achieve precise control over distributed storage network congestion. It innovatively introduces a node-level buffer prediction model, utilizing the buffer occupancy derivative combined with network round-trip latency to detect congestion risks 2-3ms in advance, significantly improving response time compared to traditional mechanisms. Based on CRUSH mapping, it accurately locates congestion sources and dynamically selects source OSDs to pause according to traffic proportion, effectively preventing innocent flows from being mistakenly paused and ensuring traffic fairness. Simultaneously, it dynamically adjusts pause times based on storage service characteristics, deeply integrating key processes such as data replication and cluster reconstruction to ensure the continuity and stability of storage services during network congestion, comprehensively improving the overall performance of the distributed storage system in a RoCE network environment.

[0102] Specifically, addressing the issues of innocent flow erroneous pauses caused by the passive response of traditional PFC with fixed thresholds and congestion propagation due to DCQCN feedback delays, this application's optional embodiments construct a node-level buffer derivative prediction model. This model utilizes the buffer occupancy change rate and network round-trip latency to detect congestion risks 2-3ms in advance, achieving a response time more than 80% earlier than traditional mechanisms. It innovatively combines CRUSH mapping with traffic share analysis, dynamically selecting the pause source OSD based on a cumulative traffic share of 75%, and using a fairness adjustment factor to prevent long-term pauses of a single node, thus achieving precise traffic control. Through a four-layer architecture (client-level CRUSH optimization selection, control plane predictive rebalancing, data plane adaptive buffer management, and network congestion control layer dynamic pausing), the system reduces tail latency from approximately 800μs to approximately 200μs in a 16:1 Incast scenario, increasing throughput by 23% while ensuring over 98% traffic fairness, thus solving the challenge of coordinating lossless transmission and low latency in distributed storage networks.

[0103] like Figure 3 As shown, the distributed storage system adopts a decentralized, distributed architecture. Its core components include clients, monitoring nodes, a metadata server (MDS), and operating system devices (OSDs). These modules collaborate via high-speed networks (such as Ethernet and InfiniBand) to achieve data storage and management. The architecture layers are explained below:

[0104] 1) Client layer.

[0105] Positioning: The entry point for users to interact with distributed storage, providing three interfaces: block, object, and file.

[0106] The core components include:

[0107] The librado library is used to encapsulate the underlying Application Programming Interface (API) and is responsible for interacting with the cluster.

[0108] The CRUSH client module is used for local caching of the CRUSH Map and calculating data storage locations. Inputs include: object name and storage pool rules (such as the number of replicas and fault domain level). The processing steps are: hashing the object name to generate a random number; traversing the CRUSH bucket level (e.g., host → rack → data center) according to the storage pool rules, and selecting the OSD based on its weight. This ensures balanced data distribution and meets redundancy policies (e.g., deploying replicas across racks).

[0109] 2) Control plane (Monitor / MDS).

[0110] Monitor: Maintains cluster metadata (OSD Map, CRUSH Map, authentication information) and ensures consistency across multiple nodes through the Paxos algorithm.

[0111] Interaction with the client: When the client connects for the first time, it pulls the latest cluster map and updates it synchronously via heartbeat.

[0112] MDS (File Scenario): Manages file system metadata (directory tree, permissions), and caches hot metadata to accelerate access.

[0113] 3) Data plane (OSD cluster).

[0114] OSD Node: Hardware components include CPU, memory, disks (Hard Disk Drive (HDD) / Solid State Drive (SSD)), and network interface (Gigabit / 10 Gigabit Ethernet). It is responsible for data storage, replication, erasure coding, consistency maintenance, and understands the cluster topology through the OSD Map.

[0115] In other words, based on the traditional distributed storage system architecture, and combined with RoCE (RDMA over Converged Ethernet) networking and predictive flow control mechanisms, the optional embodiments of this application redefine the data flow path and processing logic between the client and the OSD cluster, optimize the IO process of the distributed storage system, and improve the performance and reliability of the distributed storage system through low-latency RDMA transmission and precise congestion control. The following... Figure 4 The overall process, key steps, and algorithm model of this method are further elaborated in detail.

[0116] like Figure 4 As shown, the IO process of the optional embodiment of this application is based on a four-layer architecture design, including a client layer, a control plane layer, a data plane layer, and a network congestion control layer. Each layer achieves low-latency communication through a RoCE network. The overall process can be divided into the following stages:

[0117] The first phase, client request processing: The client application uses the librado library to convert IO requests into object operations and calculates the data storage location using a CRUSH Map cache. The CRUSH algorithm is extended by introducing a network congestion factor. This serves as the basis for OSD selection. The calculation formula is detailed in the following steps. During client initialization, the librados library automatically identifies RoCE-enabled devices via the DCBx (Data Center Bridging Exchange) protocol, establishes an RDMA connection pool, and configures parameters such as the number of queue pairs (QPs) and the send / receive queue depth for each connection, laying the foundation for low-latency transmission. Simultaneously, the local CRUSH Map cache employs a predictive update strategy, dynamically adjusting the cache update cycle based on the cluster's historical update frequency, reducing interaction overhead with the control plane.

[0118] The second phase, the control plane interaction phase, involves the client interacting with the Monitor cluster to obtain the latest cluster status and with the MDS (file system scenario) to obtain metadata. Upon initial client access, the client pulls the complete CRUSH Map and cluster status information from the Monitor; subsequent updates are synchronized via a heartbeat mechanism, with heartbeat packets also transmitted over the RoCE network. In the file system scenario, the MDS and client use RDMA read operations to obtain metadata, and also feature metadata prefetching, loading relevant metadata in advance based on client access patterns and dynamically adjusting the metadata caching strategy according to network congestion.

[0119] The third stage, the data plane transmission stage: The client directly transmits data to the OSD node via the RoCE network. After the primary OSD completes local storage, it replicates the data to the replica OSD via the RoCE network. Data transmission uses zero-copy RDMA operations. When the primary OSD replicates data to the replica OSD, a batch replication strategy is used to reduce operational overhead. In data recovery scenarios, recovery bandwidth is dynamically adjusted based on a predictive flow control mechanism, and a gradual recovery strategy is adopted to avoid impacting normal business operations with recovery traffic.

[0120] The fourth stage, predictive congestion control: Throughout the data transmission process, the predictive flow control mechanism monitors network congestion status in real time. Using a buffer derivative prediction model, it detects congestion risks 2-3ms in advance and dynamically adjusts the PFC pause time and target port to achieve precise congestion control. The predictive flow control module continuously collects data such as buffer status and network latency for real-time analysis and calculation. Once a congestion risk is detected, a PFC pause frame is immediately generated, the target port is determined, and the frame is sent. Simultaneously, the pause effect is monitored, and subsequent control strategies are adjusted based on the actual situation to form a closed-loop control system.

[0121] In summary, the optional embodiments of this application, through optimized design of each stage, achieve efficient coordination throughout the entire process from client requests to data storage, transmission, and congestion control, significantly improving the performance and reliability of the distributed storage system in the RoCE network environment.

[0122] The following provides a more detailed explanation of the four stages mentioned above.

[0123] Furthermore, the client-side request processing and RoCE connection establishment process includes:

[0124] Step 1: Initialize the RoCE network and librados library.

[0125] When the distributed storage client application starts, the librado library automatically discovers RoCE-enabled devices in the network via the DCBx protocol. An RDMA connection pool is established, with each connection configured with the following parameters: Number of QPs: 16 by default, supporting concurrent RDMA operations; Send Queue Depth: 1024; Receive Queue Depth: 1024; Maximum RDMA read request size: 64KB; Maximum RDMA write request size: 4MB.

[0126] Step 2: Local CRUSH Map caching and predictive updates.

[0127] When a client first connects to the cluster, it retrieves the complete CRUSH Map from the Monitor and caches it in local memory. A predictive update mechanism is introduced to dynamically adjust the request cycle based on historical update frequency: when the cluster is running stably (updating less than once per hour), the update cycle is 10 minutes; when the cluster is actively changing (updating more than 3 times per hour), the update cycle is shortened to 2 minutes; the local CRUSH Map cache size is limited to 10MB, and the Least Recently Used (LRU) strategy is used to evict old entries.

[0128] Step 3: OSD selection for network congestion awareness.

[0129] Extending the traditional CRUSH algorithm, because in distributed storage, the "congestion state" of OSDs is affected by buffer occupancy, network latency, and load pressure, this application's optional embodiments introduce a network congestion factor. As an important parameter for OSD selection:

[0130] ;

[0131] Where: buffer occupancy This directly reflects the "real-time pressure" of OSD data reception. When the buffer is near full, new data writing is prone to triggering packet loss or delay. Network latency. This reflects the transmission efficiency from the client to the OSD. High latency means longer data round-trip times, and this, combined with write operations, can exacerbate congestion. Load factor The combined input / output operations per second (IOPS) and bandwidth usage reflects the "busyness" of OSD processing capacity. Excessive load can indirectly lead to network queue accumulation. This indicates the weight of the buffer occupancy; , representing the network latency weight; This indicates the load factor weight, prioritizing buffer safety (highest weight) to avoid packet loss due to buffer overflow; secondly, focusing on network latency (ensuring transmission efficiency); and finally, balancing the load (preventing the spread of local overload), aligning with the distributed storage requirements of "ensuring no data loss first, then ensuring efficient transmission, and finally achieving global balance." The baseline round-trip time is 2ms by default. The average load factor for cluster OSDs is calculated as follows:

[0132] ;

[0133] Clients prioritize congestion factors When all candidate OSDs have a congestion factor exceeding 0.6, the Monitor rebalancing process is triggered, which is a data migration scheduling process: According to the new CRUSH Map, "shard migration" is initiated between OSDs. The primary OSD will gradually transfer the replicas of the overloaded shards to low-congestion nodes. During the process, predictive flow control is used to avoid the migration traffic from aggravating congestion (such as dynamically limiting migration bandwidth and prioritizing business IO).

[0134] Compared to the related CRUSH algorithm, which relies on "static weights" to select OSDs and where rebalancing is only triggered when an OSD fails, the optional embodiments of this application utilize... Real-time congestion detection and proactive rebalancing upgrades "passive fault recovery" to "proactive congestion prevention," preventing global performance collapse caused by the spread of local congestion.

[0135] Furthermore, control plane interaction and RoCE acceleration include:

[0136] Step 1: RDMA-accelerated Monitor cluster communication process. Monitor nodes communicate via RDMA through the RoCE network, reducing Paxos protocol message delivery latency from 50μs in traditional Ethernet to 10μs. A batch acknowledgment mechanism is introduced, acknowledging messages after every 100 prepared / accepted messages are collected, reducing the number of message interactions and improving Paxos protocol efficiency.

[0137] Step 2: Predictive Cluster State Maintenance. The Monitor monitors the network congestion status of OSD nodes in real time and builds a congestion prediction model.

[0138] ;

[0139] in: This is an indicator function used to determine the congestion factor of a single OSD. Whether it exceeds 0.8 (which can be understood as "the OSD is in a 'high-risk congestion' state"), return 1 if the condition is met, otherwise return 0; The total number of OSD nodes in the cluster, when the congestion probability... In such cases, CRUSH Map rebalancing is triggered in advance, with the rebalancing cycle dynamically adjusted: 5 minutes in normal state, 30 seconds in high congestion warning state, and 10 seconds in congestion state. Assuming the cluster has 100 OSDs: if 8 OSDs... ,but (Less than 0.1), no rebalancing is triggered, and a 5-minute cycle check is maintained. If 12 OSDs... ,but If the value is greater than 0.1, a CRUSH Map rebalancing is immediately triggered, and the subsequent check cycle is shortened to 30 seconds until the congestion is relieved.

[0140] Step 3: Integration of MDS and RoCE (file system scenario). MDS and the client implement RDMA read operations through the RoCE network, reducing the metadata read latency from 100μs in traditional Ethernet to 20μs.

[0141] A metadata prefetching mechanism is introduced so that when a client accesses the directory ( / home / user), the metadata of its subdirectories ( / home / user / data) is automatically prefetched.

[0142] Network congestion-aware metadata caching strategy:

[0143] Normal network conditions: Metadata cache hit rate target is 80%.

[0144] High congestion network status: Metadata cache hit rate target increased to 90%.

[0145] The Least Frequently Used (LFU) cache replacement strategy is adopted to prioritize the retention of metadata that is accessed frequently.

[0146] Cache size is dynamically adjusted: the base size is 2GB, and it increases by 100MB every 5 seconds when network congestion worsens, with a maximum of 4GB.

[0147] Further optimizations include data plane transport and RoCE features, including:

[0148] Step 1, the integration of the RoCE protocol stack and predictive flow control, aims to predict buffer congestion trends in advance during RoCE network communication on OSD nodes. This prevents packet loss and latency spikes due to buffer overflows, resulting in more stable data transmission. The OSD node deploys an enhanced RoCE protocol stack, adding a predictive flow control module to the traditional RoCE v2. The predictive flow control module monitors the rate of change of the receive buffer in real time. The calculation formula is:

[0149] ;

[0150] in: , where is the sampling time interval; Let t be the size of the buffer at time t; a sliding window averaging mechanism is introduced to reduce the impact of instantaneous fluctuations (such as misjudgments caused by sudden small flows).

[0151] ;

[0152] in, , representing a sliding window with 3 sampling points. A positive rate of change indicates the buffer is in a state of "continuously increasing occupancy" (potential congestion); a negative rate of change indicates the buffer is in a state of "releasing space" (state ease).

[0153] For example, assuming the total capacity of the OSD receive buffer is 100MB, and the sampling interval is... Sliding window :

[0154] First sample (t=10ms): The last time (t=0) was 25MB, then ;

[0155] Second sampling (t=20ms): The last time (t=10) was 30MB, then ;

[0156] Third sampling (t=30ms): The last time (t=20) it was 38MB, then .

[0157] After moving average: .

[0158] If this rate of change continues, the buffer will increase by 0.67 × 100 = 67 MB after 100 ms, which, combined with the current 45 MB, will cause an overflow (assuming a total capacity of 100 MB). At this point, the predictive flow control module can trigger actions in advance (such as notifying the sender to slow down at t=40 ms) to prevent actual overflow. Essentially, this quantifies the "buffer change trend" through mathematical modeling, enabling OSD nodes to "predict congestion," which is particularly suitable for RoCE network scenarios with low latency, high bandwidth, but sensitivity to "burst traffic," and is a key part of distributed storage data plane optimization.

[0159] Step 2: Adaptive Buffer Management (OSD) process. Nodes dynamically adjust the receive buffer size based on their load status: Low load (CPU utilization <30%): 16MB buffer size; Medium load (CPU utilization 30%-70%): 12MB buffer size; High load (CPU utilization >70%): 8MB buffer size. A buffer level warning mechanism is introduced, dividing buffer usage into three zones: Safe zone (0%-70%): Normal data reception; Warning zone (70%-90%): Predictive flow control is activated, and pause time is calculated based on the buffer derivative; Danger zone (90%-100%): New write requests are rejected and a temporary unavailable state is returned.

[0160] Step 3: RoCE-based data replication and recovery process between OSDs. RDMA-accelerated data replication: When the master OSD sends data to the replica OSD, zero-copy RDMA write operations are implemented through the RoCE network, reducing data replication latency from 200μs in traditional Ethernet to 50μs. A batch replication strategy is adopted, replicating 16 objects (approximately 64MB of data) in each batch, reducing RDMA operation overhead.

[0161] Further, predictive congestion control mechanisms include:

[0162] Step 1: Buffer Derivative Prediction Model. Based on the buffer occupancy rate and its rate of change, a congestion prediction model is constructed to calculate the remaining time. : in: This is the total size of the buffer. This represents the current size of the buffer. The sliding window average buffer is used, and the rate of change RTT is the round-trip time, measured by periodically sending probe packets. When the remaining time... When a potential congestion risk is identified, predictive flow control is triggered. This involves calculating how long it will take for the buffer to fill up and then subtracting the impact of network latency. This allows for early detection of impending congestion risks, enabling the storage system to proactively intervene before the buffer becomes full (e.g., slowing down data transmission from clients) to prevent actual packet loss and latency spikes.

[0163] Step 2: Calculate the dynamic pause time based on the prediction model to calculate the PFC pause time. : ;in: For safety, the default value is 0.2, and a historical adjustment factor is introduced. The adjustments will be made dynamically based on the effects of the first five pauses. ;

[0164] in: Let be the congestion relief time predicted for the i-th time. The final pause time is adjusted to be the actual congestion relief time for the i-th time. .

[0165] Step 3: Precise congestion source location and selective pause process. Based on the CRUSH mapping relationship, analyze the distribution of traffic sources and calculate the traffic proportion of each source OSD. :

[0166] Where: m is the total number of source OSDs sending traffic to the target OSD. A selective pause strategy is adopted, prioritizing the pause of source OSDs with a high traffic percentage: when a pause is needed, source OSDs are sorted in descending order of traffic percentage; OSDs are selected from high to low until the cumulative traffic percentage exceeds 75%; a PFC pause frame is sent to the selected source OSDs, while unselected source OSDs continue to transmit data, and a fairness adjustment factor is introduced. Avoid prolonged suspension of the same source OSD:

[0167] ;

[0168] in: This represents the total time that source OSD i has been paused in the past minute. The statistical period is used to determine the final weights. Calculation formula: The fairness adjustment factor The calculation must satisfy the following condition: when the source OSD has been paused for the past minute. If it exceeds 30 seconds, Automatically increase to 0.5 or higher, reduce the pause priority of this OSD, and avoid the problem of traffic starvation caused by long-term preemption.

[0169] By deeply integrating the low latency and high bandwidth characteristics of RoCE networks with predictive flow control mechanisms, the optional embodiments of this application achieve precise control of congestion in distributed storage networks, significantly improving the performance and reliability of the system under high concurrency and high load scenarios, and providing a more efficient storage solution for large-scale data centers.

[0170] In summary, the RoCE-based predictive flow control method (equivalent to the data flow control method in the above embodiments) proposed in the optional embodiments of this application significantly improves the performance and stability of distributed storage networks through innovative congestion prediction and precise flow control mechanisms. Compared with traditional PFC and DCQCN mechanisms, the optional embodiments of this application, through node-level buffer state prediction, detect congestion risks 2-3ms in advance, reducing the tail latency in the 16:1 Incast scenario from 830μs to 274μs, improving the response speed by more than 80%, and effectively solving the buffer accumulation problem caused by the passive response of traditional mechanisms.

[0171] Regarding traffic fairness, CRUSH mapping accurately locates congestion sources and dynamically selects pause targets based on traffic share, avoiding the impact of traditional global pauses on innocent traffic. Actual traffic fairness is improved to over 98%. Simultaneously, through dynamic adjustment of pause times and adaptive buffer management, network resource utilization is increased from 60%-70% to 75%-85%, and throughput is increased by 23%, achieving a balance between lossless transmission and high bandwidth utilization.

[0172] Addressing the characteristics of distributed storage services, this solution deeply integrates data replication and cluster reconstruction processes. Through progressive bandwidth control and priority scheduling, it ensures the continuity of core services during congestion. In a 500-node cluster, performance fluctuations during fault recovery are less than 20%, and metadata update overhead is reduced by 50%, effectively improving system stability under large-scale deployments. This solution overcomes the challenges of coordinating low latency, lossless transmission, and fairness through a four-layer architecture, providing an efficient congestion control solution for exabyte-scale data centers.

[0173] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.

[0174] This embodiment also provides a data traffic control device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0175] Figure 5 This is a structural block diagram of a data flow control device according to an embodiment of this application, such as... Figure 5 As shown, the device includes:

[0176] Prediction module 52 is used to predict the remaining time and pause time corresponding to the buffer of the first storage node based on the buffer change rate of the first storage node. The remaining time is used to indicate the time interval between the current time and the time when the buffer is filled. The pause time is used to indicate the time when the first storage node needs to pause receiving data traffic if the remaining time is less than a time threshold.

[0177] The instruction module 54 is used to select a target sender from multiple senders corresponding to the source distribution of the data traffic of the first storage node when the remaining time is less than the time threshold, and instruct the target sender to suspend sending data traffic to the first storage node based on the pause time.

[0178] The aforementioned device predicts the remaining time and pause time corresponding to the buffer of the first storage node based on the buffer change rate of the first storage node. The remaining time indicates the time interval from the current moment to the moment the buffer is filled, and the pause time indicates the time during which the first storage node needs to pause receiving data traffic when the remaining time is less than a time threshold. When the remaining time is less than the time threshold, a target sender is selected from multiple senders corresponding to the source distribution of the data traffic to the first storage node, and the target sender is instructed to pause sending data traffic to the first storage node based on the pause time. Therefore, this addresses the problem of inaccurate handling of network congestion in distributed storage systems in related technologies, achieving the technical effect of predicting and accurately handling network congestion in distributed storage systems.

[0179] In an exemplary embodiment, the apparatus further includes a determining module, configured to continuously sample the occupied size of the buffer at sampling time intervals before predicting the remaining time and pause time corresponding to the buffer of the first storage node by means of the buffer change rate of the first storage node, thereby obtaining a plurality of sampled data pairs, wherein each sampled data pair includes: the occupied size of the buffer at two times before and after the sampling time interval; and determine the buffer change rate by means of the plurality of sampled data pairs.

[0180] In an exemplary embodiment, the determining module is further configured to calculate the target buffer change rate corresponding to each sampled data pair using each sampled data pair and the sampling time interval; and input multiple consecutive target buffer change rates into a sliding window to obtain the buffer change rate.

[0181] In one exemplary embodiment, the apparatus further includes a notification module, configured to determine, after determining the buffer change rate through the plurality of sampled data pairs, whether the buffer is in an overflow state after a first target time through the buffer change rate; if the buffer is in an overflow state after the first target time, instructing the plurality of sending ends to reduce the rate of data transmission at a notification time, wherein the notification time is separated from the current time by a second target time, and the second target time is less than 1 / 2 of the first target time.

[0182] In an exemplary embodiment, the prediction module is further configured to predict the remaining time using the buffer change rate and the round-trip delay corresponding to the first storage node; and to predict the pause time using the buffer change rate and a safety factor.

[0183] In one exemplary embodiment, the prediction module is further configured to determine the difference between the total size of the buffer and the occupied size of the buffer at the current moment, and to determine the ratio of the difference to the rate of change of the buffer; and to predict the remaining time using the ratio and the round-trip delay.

[0184] In an exemplary embodiment, the prediction module is further configured to determine the difference between the total size of the buffer and the occupied size of the buffer at the current moment, and to determine the ratio of the difference to the rate of change of the buffer; and to predict the pause time using the ratio and the safety factor.

[0185] In an exemplary embodiment, the prediction module is further configured to determine a historical adjustment factor using the previous k historical adjustment data at the current time, wherein the i-th historical adjustment data in the previous k historical adjustment data includes: the pause time obtained from the i-th prediction and the actual time of the i-th unblocking of the first storage node, where k is a positive integer and i = 1, 2, 3...k; and to calculate a target pause time using the ratio and the safety factor; and to adjust the target pause time using the historical adjustment factor to obtain the pause time.

[0186] In an exemplary embodiment, the indicating module is further configured to determine the total source amount of data traffic to the first storage node, and determine each source amount corresponding to each sender; determine the traffic proportion of each sender among the plurality of senders by the ratio of each source amount to the total source amount, wherein the traffic proportion is used to indicate the source distribution; select a target traffic proportion from the plurality of traffic proportions sorted according to a target order, and take the sender corresponding to the target traffic proportion as the target sender.

[0187] In an exemplary embodiment, the instruction module is further configured to: send a preset data frame to the target sender, wherein the preset data frame is used to instruct the target sender to pause sending data traffic to the first storage node; and update the target sender by: calculating a fairness adjustment factor corresponding to each sender in the target sender, updating the traffic ratio of each sender among the plurality of senders through the fairness adjustment factor, and updating the target sender through the updated traffic ratio; and repeatedly executing the above sending and updating steps until a preset condition is met, determining that the instruction of the target sender to pause sending data traffic to the first storage node based on the pause time has been completed, wherein the preset condition includes one of the following: the pause time ends, or the remaining time is detected to be greater than or equal to the time threshold.

[0188] In one exemplary embodiment, the instruction module is further configured to determine the total time each transmitter was suspended in the most recent statistical period; and to determine the fairness adjustment factor by the ratio of the total time to the period time corresponding to the statistical period.

[0189] In one exemplary embodiment, the apparatus further includes a congestion factor determination module, configured to, when there is data traffic to be stored in a client communicating with a distributed storage system including the first storage node, calculate the buffer occupancy, network latency, and load factor of the first storage node respectively; perform a weighted summation of the buffer occupancy, the network latency, and the load factor to determine the network congestion factor of the first storage node; and instruct the client to select a second storage node in the distributed storage system corresponding to the data traffic to be stored based on the network congestion factor.

[0190] In an exemplary embodiment, the congestion factor determination module is further configured to instruct a monitoring node in a distributed storage system including the first storage node to perform shard migration of the first storage node when the network congestion factor is higher than a preset threshold; instruct the monitoring node to calculate the congestion status of multiple storage nodes in the distributed storage system using a congestion prediction model and multiple network congestion factors; and instruct the monitoring node to perform shard migration of the first storage node according to the congestion status.

[0191] In one exemplary embodiment, the apparatus further includes a response strategy determination module, configured to determine the load state of the first storage node before predicting the remaining time and pause time corresponding to the buffers of the first storage node based on the buffer change rate of the first storage node, wherein the load state includes: CPU utilization; adjusting the total size of the buffers based on the load state; and determining the response strategy of the first storage node to the data requests of the plurality of sending ends based on the load state.

[0192] In one exemplary embodiment, the response strategy determination module is further configured to: determine the total size of the buffer as a first value when the load state is in a first adjustment range; determine the total size of the buffer as a second value when the load state is in a second adjustment range; and determine the total size of the buffer as a third value when the load state is in a third adjustment range; wherein the first value is greater than the second value, the second value is greater than the third value, the maximum value in the first adjustment range is less than the minimum value in the second adjustment range, and the maximum value in the second adjustment range is less than the minimum value in the third adjustment range.

[0193] In an exemplary embodiment, the response strategy determination module is further configured to: when the load state is in a first response interval, determine whether to allow receiving data requests from the plurality of sending ends; when the load state is in a second response interval, determine whether to allow predicting the remaining time and pause time corresponding to each of the buffers; and when the load state is in a third response interval, determine to return a temporary unavailable state to the plurality of sending ends and refuse to receive the data requests from the plurality of sending ends.

[0194] For a description of the features in the embodiment corresponding to the data flow control device, please refer to the relevant description in the embodiment corresponding to the data flow control method, which will not be repeated here.

[0195] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above-described data flow control method embodiments.

[0196] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described data traffic control method embodiments when running.

[0197] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0198] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described data flow control method embodiments.

[0199] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described data flow control method embodiments.

[0200] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0201] The data traffic control method provided in this application has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A method of controlling data traffic, characterized by, The method comprises: predicting, by a buffer change rate of a first storage node, a remaining time and a pause time corresponding to a buffer of the first storage node respectively, wherein the remaining time is used to indicate a time interval from a current time to a time when the buffer is filled, and the pause time is used to indicate a time when the first storage node needs to pause receiving data flow in a case where the remaining time is less than a time threshold; in a case where the remaining time is less than the time threshold, selecting, by a source distribution of data flow of the first storage node, a target sending end from a plurality of sending ends corresponding to the source distribution, and instructing the target sending end to pause sending data flow to the first storage node based on the pause time, wherein predicting, by a buffer change rate of a first storage node, a remaining time and a pause time corresponding to a buffer of the first storage node respectively comprises: predicting the remaining time by the buffer change rate and a round-trip delay corresponding to the first storage node; and predicting the pause time by the buffer change rate and a safety coefficient, wherein predicting the remaining time by the buffer change rate and a round-trip delay corresponding to the first storage node comprises: determining a difference between a total size of the buffer and an occupied size of the buffer at the current time, and determining a ratio of the difference to the buffer change rate; predicting the remaining time by the ratio and the round-trip delay; and predicting the pause time by the ratio and the safety coefficient; wherein predicting the pause time by the ratio and the safety coefficient comprises: determining a historical adjustment factor by k previous historical adjustment data of the current time, wherein the i-th historical adjustment data in the k previous historical adjustment data comprises an i-th predicted pause time and an actual time when the first storage node is i-thly relieved from congestion, wherein k is a positive integer and i = 1, 2, 3, …, k; calculating a target pause time by the ratio and the safety coefficient; and adjusting the target pause time by the historical adjustment factor to obtain the pause time; Wherein, the remaining time The calculation formula is: ; target pause time The calculation formula is: ; history adjustment factor The calculation method is: ; pause time The calculation formula is: ; wherein, represents the total size of the buffer, represents the occupation size at the current time, represents the buffer change rate, represents the round-trip delay; is a safety factor; is the pause time obtained by the i-th prediction, is the actual time of the i-th decongestion of the first storage node.

2. The data traffic control method according to claim 1, characterized by, Before predicting, by a buffer change rate of a first storage node, a remaining time and a pause time corresponding to a buffer of the first storage node respectively, the method further comprises: sampling the occupied size of the buffer continuously at a sampling time interval to obtain a plurality of sampling data pairs, wherein each sampling data pair comprises an occupied size corresponding to the buffer at two time points before and after the sampling time interval; determining the buffer change rate by the plurality of sampling data pairs.

3. The data traffic control method according to claim 2, wherein Determining the buffer change rate by the plurality of sampling data pairs comprises: calculating a target buffer change rate corresponding to each sampling data pair by the each sampling data pair and the sampling time interval; inputting a plurality of the target buffer change rates continuously into a sliding window to obtain the buffer change rate.

4. The data traffic control method according to claim 2, characterized by, After determining the buffer change rate by the plurality of sampling data pairs, the method further comprises: determining whether the buffer is in an overflow state after a first target time based on the buffer change rate; in a case where the buffer is in the overflow state after the first target time, instructing the plurality of sending ends to reduce the speed of sending data flow at a notification time, wherein the notification time is spaced apart from the current time by a second target time, and the second target time is less than 1 / 2 of the first target time.

5. The data traffic control method of claim 1, wherein, selecting a target sending end from the plurality of sending ends corresponding to the source distribution based on the source distribution of the data flow of the first storage node, including: determining a total source amount of the data flow of the first storage node, and determining each source amount corresponding to each sending end; determining a flow proportion of each sending end in the plurality of sending ends based on a ratio of each source amount to the total source amount, wherein the flow proportion is used to indicate the source distribution; selecting a target flow proportion from the plurality of flow proportions sorted in a target order, and taking the sending end corresponding to the target flow proportion as the target sending end.

6. The data traffic control method of claim 1, wherein, indicating the target sending end to pause sending data flow to the first storage node based on the pause time, including: sending step: sending a preset data frame to the target sending end, wherein the preset data frame is used to instruct the target sending end to pause sending data flow to the first storage node; updating step: calculating a fairness adjustment factor corresponding to each sending end in the target sending end, and updating the flow proportion of each sending end in the plurality of sending ends based on the fairness adjustment factor, and updating the target sending end based on the updated flow proportion; cyclically executing the above sending step and updating step until a preset condition is met, to determine that the target sending end has completed pausing sending data flow to the first storage node based on the pause time, wherein the preset condition includes one of the following: the pause time ends, and it is detected that the remaining time is greater than or equal to the time threshold.

7. The data traffic control method according to claim 6, wherein calculating a fairness adjustment factor corresponding to each sending end in the target sending end, including: determining a total paused time of each sending end in a latest statistical period; determining the fairness adjustment factor based on a ratio of the total time to a period time corresponding to the statistical period.

8. The data traffic control method of claim 1, wherein, after indicating the target sending end to pause sending data flow to the first storage node based on the pause time, the method further includes: in a case where there is to-be-stored data flow in a client in communication with a distributed storage system including the first storage node, calculating a buffer occupancy, a network delay, and a load factor of the first storage node respectively; performing weighted summation on the buffer occupancy, the network delay, and the load factor to determine a network congestion factor of the first storage node, and instructing the client to select a second storage node corresponding to the to-be-stored data flow in the distributed storage system based on the network congestion factor.

9. The data traffic control method according to claim 8, wherein after performing the weighted summation on the buffer occupancy, the network delay, and the load factor to determine the network congestion factor of the first storage node, the method further includes at least one of the following: The monitoring node in the distributed storage system including the first storage node is instructed to perform shard migration on the first storage node when the network congestion factor is higher than a preset threshold value. The monitoring node is instructed to calculate the congestion state of a plurality of storage nodes in the distributed storage system through a congestion prediction model and a plurality of network congestion factors, and the monitoring node is instructed to perform shard migration on the first storage node according to the congestion state.

10. The data traffic control method of claim 1, wherein, Before predicting the remaining time and the pause time corresponding to the buffer of the first storage node through the buffer change rate of the first storage node, the method further comprises: determining the load state of the first storage node, wherein the load state comprises: central processing unit usage rate; adjusting the total size of the buffer through the load state, and determining the response strategy of the first storage node to the data requests of the plurality of sending ends through the load state.

11. The data traffic control method according to claim 10, wherein Adjusting the total size of the buffer through the load state comprises: determining the total size of the buffer as a first value when the load state is in a first adjustment interval; determining the total size of the buffer as a second value when the load state is in a second adjustment interval; determining the total size of the buffer as a third value when the load state is in a third adjustment interval; wherein the first value is greater than the second value, and the second value is greater than the third value; the maximum value in the first adjustment interval is less than the minimum value in the second adjustment interval, and the maximum value in the second adjustment interval is less than the minimum value in the third adjustment interval.

12. The data traffic control method according to claim 10, wherein Determining the response strategy of the first storage node to the data requests of the plurality of sending ends through the load state comprises: determining to allow receiving the data requests of the plurality of sending ends when the load state is in a first response interval; determining to predict the remaining time and the pause time corresponding to the buffer when the load state is in a second response interval; determining to return a temporary unavailable state to the plurality of sending ends and refuse to receive the data requests of the plurality of sending ends when the load state is in a third response interval.

13. A data traffic control device, characterized by comprising: Comprise: a prediction module for predicting the remaining time and the pause time corresponding to the buffer of the first storage node through the buffer change rate of the first storage node, wherein the remaining time is used to indicate the time interval from the current time to the time when the buffer is filled, and the pause time is used to indicate the time when the first storage node needs to pause receiving data flow in the case that the remaining time is less than a time threshold value; an instruction module for selecting a target sending end from a plurality of sending ends corresponding to a source distribution of data flow of the first storage node through the source distribution when the remaining time is less than the time threshold value, and instructing the target sending end to pause sending data flow to the first storage node based on the pause time, The prediction module is further configured to predict the remaining time based on the buffer change rate and a round-trip delay corresponding to the first storage node, and predict the pause time based on the buffer change rate and a safety factor. The prediction module is further configured to determine a difference between a total size of the buffer and an occupied size of the buffer at the current time, and determine a ratio of the difference to the buffer change rate. The prediction module is further configured to predict the remaining time based on the ratio and the round-trip delay, and predict the pause time based on the ratio and the safety factor. The prediction module is further configured to determine a historical adjustment factor based on k previous historical adjustment data at the current time, wherein the i-th historical adjustment data in the k previous historical adjustment data comprises a pause time obtained by i-th prediction and an actual time at which the first storage node is i-th time decongested, and k is a positive integer and i=1, 2, 3, …, k. The prediction module is further configured to calculate a target pause time based on the ratio and the safety factor, and adjust the target pause time based on the historical adjustment factor to obtain the pause time. wherein the remaining time is calculated as: ; the target pause time is calculated as: ; the history adjustment factor is calculated as: ; the pause time is calculated as: ; wherein, represents the total size of the buffer, represents the occupancy size at the current time, represents the buffer change rate, represents the round-trip delay; is a safety factor; is the pause time obtained by the i-th prediction, is the actual time of the first storage node to release congestion for the i-th time.

14. An electronic device, comprising: The computer program comprises the following steps: a memory for storing a computer program; a processor for executing the computer program to implement the steps of the data flow control method according to any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The computer program is stored in the computer readable storage medium, and when executed by the processor, the computer program implements the steps of the data flow control method according to any one of claims 1 to 12.

16. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the data flow control method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Data memory management method and system

    CN110888846A

  • Method and system for controlling network congestion of 5G communication virtualized network element

    CN112996044A

  • Network congestion control method, device, system and equipment and storage medium

    CN117834538A