Routing convergence method and device based on bidirectional forwarding detection protocol

By using a routing convergence method based on a bidirectional forwarding detection protocol, network detection parameters are dynamically adjusted, which solves the shortcomings of the traditional BGP protocol in Decode pool service adaptation and realizes stable interaction and efficient fault recovery between the Decode pool and the Prefill pool in large model inference scenarios.

CN121509322APending Publication Date: 2026-02-10CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511640912.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-11
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

The traditional BGP protocol is not compatible with Decode pool services, making it difficult for large-scale data center networks to meet the inference service needs of massive user concurrency scenarios. Enterprise private computing resources are limited in scale and cannot support the large-scale application needs of large model inference scenarios.

Method used

A routing convergence method based on a bidirectional forwarding detection protocol is adopted. By obtaining the network link bandwidth and KV Cache value between the Decode pool and the Prefill pool, network detection parameters are calculated, network link anomalies are detected and handled, and the network detection parameters are dynamically adjusted to adapt to changes in the service load of the Decode pool.

Benefits of technology

It improves the long-term operational stability of the interaction between the Decode pool and the Prefill pool under the PD separation architecture in large model inference, meets the fault recovery time requirements of the Decode pool, adapts to changes in business load, and enhances the real-time performance and reliability of network detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509322A_ABST
    Figure CN121509322A_ABST
Patent Text Reader

Abstract

The invention relates to a routing convergence method and device based on a bidirectional forwarding detection protocol, which are applied to network control equipment, and the network control equipment is used for managing interaction between a Decode pool and a Prefill pool under a PD separation architecture in large model reasoning. The method comprises the following steps: acquiring a network link bandwidth between a Decode pool and a Prefill pool and a KV Cache value of the Decode pool; calculating a network detection parameter according to the network link bandwidth, the KV Cache value and preset fault recovery time; the fault recovery time is less than the maximum fault recovery time allowed by the Decode pool; detecting a network link between the Decode pool and the Prefill pool according to the network detection parameters to obtain a detection result; and if the detection result shows that the network link is abnormal, link abnormity processing is carried out. According to the method provided by the invention, a traditional BGP protocol and a technology matched with the traditional BGP protocol can be adapted when the Decode pool service is adapted, and the delay is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of communication, in particular to a route convergence method and device based on bidirectional forwarding detection protocol, communication equipment, computer readable storage medium and computer program product. BACKGROUND

[0002] With the Deepseek open source model significantly reducing the development threshold of AI (Artificial Intelligence) application, the user-side inference demand presents an explosive growth. However, the enterprise private computing power resources generally have the problem of limited scale, which is difficult to support the inference service demand in the massive user concurrent scene. Under this background, by splitting the inference process into two stages of Prefill and Decode, the PD separation architecture can effectively adapt to the large-scale application demand of the large model inference scene, and gradually becomes the mainstream technical solution in this field.

[0003] At present, the BGP protocol is generally used to realize route exchange in the Underlay network of large-scale data center.

[0004] However, the traditional BGP protocol and the technology matched therewith are not suitable for adapting to the Decode pool business, and have significant defects. SUMMARY

[0005] Therefore, it is necessary to provide a route convergence method and device based on bidirectional forwarding detection protocol which can adapt to the Decode pool business, communication equipment, computer readable storage medium and computer program product.

[0006] In a first aspect, the present application provides a route convergence method based on bidirectional forwarding detection protocol, applied to a network control device, the network control device being used to manage the interaction between Decode pool and Prefill pool under the PD separation architecture in large model inference, comprising:

[0007] obtaining the network link bandwidth between Decode pool and Prefill pool and the KV Cache value of Decode pool;

[0008] calculating network detection parameters according to the network link bandwidth, the KV Cache value and the predetermined fault recovery time, the fault recovery time being less than the maximum fault recovery time allowed by Decode pool;

[0009] detecting the network link between Decode pool and Prefill pool according to the network detection parameters to obtain a detection result;

[0010] if the detection result represents that the network link is abnormal, performing link abnormality processing.

[0011] In one embodiment, the network detection parameter is calculated according to the network link bandwidth, the KV Cache value, and the predetermined fault recovery time, including:

[0012] Obtaining the bandwidth fluctuation coefficient of the network link between the Decode pool and the Prefill pool in the latest time period and the cache transmission coefficient of the network link between the Decode pool and the Prefill pool in the latest time period;

[0013] Obtaining the detection interval according to the bandwidth fluctuation coefficient, the cache transmission coefficient, the KV Cache value, and the network link bandwidth;

[0014] Obtaining the detection multiple according to the predetermined fault recovery time and the detection interval; the detection multiple represents the number of unresponses allowed before determining the fault;

[0015] The detection interval and the detection multiple form the network detection parameter.

[0016] In one embodiment, the bandwidth fluctuation coefficient is obtained by:

[0017] In a preset time window, obtaining the network link bandwidth between the Decode pool and the Prefill pool according to a preset period to obtain a network link bandwidth set;

[0018] Obtaining the standard deviation of the network link bandwidth and the average value of the network link bandwidth according to the network link bandwidth set;

[0019] Obtaining the ratio of the standard deviation of the network link bandwidth to the average value of the network link bandwidth to obtain the coefficient of variation;

[0020] Obtaining the bandwidth fluctuation coefficient according to the coefficient of variation.

[0021] In one embodiment, the method further includes:

[0022] Obtaining the current bandwidth change rate of the network link between the Decode pool and the Prefill pool and a plurality of predetermined change rate intervals; the plurality of change rate intervals include a first change rate interval representing that the network link is physically interrupted or seriously faulty, a second change rate interval representing that the network link is network fluctuant, and a third change rate interval representing that the network link is to be observed;

[0023] If the bandwidth change rate is in the first change rate interval, performing route switching according to a pre-constructed route list; the maximum transmittable data packet size under each route in the route list is greater than a preset threshold, and the preset threshold is a number of bytes used to reduce transmission delay;

[0024] If the bandwidth change rate is in the second change rate range, then execute the steps of obtaining the network link bandwidth between the Decode pool and the Prefill pool and the KV Cache value of the Decode pool;

[0025] If the bandwidth change rate is in the third change rate range, the detection result of the bandwidth change rate of the current network link in the future preset detection period is obtained, so as to perform route switching based on the detection result and the pre-built route list.

[0026] In one embodiment, the route list is obtained as follows:

[0027] Identify multiple candidate routing links between the Decode pool and the Prefill pool;

[0028] Obtain the link parameters for each candidate route link; the link parameters include at least one of the following: network link bandwidth parameter, network link delay parameter, network link packet loss rate parameter, and network link maximum transmission unit parameter;

[0029] Based on link parameters, the priority of each candidate route link is determined;

[0030] The candidate routing links are sorted according to priority to obtain a routing list.

[0031] In one embodiment, the priority of each candidate route link is determined based on link parameters, including:

[0032] If there are at least two types of link parameters, determine the weight of each link parameter;

[0033] For each candidate routing link, the priority of the candidate routing link is obtained by weighting the link parameters according to their respective weights.

[0034] In one embodiment, route switching based on a pre-built route list includes:

[0035] Switch the current network link to the highest priority route link in the route list;

[0036] After switching routes based on a pre-built route list, the following is also included:

[0037] The test frame loss rate is obtained by sending test frames to detect the switched routing links.

[0038] If the loss rate is less than the preset loss rate threshold, the route switch is considered successful.

[0039] If the loss rate is greater than or equal to the preset loss rate threshold, a route switch will be performed again, switching the current network link to the next highest priority route link, until the loss rate of the switched route link is less than the preset loss rate threshold.

[0040] Secondly, this application also provides a routing convergence device based on a bidirectional forwarding detection protocol, comprising:

[0041] The acquisition module is used to obtain the network link bandwidth between the Decode pool and the Prefill pool, and the KVCache value of the Decode pool;

[0042] The calculation module is used to calculate network detection parameters based on network link bandwidth, KV Cache value, and predetermined fault recovery time; the fault recovery time is less than the maximum fault recovery time allowed by the Decode pool.

[0043] The detection module is used to detect the network link between the Decode pool and the Prefill pool according to the network detection parameters and obtain the detection results;

[0044] The processing module is used to handle link anomalies if the detection results indicate that the network link is abnormal.

[0045] Thirdly, this application also provides a communication device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described above.

[0046] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.

[0047] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described above.

[0048] The aforementioned routing convergence method, apparatus, communication equipment, computer-readable storage medium, and computer program product based on a bidirectional forwarding detection protocol are applied to a network control device. This network control device manages the interaction between the Decode pool and the Prefill pool in a PD-separated architecture during large-scale model inference. It obtains the network link bandwidth between the Decode pool and the Prefill pool, and the KV Cache value of the Decode pool. Based on the network link bandwidth, KV Cache value, and a predetermined fault recovery time, it calculates network detection parameters. The fault recovery time is less than the maximum allowed fault recovery time for the Decode pool. The network link between the Decode pool and the Prefill pool is detected according to the network detection parameters to obtain the detection result. If the detection result indicates a network link anomaly, link anomaly handling is performed. By incorporating a predetermined fault recovery time into the calculation, the fault recovery time can be significantly reduced, meeting the requirements of the Decode pool. Simultaneously, the network detection parameters are updated in real time using network link bandwidth, KV cache values, and the predetermined fault recovery time. Furthermore, the network link between the Decode pool and the Prefill pool is detected according to the network detection parameters. Based on closed-loop monitoring of the network detection parameters, the network detection parameters can be dynamically adjusted, adapting to changes in the Decode pool's business load. This improves the long-term operational stability of the interaction between the Decode pool and the Prefill pool in the PD separation architecture during large model inference. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0050] Figure 1 This is an application environment diagram of a routing convergence method based on a bidirectional forwarding detection protocol in one embodiment;

[0051] Figure 2 This is a flowchart illustrating a routing convergence method based on a bidirectional forwarding detection protocol in one embodiment.

[0052] Figure 3 This is a schematic diagram illustrating the process of convergence based on a list of routes to be converged, as shown in one embodiment.

[0053] Figure 4 This is a structural block diagram of a routing convergence device based on a bidirectional forwarding detection protocol in one embodiment;

[0054] Figure 5This is an internal structural diagram of a communication device in one embodiment. Detailed Implementation

[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0056] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.

[0057] The routing convergence method based on the bidirectional forwarding detection protocol provided in this application embodiment can be applied to, for example... Figure 1 The application environment shown includes a Decode pool 101, a Prefill pool 102, a network control device 103, and a routing list 104. The network control device 103 obtains the network link bandwidth between the Decode pool and the Prefill pool, and the KVCache value of the Decode pool. Based on the network link bandwidth, KVCache value, and a predetermined fault recovery time, it calculates network detection parameters. If the fault recovery time is less than the maximum allowed fault recovery time of the Decode pool, the network link between the Decode pool and the Prefill pool is detected according to the network detection parameters, and the detection result is obtained. If the detection result indicates a network link anomaly, link anomaly handling is performed.

[0058] In one embodiment, network control device 103 can switch routes according to routing list 104, and then enable Decode pool 101 and Prefill pool 102 to interact through the switched routes.

[0059] Among them, the network control device 103 can be called a server, which can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0060] In one exemplary embodiment, such as Figure 2 As shown, a route convergence method based on a bidirectional forwarding detection protocol is provided, and this method is applied to... Figure 1Taking the network control device 103 as an example, the explanation includes the following steps 201 to 204. Wherein:

[0061] Step 201: Obtain the network link bandwidth between the Decode pool and the Prefill pool, and the KV Cache value of the Decode pool.

[0062] PD separation (also known as Prefill-Decode separation): In large model inference, the Prefill stage is mainly responsible for quickly processing input requests to prepare for subsequent decoding, while the Decode stage is responsible for generating the final output result.

[0063] Decode: Relies on KV Cache to generate output per token. It is memory and bandwidth intensive and is deployed in a separate Decode pool. It needs to interact with the Prefill pool through a high-speed network (which can also be called network equipment or network control equipment).

[0064] Prefill: Generates the initial KV cache, requires a high-performance GPU, and is deployed in the Prefill pool.

[0065] Network link bandwidth refers to the maximum amount of data that can be transmitted per unit of time on the network link between the Prefill pool and the Decode pool.

[0066] KV Cache values ​​are cached data that stores the Key and Value in the model's attention mechanism during AI inference. Their transmission requires GB-level bandwidth and is a core data type that affects inference latency.

[0067] The network control device can also be a network device, and in one embodiment, it can be a switch.

[0068] For example, an acquisition interval can be set, and the network device can then acquire the network link bandwidth between the Decode pool and the Prefill pool, as well as the KV Cache value of the Decode pool, in real time according to the interval. In addition to acquiring the KV Cache value of the Decode pool, it can also acquire real-time TPOT values, RoCEv2 protocol parameters, BGP routing data, etc. The interval can be 10ms or 20ms. In one embodiment, the network device collects the following data every 10ms via the Decode Pool Management API (Application Programming Interface): 1) Service data: KV Cache size; TPOT (Tree-based Pipeline Optimization Tool) real-time value; RoCEv2 protocol parameters, used to confirm large frame configuration, fixed at 9000 bytes; 2) Network data collection: BGP (Border Gateway Protocol) routing data: can be collected via NetFlow v9 (NetFlow Version 9 Protocol), including: route prefix, next-hop address, AS (Autonomous System) path length, bandwidth of each network link, link latency, link packet loss rate, data precision (which can include 0.1GB / s, 1ms, 0.01%); Path Maximum Transmission Unit (MTU). Maximum Transmission Unit (MTU) data: Probe frames are sent via the ICMP extended protocol every 50ms to record the maximum transmittable frame size of the path, eliminating misjudgments of MTU caused by temporary link fluctuations; at the same time, outlier filtering is performed on the collected data and it is stored in a time-series database to provide a data source for subsequent calculations.

[0069] Step 202: Calculate network detection parameters based on network link bandwidth, KV Cache value, and predetermined fault recovery time; the fault recovery time is less than the maximum fault recovery time allowed by the Decode pool.

[0070] The predetermined fault recovery time is the maximum fault recovery time allowed by the Decode pool. When the fault recovery time exceeds the predetermined fault recovery time, a link failure will result in transmission interruption. In one implementation, the predetermined fault recovery time can be 100ms.

[0071] Network detection parameters include detection interval and detection multiplier.

[0072] For example, the detection interval and detection multiplier can be calculated based on the network link bandwidth, KV cache value, and predetermined fault recovery time.

[0073] Step 203: Detect the network link between the Decode pool and the Prefill pool according to the network detection parameters, and obtain the detection results.

[0074] The detection results include three categories: physical network link interruption or severe failure, normal network fluctuations, and the need for continuous monitoring. In one embodiment, detection can be performed by sending detection frames to obtain the detection results.

[0075] For example, detection frames are sent to detect the network link between the Decode pool and the Prefill pool according to the network detection parameters to obtain the detection results.

[0076] Step 204: If the detection result indicates a network link anomaly, then perform link anomaly handling.

[0077] Link anomaly handling includes route switching based on a preset route list.

[0078] For example, if no response is received for a sent detection frame, the network link is determined to be abnormal, and routing is switched according to a preset routing list. When the network detection parameters include a detection interval and a detection multiplier, detection frames are sent according to the detection interval. If no response is received after a certain number of detections, the network link is determined to be abnormal, and routing is switched according to the preset routing list. When the network detection parameters include a detection interval and a detection multiplier, detection frames are sent according to the detection interval. If no response is received after a certain number of detections, the change rate is obtained. If the change rate indicates a physical interruption or severe fault in the network link, routing is switched according to the preset routing list; if the change rate indicates that the network link needs to be detected, the change rate over a period of time is obtained. If all change rates over a period of time indicate that the network link needs to be detected, routing is switched according to the preset routing list. In a real-time example, if the loss rate of sending detection frames through the switched network link is too high, routing is switched according to the preset routing list. It can be set that routing is switched according to the preset routing list when the loss rate is >5%.

[0079] In the aforementioned routing convergence method based on the bidirectional forwarding detection protocol, the introduction of a predetermined fault recovery time into the calculation significantly reduces the fault recovery time, meeting the requirements of the Decode pool. Simultaneously, network detection parameters are updated in real-time using network link bandwidth, KV cache values, and the predetermined fault recovery time. Furthermore, the network link between the Decode pool and the Prefill pool is detected according to these network detection parameters. Closed-loop monitoring based on these parameters enables dynamic adjustment of the network detection parameters, adapting to changes in the Decode pool's service load. This improves the long-term operational stability of the interaction between the Decode pool and the Prefill pool in the PD separation architecture during large-scale model inference.

[0080] In one exemplary embodiment, network detection parameters are calculated based on network link bandwidth, KV cache value, and predetermined fault recovery time, including:

[0081] Get the bandwidth fluctuation coefficient of the network link between the Decode pool and the Prefill pool in the most recent time period and the cache transfer coefficient of the network link between the Decode pool and the Prefill pool in the most recent time period;

[0082] The detection interval is obtained based on the bandwidth fluctuation coefficient, cache transmission coefficient, KV cache value, and network link bandwidth.

[0083] The detection multiple is obtained based on the predetermined fault recovery time and detection interval; the detection multiple represents the number of unresponsive events allowed before a fault is determined.

[0084] The detection interval and detection factor are combined to form the network detection parameters.

[0085] Among them, the bandwidth fluctuation coefficient is a quantitative indicator used to measure the degree of fluctuation of network bandwidth per unit time.

[0086] The cache transfer coefficient is an indicator used to measure the data transfer efficiency of a cache system.

[0087] The detection interval can be obtained by substituting the bandwidth fluctuation coefficient, buffer transmission coefficient, KV cache value, and network link bandwidth into the detection interval calculation formula. The detection interval calculation formula is as follows:

[0088]

[0089] in, For the detection interval, KV Cache size (GB) This refers to the link bandwidth (GB / s). This is a buffer transmission coefficient (with a value of 1.2-1.5 to adapt to the RoCEv2 protocol overhead). This is the bandwidth fluctuation coefficient (with a value of 0.8-1.2).

[0090] The predetermined fault recovery time and detection interval can be substituted into the detection multiple calculation formula to obtain the detection multiple. The detection multiple formula is as follows:

[0091]

[0092] in, To detect multiples, The maximum allowable fault recovery time is ≤100ms. The function is a floor function; it automatically decreases when the KV cache increases to improve detection sensitivity, and adjusts for bandwidth fluctuations. Adjustments were made to reduce the false alarm rate.

[0093] In one exemplary embodiment, the bandwidth fluctuation coefficient is obtained in the following manner:

[0094] Within a preset time window, the network link bandwidth between the Decode pool and the Prefill pool is obtained according to a preset period to obtain a set of network link bandwidths.

[0095] Calculate the standard deviation and average value of network link bandwidth based on the network link bandwidth set;

[0096] The coefficient of variation is obtained by comparing the standard deviation of network link bandwidth with the average network link bandwidth.

[0097] The bandwidth fluctuation coefficient is obtained based on the coefficient of variation.

[0098] The standard deviation characterizes the degree of dispersion (or fluctuation) of the network link bandwidth set from the average value.

[0099] The average value is calculated by dividing the sum of all data in the network link bandwidth set by the number of data points, reflecting the average level of this set of data.

[0100] The coefficient of variation (CV) is a relative fluctuation indicator of network link bandwidth. It is calculated by dividing the standard deviation by the mean. Its core value is to eliminate the influence of bandwidth magnitude. Standardization measures the degree of bandwidth fluctuation and is more meaningful for comparison than looking at the standard deviation alone.

[0101] The preset time window can be 100ms or 200ms, and the corresponding preset period can be 10ms. The preset period must be shorter than the preset time window. In one embodiment, the preset time window is 100ms and the preset period is 10ms. Then, the network link bandwidth between 10 Decode pools and Prefill pools can be obtained, and the network link bandwidth between these 10 Decode pools and Prefill pools can be combined into a network link bandwidth set.

[0102] The formula for obtaining the coefficient of variation is:

[0103]

[0104] The formula for obtaining the bandwidth fluctuation coefficient is:

[0105]

[0106] In one exemplary embodiment, the route list is obtained as follows:

[0107] Identify multiple candidate routing links between the Decode pool and the Prefill pool;

[0108] Obtain the link parameters for each candidate route link; the link parameters include at least one of the following: network link bandwidth parameter, network link delay parameter, network link packet loss rate parameter, and network link maximum transmission unit parameter;

[0109] Based on link parameters, the priority of each candidate route link is determined;

[0110] The candidate routing links are sorted according to priority to obtain a routing list.

[0111] Among them, the network link bandwidth parameter includes the actual network link bandwidth of the candidate route link at the current moment and the maximum network link bandwidth of the candidate route link.

[0112] Network link latency (or delay) is a core parameter for measuring network link response speed. It refers to the total time from when data is sent from the sender to when it is fully received by the receiver. The lower the value, the more sensitive the link response, making it a key quality indicator for scenarios such as real-time communication and online interaction. Network link latency includes the actual network link latency of the candidate route link at the current moment and the maximum network link latency of the candidate route link.

[0113] Network link packet loss rate is a core parameter for measuring the reliability of network link data transmission. It refers to the proportion of data packets lost during transmission out of the total number of data packets sent. The network link packet loss rate parameter includes the current actual network link packet loss rate of the candidate route link and the maximum packet loss rate of the candidate route link.

[0114] The maximum transmission unit (MTU) parameter of a network link refers to the maximum size of a data packet (in bytes) that a network link can transmit. When the RoCEv2 protocol transmits KV cache, an MTU ≥ 9000 bytes is required to avoid fragmentation.

[0115] In one embodiment, the priority of each candidate route link can be obtained by setting different weights for link parameters and weighting the link parameters according to the weight of each link parameter. Then, the links are sorted in descending order of priority to obtain a route list.

[0116] In one exemplary embodiment, the priority of each candidate routing link is determined based on link parameters, including:

[0117] If there are at least two types of link parameters, determine the weight of each link parameter;

[0118] For each candidate routing link, the priority of the candidate routing link is obtained by weighting the link parameters according to their respective weights.

[0119] The sum of the weights of each link parameter is 1.

[0120] For example, when the link parameters include network link bandwidth, network link delay, network link packet loss rate, and network link maximum transmission unit (MTBF) parameters, the weights of each link parameter satisfy the following:

[0121]

[0122] in, Indicates the weight of the maximum transmission unit (MTU) of the network link. In this application, the MTU of each route in the routing list must be ≥9000 bytes. Since the MTU size can reduce routing latency, the MTU weight here will be relatively large. Indicates the weight of the network link delay parameter. Indicates the network link bandwidth weight. This represents the weight of the network link packet loss rate.

[0123] This is the weighting factor for the network link MTU, with its lower limit set at 0.4 (i.e., ...). The MTU priority (≥0.4) is essentially determined by the core priority of MTU in low-latency, high-reliability real-time transmission scenarios.

[0124] In one embodiment, the KV Cache transmission latency between TPOT and the Decode pool is monitored in real time. If TPOT latency is ≥50ms or the transmission latency increases by ≥30%, the BFD parameters (also known as network detection parameters) and the priority of candidate route links are recalculated. Simultaneously, the weighting coefficients can be optimized daily based on historical data. , , , .

[0125] For example, the priority of each link parameter can be obtained by weighting it according to its weight and substituting it into the formula:

[0126] ,

[0127] in, The actual MTU (in bytes) of the network link. , This represents the actual network link delay (ms). For maximum network link latency, This represents the actual network link bandwidth (GB / s). For maximum network link bandwidth, This represents the actual packet loss rate of the network link. This represents the maximum packet loss rate.

[0128] In one exemplary embodiment, route switching based on a pre-built route list includes:

[0129] Switch the current network link to the highest priority route link in the route list;

[0130] After switching routes based on a pre-built route list, the following is also included:

[0131] The test frame loss rate is obtained by sending test frames to detect the switched routing links.

[0132] If the loss rate is less than the preset loss rate threshold, the route switch is considered successful.

[0133] If the loss rate is greater than or equal to the preset loss rate threshold, a route switch will be performed again, switching the current network link to the next highest priority route link, until the loss rate of the switched route link is less than the preset loss rate threshold.

[0134] Sending test frames can refer to MTU probing, or it can refer to sending 9000 bytes of RoCEv2 probe frames each time and then obtaining the test frame loss rate.

[0135] The packet loss rate threshold can be 5%, which means that for every 100 packets sent, no more than 5 will be lost. When the loss rate is less than 5%, it indicates that the Decode pool and the Prefill pool can communicate normally through the switched route, but the communication quality cannot be guaranteed. This is suitable for scenarios where packet loss is not a major concern. If communication quality needs to be guaranteed (e.g., in real-time communication scenarios), the loss rate needs to be reduced to 0.5% or 0.1%.

[0136] For example, the highest priority route is designated as the primary route, and the other routes as backup routes, resulting in a route list. The current network link is then switched to the highest priority route in the list. When switching from the primary route to the highest priority backup route, a route update message is sent via BGP. The primary route is then directly switched to the highest priority backup route, bypassing the traditional BGP AS path comparison and route priority election process. A 9000-byte RoCEv2 test frame is then sent to obtain the test frame loss rate. If the loss rate is less than 5%, the MTU representing the highest priority backup route matches the current communication between the Decode and Prefill pools, and the switch is successful. If the loss rate is greater than 5%, the MTU representing the highest priority backup route does not match the current communication between the Decode and Prefill pools. The highest priority backup route is then switched to the second highest priority backup route, and another 9000-byte RoCEv2 test frame is sent to obtain the test frame loss rate. This process continues until the loss rate of the switched route is less than 5%.

[0137] In one exemplary embodiment, such as Figure 3 As shown, the convergence process based on the list of routes to be converged includes steps 301 to 304. Wherein:

[0138] Step 301: Obtain the current bandwidth change rate and multiple predetermined change rate intervals of the network link between the Decode pool and the Prefill pool; the multiple change rate intervals include a first change rate interval representing a physical interruption or severe failure of the network link, a second change rate interval representing network fluctuations of the network link, and a third change rate interval representing a network link to be observed.

[0139] Among them, bandwidth change rate This is a quantitative indicator that measures how quickly network bandwidth changes over time. Its core function is to reflect the increase or decrease in bandwidth per unit of time, used to determine if there are sudden bandwidth fluctuations. It is a key indicator for supplementing the bandwidth fluctuation coefficient and ensuring the stability of real-time transmission. Here, it measures the bandwidth descent rate. Several predetermined rate-of-change intervals are included, including the first rate-of-change interval: This indicates a bandwidth drop of over 90%, signifying a physical interruption or severe failure of the network link. Under normal circumstances, bandwidth cannot drop by more than 90%. A bandwidth drop of over 90% means the link's transmission capacity is essentially lost, consistent with the nature of a physical interruption / severe failure; Second rate of change range: This indicates a bandwidth drop of less than 30%, signifying network fluctuations. Setting the limit to 30% can filter out normal fluctuations and prevent excessive response. The third rate of change range: This indicates a bandwidth drop between 30% and 90%, signifying that the network link is under observation. When the bandwidth drop exceeds 30%, the available bandwidth suddenly decreases, large data packet transmissions become queued, and KV Cache transmission latency fluctuates and then drops below 90%, indicating that the bandwidth is available. Therefore, a bandwidth drop between 30% and 90% is considered under observation.

[0140] Step 302: If the bandwidth change rate is within the first change rate range, then perform route switching according to the pre-built route list; the maximum data packet size that can be transmitted under each route in the route list is greater than a preset threshold, which is the number of bytes used to reduce transmission latency.

[0141] The preset threshold is 9000 bytes. For the RoCEv2 protocol, a large 9000-byte frame is needed to reduce packet fragmentation during communication between the Decode pool and the Prefill pool, thus reducing transmission latency. However, this can lead to the model being unable to acquire data simultaneously during inference, resulting in semantic inconsistencies. Furthermore, after confirming a route switch, a 9000-byte detection frame needs to be sent, and the adaptation status is confirmed based on the packet loss rate of the detection frame. Therefore, a preset threshold of 9000 bytes is optimal. Consequently, the maximum transmittable packet size for each route in the route list must be greater than 9000 bytes.

[0142] For example, if If the network link is physically interrupted or severely faulty, then a route switch is performed based on a pre-built route list.

[0143] Step 303: If the bandwidth change rate is within the second change rate range, then execute the steps of obtaining the network link bandwidth between the Decode pool and the Prefill pool and the KV Cache value of the Decode pool.

[0144] For example, if This indicates that bandwidth fluctuations within 30% are considered normal network fluctuations. Instead of switching routes, the process proceeds to obtain the network link bandwidth between the Decode pool and the Prefill pool, as well as the KVCache value of the Decode pool. The latest detection interval and detection multiplier are then calculated to improve tolerance for network bandwidth fluctuations and avoid over-response.

[0145] Step 304: If the bandwidth change rate is in the third change rate range, the detection result of the bandwidth change rate of the current network link in the future preset detection period is obtained, so as to perform route switching based on the detection result and the pre-built route list.

[0146] For example, if Then, the detection result of the bandwidth change rate of the current network link in the next three periods (which can be 30ms or 5 periods: 50ms) is used. If the detection result of all three periods is not found, the result is considered invalid. In this case, it is treated as a link interruption (indicating a physical break or severe failure of the network link), and route switching is performed according to a pre-built routing list. If the results of the three detections show... If the situation is as described above, it indicates that the bandwidth has been restored. The steps to obtain the network link bandwidth between the Decode pool and the Prefill pool and the KVCache value of the Decode pool are executed, and then the latest detection interval and detection multiplier are calculated.

[0147] In this embodiment, by determining whether to converge based on the list of routes to be converged based on the bandwidth change rate being within a certain range, it is possible to achieve second-level switching of qualified routes without service interruption; the link adaptability is continuously improved, supporting the highly reliable and low-latency operation of large model inference tasks.

[0148] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.

[0149] Based on the same inventive concept, this application also provides a routing convergence device based on the bidirectional forwarding detection protocol for implementing the routing convergence method based on the bidirectional forwarding detection protocol described above. The solution provided by this device is similar to the implementation described in the above method. Therefore, the specific limitations of one or more embodiments of the routing convergence device based on the bidirectional forwarding detection protocol provided below can be found in the limitations of the routing convergence method based on the bidirectional forwarding detection protocol described above, and will not be repeated here.

[0150] In one exemplary embodiment, such as Figure 4 As shown, a routing convergence device based on a bidirectional forwarding detection protocol is provided, comprising: an acquisition module 401, a calculation module 402, a detection module 403, and a processing module 404, wherein:

[0151] The acquisition module 401 is used to acquire the network link bandwidth between the Decode pool and the Prefill pool and the KV Cache value of the Decode pool;

[0152] The calculation module 402 is used to calculate network detection parameters based on network link bandwidth, KV Cache value, and predetermined fault recovery time; the fault recovery time is less than the maximum fault recovery time allowed by the Decode pool.

[0153] The detection module 403 is used to detect the network link between the Decode pool and the Prefill pool according to the network detection parameters and obtain the detection results;

[0154] The processing module 404 is used to perform link anomaly processing if the detection result indicates that the network link is abnormal.

[0155] In an optional embodiment, the calculation module 402 is further configured to:

[0156] Get the bandwidth fluctuation coefficient of the network link between the Decode pool and the Prefill pool in the most recent time period and the cache transfer coefficient of the network link between the Decode pool and the Prefill pool in the most recent time period;

[0157] The detection interval is obtained based on the bandwidth fluctuation coefficient, cache transmission coefficient, KV cache value, and network link bandwidth.

[0158] The detection multiple is obtained based on the predetermined fault recovery time and detection interval; the detection multiple represents the number of unresponsive events allowed before a fault is determined.

[0159] The detection interval and detection factor are combined to form the network detection parameters.

[0160] In an optional embodiment, the calculation module 402 is further configured to:

[0161] Within a preset time window, the network link bandwidth between the Decode pool and the Prefill pool is obtained according to a preset period to obtain a set of network link bandwidths.

[0162] Calculate the standard deviation and average value of network link bandwidth based on the network link bandwidth set;

[0163] The coefficient of variation is obtained by comparing the standard deviation of network link bandwidth with the average network link bandwidth.

[0164] The bandwidth fluctuation coefficient is obtained based on the coefficient of variation.

[0165] In an optional embodiment, the calculation module 402 is further configured to:

[0166] Obtain the current bandwidth change rate and multiple predetermined change rate intervals of the network link between the Decode pool and the Prefill pool; the multiple change rate intervals include a first change rate interval characterizing the network link as a physical interruption or severe failure, a second change rate interval characterizing the network link as a network fluctuation, and a third change rate interval characterizing the network link as to be observed.

[0167] If the bandwidth change rate is within the first change rate range, then route switching is performed according to the pre-built route list; the maximum data packet size that can be transmitted under each route in the route list is greater than a preset threshold, which is the number of bytes used to reduce transmission latency;

[0168] If the bandwidth change rate is in the second change rate range, then execute the steps of obtaining the network link bandwidth between the Decode pool and the Prefill pool and the KV Cache value of the Decode pool;

[0169] If the bandwidth change rate is in the third change rate range, the detection result of the bandwidth change rate of the current network link in the future preset detection period is obtained, so as to perform route switching based on the detection result and the pre-built route list.

[0170] In an optional embodiment, the processing module 404 is further configured to:

[0171] Identify multiple candidate routing links between the Decode pool and the Prefill pool;

[0172] Obtain the link parameters for each candidate route link; the link parameters include at least one of the following: network link bandwidth parameter, network link delay parameter, network link packet loss rate parameter, and network link maximum transmission unit parameter;

[0173] Based on link parameters, the priority of each candidate route link is determined;

[0174] The candidate routing links are sorted according to priority to obtain a routing list.

[0175] In an optional embodiment, the processing module 404 is further configured to:

[0176] If there are at least two types of link parameters, determine the weight of each link parameter;

[0177] For each candidate routing link, the priority of the candidate routing link is obtained by weighting the link parameters according to their respective weights.

[0178] In an optional embodiment, the processing module 404 is further configured to:

[0179] Switch the current network link to the highest priority route link in the route list;

[0180] After switching routes based on a pre-built route list, the following is also included:

[0181] The test frame loss rate is obtained by sending test frames to detect the switched routing links.

[0182] If the loss rate is less than the preset loss rate threshold, the route switch is considered successful.

[0183] If the loss rate is greater than or equal to the preset loss rate threshold, a route switch will be performed again, switching the current network link to the next highest priority route link, until the loss rate of the switched route link is less than the preset loss rate threshold.

[0184] Each module in the aforementioned routing convergence device based on the bidirectional forwarding detection protocol can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the communication device in hardware form or independent of it, or stored in the memory of the communication device in software form, so that the processor can call and execute the corresponding operations of each module.

[0185] In one exemplary embodiment, a communication device is provided, which may be a server, and its internal structure diagram may be as follows. Figure 5As shown, the communication device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores routing list data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When executed by the processor, the computer program implements a routing convergence method based on a bidirectional forwarding detection protocol.

[0186] Those skilled in the art will understand that Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the communication device to which the present application is applied. Specific communication devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0187] In one exemplary embodiment, a communication device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the routing convergence based on the bidirectional forwarding detection protocol of the above embodiments.

[0188] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the routing convergence based on the bidirectional forwarding detection protocol of the above embodiments.

[0189] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0190] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0191] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0192] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A routing convergence method based on a bidirectional forwarding detection protocol, characterized in that, An application to a network control device, the network control device being used to manage the interaction between the Decode pool and the Prefill pool under a PD separation architecture in large model inference, the method comprising: Get the network link bandwidth between the Decode pool and the Prefill pool, and the KV cache value of the Decode pool; The network detection parameters are calculated based on the network link bandwidth, the KV Cache value, and the predetermined fault recovery time; the fault recovery time is less than the maximum fault recovery time allowed by the Decode pool. The network link between the Decode pool and the Prefill pool is detected according to the network detection parameters, and the detection results are obtained. If the detection result indicates a network link anomaly, then link anomaly handling is performed.

2. The method according to claim 1, characterized in that, The calculation of network detection parameters based on the network link bandwidth, the KVCache value, and the predetermined fault recovery time includes: Obtain the bandwidth fluctuation coefficient of the network link between the Decode pool and the Prefill pool in the most recent time period and the cache transfer coefficient of the network link between the Decode pool and the Prefill pool in the most recent time period; The detection interval is obtained based on the bandwidth fluctuation coefficient, the cache transmission coefficient, the KV Cache value, and the network link bandwidth. The detection multiple is obtained based on the predetermined fault recovery time and the detection interval; the detection multiple represents the number of unresponsive events allowed before a fault is determined. The network detection parameters are composed of the detection interval and the detection multiplier.

3. The method according to claim 2, characterized in that, The bandwidth fluctuation coefficient is obtained in the following way: Within a preset time window, the network link bandwidth between the Decode pool and the Prefill pool is obtained according to a preset period to obtain a set of network link bandwidths. Based on the network link bandwidth set, calculate the standard deviation and average value of the network link bandwidth; The coefficient of variation is obtained by comparing the standard deviation of the network link bandwidth with the average value of the network link bandwidth. The bandwidth fluctuation coefficient is obtained based on the coefficient of variation.

4. The method according to claim 1, characterized in that, The method includes: Obtain the current bandwidth change rate and a predetermined number of change rate intervals for the network link between the Decode pool and the Prefill pool; the number of change rate intervals include a first change rate interval characterizing the network link as physically interrupted or severely faulty, a second change rate interval characterizing the network link as experiencing network fluctuations, and a third change rate interval characterizing the network link as to be observed; If the bandwidth change rate is within the first change rate range, then route switching is performed according to the pre-built route list; the maximum data packet size that can be transmitted under each route in the route list is greater than a preset threshold, the preset threshold being the number of bytes used to reduce transmission latency; If the bandwidth change rate is within the second change rate range, then the steps of obtaining the network link bandwidth between the Decode pool and the Prefill pool and the KV Cache value of the Decode pool are executed. If the bandwidth change rate is within the third change rate range, the detection result of the bandwidth change rate of the current network link in the future preset detection period is obtained, so as to perform route switching based on the detection result and the pre-built route list.

5. The method according to claim 4, characterized in that, The route list is obtained in the following way: Identify multiple candidate routing links between the Decode pool and the Prefill pool; Obtain the link parameters for each candidate route link; the link parameters include at least one of the following: network link bandwidth parameter, network link delay parameter, network link packet loss rate parameter, and network link maximum transmission unit parameter; Based on the link parameters, the priority of each candidate route link is determined; The candidate routing links are sorted according to the stated priority to obtain the routing list.

6. The method according to claim 5, characterized in that, The step of determining the priority of each candidate route link based on the link parameters includes: If there are at least two types of link parameters, determine the weight of each link parameter; For each candidate routing link, the priority of the candidate routing link is obtained by weighting the link parameters according to their respective weights.

7. The method according to any one of claims 4-6, characterized in that, The route switching based on the pre-built route list includes: Switch the current network link to the highest priority route link in the route list; After performing route switching based on a pre-built route list, the process also includes: The test frame loss rate is obtained by sending test frames to detect the switched routing links. If the loss rate is less than the preset loss rate threshold, the route switch is determined to be successful. If the loss rate is greater than or equal to the preset loss rate threshold, then another route switch is performed, switching the current network link to the next highest priority route link, until the loss rate of the switched route link is less than the preset loss rate threshold.

8. A routing convergence device based on a bidirectional forwarding detection protocol, characterized in that, The device includes: The acquisition module is used to obtain the network link bandwidth between the Decode pool and the Prefill pool, and the KVCache value of the Decode pool; The calculation module is used to calculate network detection parameters based on the network link bandwidth, the KV Cache value, and a predetermined fault recovery time; the fault recovery time is less than the maximum fault recovery time allowed by the Decode pool. The detection module is used to detect the network link between the Decode pool and the Prefill pool according to the network detection parameters, and obtain the detection result; The processing module is used to perform link anomaly processing if the detection result indicates a network link anomaly.

9. A communication device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.