A network resource scheduling method, device and equipment

CN121547396BActive Publication Date: 2026-09-29INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610064795.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-19
Publication Date
2026-09-29
Estimated Expiration
2046-01-19

AI Technical Summary

Technical Problem

[0004]本申请提供了一种网络资源调度方法、装置和设备,以至少解决相关技术中RoCE网络资源闲置与TCP/IP网络高负载性能不足的问题

Benefits of technology

[0008]通过本申请,通过构建融合资源池,消除了RoCE与TCP/IP网络资源的隔离状态;进而通过基于业务传输需求特征确定主传输协议类型,实现了业务需求与网络协议特性的精准匹配;最后通过从融合池中动态分配对应链路资源,完成了跨协议资源的统一调度与共享。解决了RoCE网络资源闲置与TCP/IP网络高负载性能不足的问题,达到了提升数据中心网络整体资源利用率与业务性能的有益效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121547396B_ABST
    Figure CN121547396B_ABST
Patent Text Reader

Abstract

The application discloses a network resource scheduling method, device and equipment, relates to the technical field of network resources, and comprises the following steps: constructing a fusion resource pool, eliminating the isolation state of RoCE and TCP / IP network resources, determining a main transmission protocol type based on service transmission demand characteristics, realizing accurate matching of service demand and network protocol characteristics, and finally completing unified scheduling and sharing of cross-protocol resources by dynamically allocating corresponding link resources from the fusion pool. The problems of idle RoCE network resources and insufficient high-load performance of TCP / IP network are solved, and the overall resource utilization and service performance of the data center network are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of network resource technology, and in particular to a network resource scheduling method, apparatus and device. Background Technology

[0002] In current data center network architectures, RoCE (Remote Direct Memory Access over Converged Ethernet), with its low latency, high bandwidth, and zero-copy characteristics, has become the preferred transmission solution for high-performance services (such as distributed databases and AI (Artificial Intelligence) training). Meanwhile, the highly compatible and mature TCP / IP (Transmission Control Protocol / Internet Protocol) network widely serves general-purpose services such as file transfer and web services. However, RoCE and TCP / IP networks are typically deployed and managed independently, resulting in the isolation of their physical ports, bandwidth, queues, and other hardware resources, hindering unified scheduling and sharing across protocols. This not only leads to idle and wasted resources in the RoCE network during off-peak periods but also makes it difficult for the TCP / IP network to leverage idle RoCE resources to improve performance under high load, creating a dual bottleneck of resource utilization and service performance. Furthermore, existing services and network resources often employ static binding or simple policy routing, lacking adaptive scheduling mechanisms based on service characteristics, making it difficult to achieve global resource optimization and performance guarantees.

[0003] Therefore, how to provide a solution to the above-mentioned technical problems is a problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] This application provides a network resource scheduling method, apparatus, and device to at least solve the problems of idle RoCE network resources and insufficient performance of TCP / IP networks under high load in related technologies.

[0005] This application provides a network resource scheduling method, comprising: responding to a service access request, determining the main transmission protocol type of the service based on the transmission requirement characteristics of the service; allocating main link resources corresponding to the main transmission protocol type to the service in a converged resource pool, wherein the converged resource pool stores link resources of at least two transmission protocol types; and establishing a network connection based on the allocated main link resources to perform the transmission of service data.

[0006] This application also provides a network resource scheduling device, comprising: a policy determination module, configured to determine the main transmission protocol type of the service based on the transmission requirement characteristics of the service in response to a service access request; a resource allocation module, configured to allocate main link resources corresponding to the main transmission protocol type to the service in a converged resource pool, wherein the converged resource pool stores link resources of at least two transmission protocol types; and a transmission execution module, configured to establish a network connection based on the allocated main link resources to execute the transmission of service data.

[0007] This application also provides a network device, comprising: a communication unit for receiving service access requests and performing service data transmission based on an established network connection; a storage unit for storing a fused resource pool containing link resources of at least two transmission protocol types; and a processing unit connected to the communication unit and the storage unit, respectively, for: responding to a service access request received by the communication unit, determining the primary transmission protocol type of the service based on the transmission requirement characteristics of the service; allocating primary link resources corresponding to the primary transmission protocol type for the service from the fused resource pool of the storage unit; and controlling the communication unit to establish a network connection based on the allocated primary link resources to perform service data transmission.

[0008] This application eliminates the isolation between RoCE and TCP / IP network resources by constructing a converged resource pool; furthermore, it achieves precise matching between business requirements and network protocol characteristics by determining the primary transport protocol type based on business transmission demand characteristics; finally, it completes unified scheduling and sharing of cross-protocol resources by dynamically allocating corresponding link resources from the converged pool. This solves the problems of idle RoCE network resources and insufficient performance of TCP / IP networks under high load, achieving the beneficial effect of improving the overall resource utilization and business performance of the data center network. Attached Figure Description

[0009] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 A flowchart illustrating the steps of a network resource scheduling method provided in this application embodiment.

[0011] Figure 2 This is a schematic diagram of the structure of a network resource scheduling system provided in an embodiment of this application.

[0012] Figure 3A flowchart illustrating the steps of another network resource scheduling method provided in this application embodiment.

[0013] Figure 4 This is a schematic diagram of the structure of a network resource scheduling device provided in an embodiment of this application. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0015] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0016] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0017] Please refer to Figure 1 This application provides a network resource scheduling method, which can be executed by a network device, including but not limited to network interface cards (NICs) and switches. The method is described in detail below, based on its execution flow, including: S101: In response to a service access request, determining the main transmission protocol type of the service based on its transmission requirement characteristics.

[0018] First, it should be noted that this embodiment designs a unified transmission interface, unified_transport_api, which includes the following core functions: ut_init(): Initializes the interface, loads the RoCE and TCP / IP protocol stack drivers, and establishes communication with the converged resource pool; ut_connect(dst_ip, dst_port, transport_type): Establishes a connection, where transport_type specifies RoCE priority, TCP / IP priority, or automatic selection; ut_send(buf, len, flags): Sends data, where flags include reliable transmission flags, low latency flags, high throughput flags, or other service flags; ut_recv(buf, len, flags): Receives data; ut_close(): Closes the connection and releases resources. After receiving a service access request, the function calls of unified_transport_api are dynamically mapped to the corresponding Verbs interface or Socket interface calls according to the preset mapping rules and the current network status.

[0019] For example, when the parameter of transport_type is RoCE priority and the RoCE link is normal, ut_connect is mapped to ibv_create_qp (create RoCE hardware queue), and ut_send is mapped to ibv_post_send; when the RoCE link fails or the parameter of transport_type is TCP / IP priority, ut_connect is mapped to socket() and connect(), and ut_send is mapped to send(); the above mapping rules are stored in a dynamically updatable interface mapping table, ensuring the flexibility and reliability of the calls.

[0020] Specifically, a service access request refers to a request triggered when an upper-layer application initiates a network transmission service application by calling a unified application programming interface (e.g., the `ut_connect` function). Upon receiving this service access request, the transmission requirement characteristics of the service can be obtained by analyzing the `flags` parameters carried in the `ut_send` / `ut_recv` calls or by collecting real-time behavior data during the initial data transmission phase. Transmission requirement characteristics refer to the performance requirements and behavioral patterns exhibited by the service during data transmission. Selecting a matching transmission protocol type for different transmission requirement characteristics can optimize transmission performance. In this embodiment, the primary transmission protocol type refers to the network transmission protocol that should be prioritized for the current service, used to achieve better performance in data transmission, such as lower latency, higher throughput, or better CPU efficiency.

[0021] For ease of understanding, this application uses two transmission protocol types as examples to describe the network resource scheduling method in detail. These two transmission protocol types include RoCE and TCP / IP; however, this method is not limited to these two protocols. This step precisely maps business-level requirements to network protocol selection, providing a decision-making basis for subsequent resource matching.

[0022] It is understandable that when a service initiates a connection or begins data transmission, its transmission requirements are captured by analyzing interface parameters or through real-time monitoring. Since different transmission protocols (such as RoCE and TCP / IP) have different focuses in terms of latency, throughput, and CPU overhead, there is a correspondence between transmission requirements and the optimal transmission protocol type. This step provides a solution to the problems of poor compatibility and difficulty in leveraging the hardware acceleration capabilities of RoCE network cards in related technologies through service characteristic perception and protocol decision-making. Protocol selection in related technologies often relies on manual static configuration, which can easily lead to the high-performance RoCE protocol being used for services that do not require its features, thus causing compatibility issues, while its hardware acceleration potential is not realized in truly needed scenarios. This step achieves dynamic and precise matching between service requirements and transmission protocols, ensuring that services with high latency and throughput requirements receive priority access to RoCE hardware acceleration support, while guiding services with massive small connections and high compatibility requirements, which are more suitable for traditional protocol stacks, to TCP / IP. This effectively avoids resource mismatch and compatibility conflicts while improving overall performance.

[0023] S102: In the converged resource pool, allocate main link resources corresponding to the main transmission protocol type for the service. The converged resource pool stores link resources of at least two transmission protocol types.

[0024] Prior to this step, the process includes constructing a converged resource pool. It's understood that the purpose of constructing the converged resource pool is to integrate link resources from different transport protocols, achieving unified management and efficient scheduling. It's not simply about partitioned storage, but rather about achieving deep integration and flexible access to multi-protocol resources through resource abstraction, dynamic allocation mechanisms, and protocol co-optimization. Specifically, it integrates heterogeneous resources such as RoCE hardware queues, ports, and TCP / IP sockets and connection handles through abstraction. In this embodiment, the design of at least two transport protocol types clarifies the heterogeneous characteristics of the converged resource pool, providing diverse transmission path options for services. The main link resource is a specific resource instance, such as a RoCE queue or a TCP / IP socket, allocated from the converged resource pool based on the decision results.

[0025] When constructing the integrated resource pool, a combination of proactive detection and passive reporting is used to dynamically monitor resource status, and a resource status database is dynamically maintained. Proactive detection includes periodically collecting metrics such as bandwidth, latency, and utilization of links of various transmission protocol types. Considering the characteristics of transmission protocols, different collection periods are used for RoCE resources and TCP / IP resources. Specifically, since the RoCE protocol focuses on low-latency, high-bandwidth data transmission and is suitable for scenarios with extremely high network transmission performance requirements, the collection period for RoCE resource parameters can be set to be shorter to more promptly detect changes in network status and ensure low-latency transmission requirements. The TCP / IP protocol stack, on the other hand, is widely used and performs well in terms of versatility and stability. Its resource parameter collection period can be flexibly adjusted according to business scenarios; in scenarios where real-time requirements are not high, the collection period can be set to be relatively longer.

[0026] Furthermore, when the RoCE link is interrupted, the number of data packet retransmissions exceeds the threshold, the bandwidth utilization drops sharply, or there is an abnormal state, a passive reporting mechanism is triggered. The fault information is encapsulated into a data packet of a specific format and reported through the management channel so that the information can be parsed based on the reported data and the resource pool status table can be updated to mark the faulty network card. The reported data contains detailed fault diagnosis information.

[0027] For resource parameters collected from different transmission protocol types, a partitioned time-series strategy can be used for storage. First, the storage resource pool is divided into multiple partitions according to the number of transmission protocol types, with each partition corresponding to and storing resource parameters for one transmission protocol type. For example, assuming the transmission protocol types include RoCE and TCP / IP, two partitions are created in the merged resource pool: partition A1 and partition A2. RoCE resource parameters are stored in partition A1, and TCP / IP resource parameters are stored in partition A2.

[0028] Different transport protocol types can correspond to different storage durations for their resource parameters. For example, for RoCE resources, which are suitable for low-latency scenarios, the data storage period for the first partition A1 can be set to 1-5 minutes for rolling storage, retaining key operational metrics (such as link bandwidth and latency jitter) to quickly respond to failover requirements. For TCP / IP resources, since TCP / IP emphasizes stable transmission, the data storage period for the second partition A2 can be extended to 30 minutes to 1 hour, recording complete session connection and throughput data. Data updates in each partition can employ cyclic overwrite storage, automatically overwriting historical data after the preset storage period, or combined with an incremental update mechanism, only writing to the changed resource status in real time, reducing database load and ensuring efficient storage and fast retrieval of RoCE and TCP / IP resource status data.

[0029] It is understandable that in the converged resource pool, each transmission protocol type may contain multiple available link resources. For example, under the RoCE transmission protocol type, there may be a first link resource R1, a second link resource R2, and a third link resource R3. R1, R2, and R3 may correspond to different channels of the same port. When resources need to be allocated for a service, the most suitable resource is selected from the resource set of the corresponding protocol type based on the specific transmission requirements and strategies of the service. This resource is then designated as the primary link resource for that service, and a corresponding resource allocation record is generated.

[0030] This step addresses the issues of inconsistent scheduling and uneven resource utilization of heterogeneous transmission protocol resources by constructing a unified, integrated resource pool and implementing a refined resource allocation strategy. In related technologies, RoCE and TCP / IP protocols are isolated and managed independently, leading to wasted high-performance resources during idle periods and insufficient general-purpose resources to meet business demands under high loads. This embodiment constructs a dynamically aware, multi-protocol integrated resource pool, achieving a mapping from abstract protocol decisions to specific resource instances. This allows services to obtain the most suitable link resources based on their transmission requirements. This not only improves the utilization of high-performance hardware resources (such as RoCE network cards) but also enables cross-protocol load balancing based on global resource status, achieving an optimal balance between transmission performance and resource efficiency at the system level. Furthermore, this integrated resource pool architecture has good scalability and can flexibly adapt to more types of transmission protocols.

[0031] S103: Establish a network connection based on the allocated primary link resources to perform the transmission of service data.

[0032] In this embodiment, a network connection channel is established between the service and the allocated main link resources. Then, through this channel, the corresponding underlying hardware or system functions are called to complete the sending and receiving of service data.

[0033] Specifically, the sending process is as follows: Encapsulated business data is read from the data buffer, and the corresponding hardware interface is called based on the main transport protocol type. For example, if the main transport protocol type is RoCE, then ibv_post_send is called; if it is TCP / IP, then rte_eth_tx_burst (the DPDK send function) or send() is called. After sending is complete, the data buffer status is updated, and it is marked as sent.

[0034] The receiving process is as follows: the data received by the network card is obtained through hardware interrupt or polling (DPDK mode), the transmission control header is decapsulated, the checksum is verified, the data is written to the buffer and marked as ready to be read, a data ready notification is sent to the service layer, and finally the data is delivered through the ut_recv callback function.

[0035] During connection establishment and transmission, fine-grained control is exercised over underlying network resources, including: Link configuration: configuring the RoCE port's speed and duplex mode via ibv_query_port / ibv_modify_port; configuring the TCP / IP port's MTU and link negotiation mode via rte_eth_link_set_up; Queue management: creating / destroying RoCE hardware queues (ibv_create_qp / ibv_destroy_qp), configuring queue depth and priority; managing TCP / IP Socket queues; Interrupt control: enabling / disabling network card interrupts, configuring interrupt triggering conditions (such as triggering an interrupt when the number of received data packets reaches a threshold).

[0036] As an optional embodiment, the operating status of the network card hardware can also be collected in real time, including: link status: link status S_link, speed, and duplex mode of the RoCE port; link negotiation result and MTU (Maximum Transmission Unit) value of the TCP / IP port; queue status: utilization rate of the RoCE hardware queue U_queue, send / receive completion queue pointer; buffer occupancy rate of the TCP / IP socket queue; error status: error counter of the network card (CRC (Cyclic Redundancy Check) error, frame loss, DMA (Direct Memory Access) error, etc.), and a hardware abnormality alarm is triggered when the number of errors exceeds the threshold.

[0037] This embodiment fully leverages the performance potential of the selected transport protocol by directly calling the underlying hardware interface matching the protocol type (such as RoCE's ibv_post_send or TCP / IP's DPDK function), especially RoCE's hardware offloading and low-latency advantages. Fine-grained hardware control (including link configuration, queue management, and interrupt optimization) is implemented during connection establishment and data transmission, dynamically adjusting underlying resource behavior according to business needs to further improve transmission efficiency and system response speed. A unified receive callback mechanism (such as ut_recv) shields the differences in underlying protocols, providing a simple and consistent data receiving interface for upper-layer services and reducing the complexity of business layer adaptation. This embodiment ensures seamless integration from resource allocation to actual transmission, improving single-service transmission performance while also enhancing the system's overall processing capability and stability in multi-protocol mixed scenarios.

[0038] In an exemplary embodiment, determining the primary transmission protocol type of a service based on its transmission demand characteristics includes: collecting the transmission demand characteristics of the service; classifying the service based on the transmission demand characteristics to obtain a service classification result; and determining the primary transmission protocol type of the service based on the service classification result.

[0039] Specifically, transmission demand characteristics reflecting the performance requirements and behavioral patterns of a service can be dynamically obtained by listening to network call interface parameters initiated by the service, sampling traffic characteristics at the initial stage of transmission, or matching from a pre-configured policy table in combination with the service identifier.

[0040] By clustering or matching transmission demand characteristics, different services can be categorized into groups with similar transmission needs. This classification method helps abstract complex, multi-dimensional transmission demands into a limited set of typical scenarios, thereby simplifying the complexity of subsequent protocol decisions and supporting consistent and efficient resource scheduling strategies for similar services. During classification, the distribution patterns of features across dimensions such as latency, throughput, and connection scale can be used to achieve fine-grained or coarse-grained differentiation of service scenarios. For example, based on analysis of latency, throughput, and connection scale, services can be categorized into low-latency high-throughput, high-concurrency connection, and balanced performance categories. At a finer-grained level, the low-latency high-throughput category can be further distinguished into extremely latency-sensitive and high-throughput-dominant types, while at a coarse-grained level, these subcategories can be merged into a unified high-performance category. This classification method not only abstracts heterogeneous service demands into a limited set of decision-making scenarios, reducing the complexity of subsequent protocol matching and resource scheduling, but also provides consistent and efficient resource allocation strategies for similar services, thereby achieving a balance between resource optimization and performance assurance at the system level.

[0041] The classification results provide a clear decision-making basis for protocol selection. Based on the mapping relationship between different categories and protocol characteristics, the most suitable transmission protocol type can be quickly and accurately matched for each type of service, avoiding the overhead of complex evaluation for each service, improving decision-making efficiency, and ensuring the rationality and consistency of protocol selection. In implementation, queries can be performed based on a predefined classification and protocol mapping table, or dynamic matching can be achieved through lightweight weight calculation or a policy rule engine, thereby realizing the automation and intelligence of protocol selection.

[0042] In one exemplary embodiment, transmission requirement characteristics include data volume characteristics, performance requirement characteristics, and connection characteristics; classifying services includes classifying services based on data volume characteristics, performance requirement characteristics, and connection characteristics.

[0043] The service classification results include a first type, a second type, and a third type; the performance requirement characteristics of the transmission requirement characteristics corresponding to the first type meet the first preset condition, the connection characteristics of the transmission requirement characteristics corresponding to the second type meet the second preset condition, and the transmission requirement characteristics of the third type do not meet the first and second preset conditions.

[0044] In this embodiment, the extraction of service transmission requirement features includes three dimensions: data volume features, performance requirement features, and connection features. Data volume features are defined by parameters such as the length of data transmitted in a single transmission and the transmission frequency per unit time. For example, the `len` parameter records the size of a single transmission, and the transmission frequency is recorded by counting the number of `ut_send` calls. Performance requirement features are defined by parsing identifying parameters in the service calls, such as the `flags` parameter containing requirements for low latency, high throughput, or high reliability. Connection features describe the form of the network session, and their metrics include the connection duration from `ut_connect` to `ut_close`, and the number of concurrent `ut_connect` connections. All collected feature data is stored in the service feature database and managed according to service ID.

[0045] When classifying business operations, K-means clustering can be used, but is not limited to, this algorithm. K-means clustering is a data-driven method that automatically groups data points based on distances between them. It takes a multi-dimensional feature vector composed of data volume, performance requirements, and connectivity features as input, without requiring a pre-set fixed classification threshold. In its implementation, a feature vector is constructed for each business operation. For example, numerical indicators such as average bytes transmitted per transmission, performance strength values ​​parsed from identifiers, and average concurrent connections are combined. Cluster centers are determined through iterative calculations, and a classification label is ultimately output for each business operation.

[0046] The business classification results mainly include three categories with clear criteria. The first type is RoCE-priority, determined by performance requirements meeting specific conditions (the first preset condition in this embodiment). For example, a business explicitly marked as low latency with a latency requirement below 100 microseconds, or marked as high throughput with a throughput requirement exceeding 10Gbps, is classified into this category. The second type is TCP / IP-priority, determined by connection characteristics meeting specific conditions (the second preset condition). For example, a business with more than 1000 concurrent connections and a single data transmission length of less than 1KB, such as web services or lightweight API calls, falls into this category. The third type is automatic adaptation, encompassing all other businesses that do not meet the specific conditions of the first and second types, such as regular file transfers or log synchronization tasks. The classification results are updated to the business classification table in real time.

[0047] By setting specific and quantifiable predefined conditions (such as requiring latency below 100 microseconds or throughput exceeding 10Gbps for the first type, and concurrent connections exceeding 1000 with single data transmissions less than 1KB for the second type), clear and objective business classification criteria are provided, avoiding inconsistencies caused by subjective judgment. Secondly, this classification framework based on multi-dimensional features and explicit conditions enables automated classification algorithms (such as K-means clustering) to have clear objectives and interpretability. The clustering results can be directly correlated with the technical characteristics of the business (such as a preference for RoCE or TCP / IP protocols), thereby achieving an efficient and accurate mapping from business characteristics to resource scheduling strategies. Finally, through structured feature definitions and type classifications, this solution provides a reliable foundation for intelligent and differentiated scheduling of network resources (such as high-performance RoCE network cards and general-purpose TCP / IP stacks), improving overall resource utilization efficiency and service quality.

[0048] In an exemplary embodiment, determining the primary transmission protocol type of a service based on the service classification result includes: if the service classification result is a first type or a second type, determining the primary transmission protocol type of the service according to a preset priority rule; if the service classification result is a third type, determining the primary transmission protocol type of the service by comparing the scheduling weights of link resources of different transmission protocol types.

[0049] Specifically, when the service is of type 1, the primary transmission protocol type of the service is determined to be type 1 protocol; when the service is of type 2, the primary transmission protocol type of the service is determined to be type 2 protocol.

[0050] This embodiment determines the primary transmission protocol type based on the service classification result. Specifically, if the service classification result is a first type or a second type, the primary transmission protocol type will be directly determined according to a preset priority rule; if the service classification result is a third type, it will be dynamically determined by comparing the real-time scheduling weights of link resources corresponding to different transmission protocol types. Specifically, the primary transmission protocol type for first-type services is determined as the first protocol type, such as the RoCE protocol; the primary transmission protocol type for second-type services is determined as the second protocol type, such as the TCP / IP protocol.

[0051] When outputting specific transmission strategies, differentiated solutions are adopted for different service types. For Type A services (Type 1), resource allocation prioritizes RoCE links, ensuring low latency during failover, using the RoCEVerbs interface for protocol adaptation, and forcibly enabling kernel bypass. For Type B services (Type 2), resource allocation prioritizes TCP / IP links, ensuring connection stability during failover, using the Socket interface for protocol adaptation, and enabling kernel bypass as needed. For Type C services (Type 3), resource allocation uses an automatic selection mechanism, making decisions based on real-time link scheduling weights. Failover uses the system default strategy, protocol adaptation interfaces are automatically mapped, and kernel bypass functionality is dynamically adjusted according to actual performance requirements.

[0052] In this embodiment, a preset priority rule is applied to the first and second types of services, which is a deterministic and efficient path. For the third type of service, a dynamic decision-making process is triggered. The resource scheduling unit calculates the weight based on the real-time link status data, and the protocol type with the higher weight will be selected. This allows the third type of service to adaptively utilize the current optimal resources. As another optional embodiment, when the weight of the preferred RoCE link for the first type of service is lower than a certain security threshold, it is allowed to temporarily downgrade to select other protocol links.

[0053] This embodiment effectively optimizes the balance between overall resource utilization efficiency and business service quality through this refined scheduling strategy.

[0054] In an exemplary embodiment, establishing a network connection based on the allocated primary link resources to perform service data transmission includes: when the primary transport protocol type of the service is a second protocol type, performing service data transmission through a kernel bypass mode; the kernel bypass mode includes: bypassing the network protocol stack of the operating system kernel, implementing network protocol processing in user space, and utilizing the hardware acceleration function of the network interface card (NIC) to transmit data. Utilizing the hardware acceleration function of the NIC to transmit data includes: directly accessing the user-space data buffer through the NIC's direct memory access engine; enabling the NIC's hardware verification and offloading functions. Data fragmentation and reassembly operations are performed by the NIC's hardware.

[0055] When a service is assigned a secondary transport protocol type (such as TCP / IP), data transmission is performed using kernel bypass mode. The core of kernel bypass mode is to bypass the traditional network protocol stack of the operating system kernel, instead implementing the necessary network protocol processing logic in user space and fully utilizing the network interface card's (NIC) hardware acceleration capabilities to complete data transmission. Specifically, utilizing hardware acceleration includes directly reading and writing user-space data buffers through the NIC's direct memory access engine, enabling the NIC's hardware checksum and offload functions, and having the NIC hardware perform data fragmentation and reassembly operations.

[0056] This embodiment is implemented based on DPDK technology. The process begins with initializing the DPDK environment, including loading a specific driver and binding the network interface card (NIC) port to the DPDK user-space driver. Subsequently, large pages of memory are allocated in user space as a data buffer to avoid data copying between kernel and user space. By configuring the takeover of NIC interrupts, the interrupt handler is moved to user space to reduce context switching overhead. This implements a lightweight user-space TCP protocol stack based on DPDK to support complete TCP connection management, thus bypassing the kernel's standard protocol processing flow. This mode provides a bypass switch interface, allowing dynamic enabling or disabling of the kernel bypass function; it is enabled by default and automatically disabled under specific conditions such as RoCE NIC failure.

[0057] To further improve performance, this embodiment utilizes the hardware resources of the RoCE network card to accelerate TCP / IP data transmission. Its core optimizations are reflected in three aspects: First, DMA data transmission: through the RoCE network card's DMA engine, data in user-space big page memory is directly transferred to the network card's hardware buffer, avoiding the involvement of the central processing unit; second, hardware checksum: by configuring and enabling the network card's TCP checksum offload function, the hardware replaces software in calculating the checksum of the protocol header and data; third, fragmentation and reassembly: for data packets exceeding the maximum transmission unit, the network card hardware automatically performs fragmentation and reassembly at the receiving end, thereby reducing the processing burden on the user-space protocol stack.

[0058] This embodiment provides a high-performance accelerated transmission solution for services assigned to TCP / IP paths. Related technologies using the kernel protocol stack for TCP / IP transmission suffer from inherent bottlenecks such as data copying and frequent context switching. This embodiment significantly reduces CPU intervention and system overhead by moving protocol processing to user space and deeply integrating the network interface card's (NIC) hardware offloading capabilities, thereby substantially reducing transmission latency and increasing throughput. This not only ensures the transmission performance of non-RoCE path services and makes protocol scheduling and failover more flexible, but also fully leverages the potential of the same NIC hardware across multiple protocol modes, maximizing resource utilization and enhancing the overall system's ability to handle high load and high availability requirements.

[0059] In an exemplary embodiment, the network resource scheduling method further includes: obtaining performance indicators of the kernel bypass mode, including transmission efficiency improvement rate and CPU utilization rate; and controlling the kernel bypass mode to be started or stopped based on the performance indicators.

[0060] The transmission efficiency improvement rate of the kernel bypass mode is obtained using the first relational expression, which is: ,in, To improve transmission efficiency, For the transmission latency of the kernel protocol stack, To bypass the transmission delay after the operating system kernel.

[0061] Specifically, if the transmission efficiency improvement rate is lower than the first performance threshold, or the CPU utilization rate is higher than the first resource threshold, the kernel bypass mode is turned off; if the transmission efficiency improvement rate is not lower than the second performance threshold and the CPU utilization rate is lower than the second resource threshold, the kernel bypass mode is kept on; wherein, the first performance threshold is less than or equal to the second performance threshold, and the first resource threshold is greater than or equal to the second resource threshold.

[0062] In this embodiment, performance metrics during the operation of this mode are continuously acquired, and its activation or deactivation is dynamically controlled based on these metrics. Performance metrics include, but are not limited to, transmission efficiency improvement rate and CPU utilization. The transmission efficiency improvement rate is calculated using a quantification formula (first relation), which is: Where E_bypass represents the transmission efficiency improvement rate, and T_kernel represents the transmission latency of the traditional kernel protocol stack. T_copy represents the data copying time between kernel mode and user mode; T_stack represents the kernel protocol stack processing time (TCP handshake, routing, forwarding, etc.); T_int represents the kernel interrupt handling time; and T_bypass represents the transmission delay after bypassing the operating system kernel. T_dma is the DMA data transfer time; T_hw is the network card hardware processing time (checksum, fragmentation, etc.); and T_user is the user-space lightweight protocol stack processing time.

[0063] Real-time monitoring of transmission performance after kernel bypass is performed, collecting the following performance metrics, including but not limited to transmission latency calculated from the difference between network interface card (NIC) hardware and user-space timestamps, throughput calculated based on NIC statistics, CPU utilization of the logical core running the user-space protocol stack, and packet loss rate calculated by comparing the number of sent and acknowledged data packets. Based on the collected performance metrics, a decision rule based on dual thresholds is used to control the kernel bypass mode. If the transmission efficiency improvement rate is lower than the first performance threshold, or the CPU utilization rate is higher than the first resource threshold, the kernel bypass mode is disabled; if the transmission efficiency improvement rate is not lower than the second performance threshold and the CPU utilization rate is lower than the second resource threshold, the mode is kept enabled. The first performance threshold is less than or equal to the second performance threshold, and the first resource threshold is greater than or equal to the second resource threshold. This threshold relationship is designed to create a hysteresis interval to prevent control commands from oscillating frequently near the critical point, thereby ensuring the smoothness of system state transitions.

[0064] In practical implementation, explicit numerical examples can be set for these thresholds. For example, the first performance threshold can be set to... The second performance threshold is Set the first resource threshold to The second resource threshold is The corresponding control logic is as follows: when the efficiency improvement rate E_bypass is detected to be lower than... Or the CPU utilization rate is higher than When E_bypass is not less than 1, automatically disable the kernel bypass function and switch back to the traditional kernel protocol stack; And the CPU utilization rate is lower than When this happens, the kernel remains off-limits. If the indicator is in a lag zone, for example, the rate of increase is... And the occupancy rate is If the system remains in its current operating state, the system will continue to operate as before. This embodiment establishes a dynamic control mechanism based on real-time performance feedback, enabling the kernel bypass mode to be actively utilized when the benefits are significant, and to intelligently fall back to the standard kernel path when the performance gains are insufficient or the resource consumption is excessive, thereby achieving an optimal balance between improving transmission performance and ensuring the overall stability and resource efficiency of the system.

[0065] In one exemplary embodiment, the primary transmission protocol type of a service is determined by comparing the scheduling weights of link resources of different transmission protocol types, including: obtaining the real-time status of link resources corresponding to multiple transmission protocol types in the converged resource pool; calculating the scheduling weight of each transmission protocol type based on the real-time status of the link resources of the transmission protocol type; and determining the transmission protocol type with the highest scheduling weight as the primary transmission protocol type of the service.

[0066] It is understandable that, for the third type of service, the primary transmission protocol type is determined by comparing the scheduling weights of the link resources corresponding to different transmission protocol types. First, real-time status information of all link resources is obtained from the converged resource pool. This status information includes, but is not limited to, key indicators such as available bandwidth, current latency, and load rate for each link. Each transmission protocol type (such as RoCE and TCP / IP) in the converged resource pool corresponds to one or more link resources. Based on the real-time status of all links under each protocol type, a comprehensive scheduling weight is calculated. Then, the scheduling weights of all protocol types are compared, and the transmission protocol type with the highest weight value is determined as the current primary transmission protocol type for that service. Simultaneously, within this selected protocol type, the specific link with the highest weight value can be designated as the primary link resource. This decision-making process is dynamic; real-time changes in network status will cause changes in scheduling weights. Therefore, the same service accessing the system at different times may be assigned different primary transmission protocol types and primary link resources.

[0067] The dynamic weight comparison mechanism designed in this embodiment provides highly adaptive protocol selection capabilities for the third type of service. This allows service traffic to automatically flow to relatively idle or higher-performance resource areas, thereby achieving dynamic load balancing and maximizing utilization of the entire resource pool. Specifically, when the overall TCP / IP network load is high, causing its scheduling weight to decrease, idle RoCE links will receive higher weights due to their high availability bandwidth and low latency, making them available for the third type of service. This improves the transmission experience of this type of service and fully utilizes idle high-performance hardware resources. Conversely, when the RoCE network is busy, the third type of service will automatically tend to choose relatively idle TCP / IP links, thus avoiding exacerbating congestion on the high-performance network. This intelligent scheduling strategy, which dynamically adjusts according to system status, effectively optimizes overall resource utilization efficiency.

[0068] In one exemplary embodiment, calculating the scheduling weight of a transmission protocol type based on the real-time status of link resources of that transmission protocol type includes: calculating the scheduling weight of the transmission protocol type using a second relation, wherein the second relation is... ,in, For link scheduling weights, For the available bandwidth of the link, This represents the total bandwidth of the link. For link load rate, The average link delay, , , All are weighting coefficients. .

[0069] In this embodiment, W_link represents the final calculated link scheduling weight, which is a dimensionless value used for horizontal comparison; B_avail represents the currently available bandwidth of the link; B_total represents the total bandwidth of the link, and the ratio of the two reflects the remaining proportion of bandwidth; T_delay represents the current average transmission delay of the link; and U_load represents the current load rate of the link. , , These are three configurable weighting coefficients, which sum to 1, representing the relative importance of the three optimization objectives—available bandwidth ratio, low latency, and low load rate—in the overall decision-making process. It is a very small positive integer whose purpose is to prevent the calculation error of zero denominator when the delay T_delay is zero.

[0070] In its implementation, the resource scheduling unit periodically calculates the W_link value for each active link in the resource pool using this formula. The real-time variables required for the formula are provided by the resource awareness unit, such as obtaining B_avail and B_total by querying the network interface card counter, obtaining T_delay by sending measurement packets or analyzing timestamps, and comprehensively deriving U_load through methods such as the ratio of used bandwidth to total bandwidth. Weighting coefficients. , , As a key configuration parameter of the system, it determines the tendency of the scheduling strategy. For example, in scenarios that emphasize throughput, a larger value can be set. In scenarios where low latency is emphasized, a larger value can be set. These coefficients can be dynamically adjusted during system operation, and the calculated W_link values ​​will be cached and used for subsequent scheduling decisions.

[0071] As an optional implementation, to ensure system reliability and business continuity, a certain proportion of redundant resources can be reserved in the converged resource pool in advance. When the RoCE network fails (such as link interruption or network card failure), the switching mechanism is quickly triggered to seamlessly migrate business traffic to the reserved redundant resources, preventing business interruption due to single point of failure. At the same time, the existence of redundant resources can also cope with sudden peak business traffic and ensure that the system always maintains stable and efficient operation.

[0072] In an exemplary embodiment, during the process of establishing a network connection based on the allocated primary link resources to perform the transmission of service data, the network resource scheduling method further includes: detecting faults in the primary link resources; and when a fault is detected and the switching conditions are met, triggering a switching operation to forward the service data to the backup link resources of the backup transmission protocol type allocated for the service for transmission.

[0073] When a service transmits service data through the main link resource, the process also includes steps such as performing fault detection on the main link resource and switching over when conditions are met. First, multi-dimensional proactive monitoring and fault detection of the health status of the main link resource, especially the RoCE network link, are performed, and fault events are recorded in a fault event table. This step specifically includes: periodically sending link-layer probe packets to detect connectivity; if no response is received for more than three consecutive probe cycles, or the link packet loss rate exceeds a preset threshold, or the bandwidth utilization rate remains below a certain threshold... If accompanied by a large number of retransmissions, it is determined to be a link failure; monitor the physical status and error counter of the port. If the port status is Down, or the error counter increments by more than 100 times within 10 seconds, it is determined to be a port failure; monitor the depth, packet loss count, and latency indicators of the hardware queue in real time. If the queue remains full for more than 5 seconds, or the sudden packet loss rate exceeds [a certain threshold], it is determined to be a link failure. If the average latency exceeds twice the Service Level Agreement (SLA) threshold, it is considered a queue failure. Simultaneously, for the RoCE protocol stack, indicators such as priority flow control frame backlog, false triggering of explicit congestion notification markers, and remote direct memory access request timeouts are checked. If five consecutive flow control credit values ​​are zero, or the explicit congestion notification marker rate abnormally increases to exceed normal traffic levels... Or the remote direct memory access request timeout rate exceeds If the monitoring conditions in any dimension are met, a protocol failure is determined. A failure alarm is triggered when the monitoring conditions in any dimension are met. In this embodiment, multi-dimensional data is combined for correlation analysis to locate the failure. For example, when a link failure is accompanied by a protocol timeout, the link is repaired first.

[0074] Secondly, upon detecting a fault event, the handover decision-making process begins. This step does not respond immediately to all faults; instead, it assesses whether the fault meets preset handover conditions. A fault risk value is calculated by considering the probability of the fault occurring, the estimated recovery time, and the impact on current services. Only when this risk value exceeds a preset handover threshold is the handover condition deemed met, and subsequent operations are triggered.

[0075] Finally, the handover operation is performed. Once the handover conditions are met, the data flow through the primary transport protocol link is paused, the target protocol is adapted through a protocol conversion mechanism, and the service data flow is seamlessly forwarded to a backup link resource of another transport protocol type pre-allocated for the service to continue transmission.

[0076] This embodiment improves the overall stability and availability of the system by using refined fault monitoring and intelligent switching condition judgment to quickly respond to serious faults and ensure business continuity, while avoiding unnecessary switching due to momentary jitter.

[0077] The recorded fault event table plays the following roles in system operation and optimization: First, it supports rapid fault switching decisions. Based on the severity of the events recorded in the table and the associated service priority, traffic switching can be performed immediately. For example, service traffic on a faulty link can be quickly migrated to other available links within the same subnet. The service priorities are pre-defined by the system based on the service type. For instance, latency-sensitive real-time services are given high priority, enjoying priority in resource allocation and fault switching to ensure the continuity of critical services, while tasks with low latency requirements are assigned lower priority.

[0078] Secondly, the fault event table provides a data foundation for dynamic resource scheduling. By analyzing historical fault data, resource allocation and various strategies can be proactively adjusted to avoid allocating new services to high-fault areas and optimize overall performance. Specific adjustment methods include, but are not limited to: if data shows that the RoCE link has a high failure rate during a specific period, its resource allocation priority is reduced, and TCP / IP links are prioritized instead, with resource allocation only implemented when network load is low. RoCE will only be considered when necessary; if low-latency services frequently experience latency exceeding limits after switching, the failover strategy will be adjusted, adopting dynamic priority and reserving additional bandwidth for core services; if protocol adaptation based on the RoCEVerbs interface frequently encounters compatibility issues, a hybrid adaptation scheme of RoCEv2 and Socket interfaces will be introduced, dynamically selecting the interface based on application compatibility test results; if forcibly enabling kernel bypass mode leads to a decrease in system stability, the strategy will be changed to dynamic adjustment based on service type and system load, for example, enabling it for storage I / O intensive services to improve performance, and temporarily disabling it when network traffic is sudden and unstable to ensure system stability.

[0079] Furthermore, the fault event table also serves root cause analysis and early warning. By performing correlation analysis on the timestamps, fault types, and impact ranges of events in the table, fault prediction models can be constructed, thereby triggering preventative maintenance measures in advance to prevent problems before they occur. In summary, through the systematic recording and application of fault events, this embodiment not only achieves rapid fault response but also realizes a shift from passive handling to proactive optimization, significantly improving the system's self-healing capabilities, resource utilization efficiency, and long-term operational stability.

[0080] In one exemplary embodiment, when a fault is detected and the switching conditions are met, a switching operation is triggered, including: when a fault is detected, calculating a fault risk value using a third relation; when the fault risk value is greater than a preset switching threshold, triggering the switching operation; the third relation is... ;in, As the first risk factor, As the second risk factor, The third risk factor, The probability of failure occurring. To estimate the fault recovery time, This is the business timeout threshold. For the amount of business data affected, Total business data volume This represents the fault risk value.

[0081] In this embodiment, R_fault represents the calculated fault risk value, a scalar between 0 and 1, with a higher value indicating a higher risk. P_fault represents the probability of the fault occurring, which is a prior or posterior probability based on statistics or prediction. T_recover represents the estimated fault recovery time, and T_timeout represents the maximum timeout that the service can tolerate; the ratio of the two reflects the urgency of the situation. D_impact represents the amount of service data affected by this fault, and D_total represents the total amount of service data; the ratio of the two reflects the scope of the impact. As the first risk factor, As the second risk factor, The third risk coefficient is assigned different importance to three dimensions: failure probability, time urgency, and impact scope, with a sum of 1. The decision rule is: when the calculated R_fault is greater than a preset switching threshold (e.g., >0.6), the switching condition is met. This allows failures of different businesses and scales to be compared under the same standard. The preset switching threshold (e.g., 0.6) is an empirical value derived from extensive stress testing and business scenario analysis, used to balance the sensitivity and stability of the switching.

[0082] In one exemplary embodiment, triggering a handover operation includes: pausing the transmission of service data through the primary link resource; converting the service data to be sent from the protocol format of the primary transmission protocol type to the protocol format of the backup transmission protocol type; and sending the protocol-converted service data through the established backup link resource of the backup transmission protocol type.

[0083] First, data transmission through the failed primary link is suspended to immediately block new data flows on the failed link and stabilize the current transmission state, preventing data loss or out-of-order delivery. Second, the service data to be sent is converted from its primary transport protocol format to a backup transport protocol format, involving a complete reconstruction of the protocol header information and data encapsulation format. Finally, the converted service data is sent through a pre-established backup link resource belonging to the backup transport protocol type, thereby restoring the service flow transmission on the new link.

[0084] The specific implementation of protocol format conversion lies in the bidirectional mapping between RoCE and TCP / IP protocols. During the conversion from RoCE to TCP / IP, fields such as the queue pair number (QPN), packet sequence number (PSN), and queue key (QKey) in the RoCE header are extracted and mapped to the corresponding port number, sequence number, and checksum in the TCP header. Simultaneously, the RDMA data frame is repackaged into a complete TCP data segment. During the reverse conversion, information such as the port number and sequence number in the TCP header is extracted, mapped back to the parameters required by the RoCE protocol, and the TCP data segment is reassembled into an RDMA data frame with scatter-gather list (SGE) elements. To ensure the reliability of the conversion process, a state machine with four states—idle, in progress, completed, and abnormal—is designed. State transitions manage the entire process, and a retry mechanism is automatically triggered when conversion fails, with a maximum of three retries.

[0085] At the execution level, once a switchover is deemed necessary, data submission to the hardware queue of the original failed link is immediately suspended, and the sequence number of the last successfully acknowledged data packet is recorded as a breakpoint. Subsequently, all unacknowledged data to be sent is read and converted to the correct protocol format according to the aforementioned mapping rules. After conversion, a pre-existing backup transmission link connection, such as a TCP / IP socket connection, is activated, and the converted data packets are sent out by calling a standard sending interface. The entire process is designed to be completed in a very short time, thus achieving rapid failover and service recovery transparent to upper-layer services.

[0086] In an exemplary embodiment, the network resource scheduling method further includes: after triggering a switching operation, continuously monitoring the recovery status of the primary link resource of the primary transport protocol type; when it is confirmed that the primary link resource has recovered stably, gradually migrating the service data from the backup link resource of the backup transport protocol type back to the primary link resource of the primary transport protocol type.

[0087] After the service is switched to backup link resources, the network resource scheduling method also includes continuously monitoring the recovery status of the primary link resources of the primary transport protocol type, and gradually migrating the service data back after confirming stable recovery. Monitoring the recovery status is a multi-faceted verification process. First, physical connectivity is checked by periodically sending link-layer probe packets to the target node. If a certain number of normal responses are received consecutively and the response time is within a threshold range, physical link recovery is preliminarily determined. Second, the status of the RoCE protocol control plane is checked, such as confirming whether the queue pair status has recovered from error to ready, and verifying whether mechanisms such as priority flow control and explicit congestion notification can interact normally. Next, actual transmission performance indicators are monitored, such as when the packet loss rate of data transmission is lower than a preset threshold and the throughput recovers to the pre-failure level. At this point, performance can be considered to have returned to normal. Finally, application-layer verification is performed by simulating small-volume business data or actually taking over a very small amount of business traffic to ensure that the business can send and receive data normally without errors, thereby ultimately confirming that the main link resources have been fully restored to stability.

[0088] Once the primary link resources are confirmed to be restored, a gradual data migration process is initiated. This process is typically launched within a low-load service window to minimize the impact on services. At the start of the migration, the necessary communication context on the primary link resources is first rebuilt, such as re-establishing RoCE queue pairs. Subsequently, a small portion of service traffic, such as... The system switches back to the primary link from the backup link resource and continuously monitors it for a short period, such as 200 milliseconds, to confirm that no anomalies occur. After confirming that the migration of this batch of traffic is stable, the migration ratio can be dynamically calculated and gradually increased based on real-time network load and resource pool capacity to achieve intelligent load adjustment. After each increase in the migration ratio, a similar stable observation period is maintained. This gradual process continues until... All business traffic was successfully migrated back to the primary link resources. Throughout the migration process, a format conversion from the backup transport protocol to the primary transport protocol was performed synchronously, such as converting TCP / IP data back to RoCE format. Once all traffic had successfully migrated and was running stably, the resources occupied by the backup link were released, and the upper layers of the service were notified of the migration completion. This series of steps, through careful verification and gradual traffic switching, ensured high reliability and business continuity during the fault recovery process.

[0089] In an exemplary embodiment, the network resource scheduling method further includes: releasing the primary link resources and backup link resources allocated for the service when the service data transmission ends; and marking the status of the released link resources in the converged resource pool as idle.

[0090] When a service actively closes a connection or the system determines that a session has ended, the primary and backup link resources allocated to that service are immediately released, and the status of these released resources in the converged resource pool is remarked as idle. Specifically, after the upper-layer service completes transmission, it calls interfaces such as ut_close to close the connection. Subsequently, the transmission execution and control module stops related data transmission and reception operations and performs specific resource release actions, such as destroying RoCE queue pairs or closing TCP / IP socket connections. At the same time, this module notifies the network resource pool management module to reclaim resources. The resource pool management module then updates its internal resource allocation table and global status view, clearing the occupant information from the status fields of released hardware queues, ports, memory, and connection handles, and marking them as idle, so that they return to the shared resource pool and become system-available resources for subsequent service requests.

[0091] This embodiment ensures the sustainable recycling of resources in the integrated resource pool, solves the problem of resource idleness and waste that may occur during long-term operation, improves the fairness of resource scheduling and long-term operating efficiency, and provides basic support for the healthy, stable and high-performance service of the entire system.

[0092] Reference Figure 2 It includes a business layer, a protocol encapsulation and adaptation module, a business requirement awareness module, a network resource pool management module, a fault detection and switching module, a kernel bypass optimization module, a transmission execution and control module, and a network interface card (NIC) hardware layer. Specifically, the network resource pool management module includes a resource awareness unit, a resource allocation unit, and a resource scheduling unit; the fault detection and switching module includes a fault awareness unit, a fault decision unit, a switching execution unit, and a recovery and migration unit; the protocol encapsulation and adaptation module includes an interface adaptation unit, a protocol conversion unit, and a data encapsulation unit; the kernel bypass optimization module includes a kernel bypass unit, a hardware acceleration unit, and a performance monitoring unit; the business requirement awareness module includes a business feature acquisition unit, a requirement classification unit, and a policy output unit; and the transmission execution and control module includes a data transmission and reception unit, a hardware control unit, and a status feedback unit.

[0093] Combination Figure 3 The above solution will be illustrated with a specific embodiment.

[0094] Step 1: System Initialization and Resource Pool Construction: Start the network resource pool management module. The resource awareness unit collects resource data through active probing (sending rping to RoCE links and ICMP Echo to TCP / IP links): Total bandwidth of RoCE links is 100Gbps, available bandwidth is 80Gbps, latency is 30μs, and port status is normal; Total bandwidth of TCP / IP links is 25Gbps, available bandwidth is 22Gbps, latency is 100μs, and port status is normal; The converged resource pool is divided into RoCE resource partitions and TCP / IP resource partitions; Start the protocol encapsulation and adaptation module and load the unified_transport_api interface; Start DPDK, bind the network card port to the igb_uio driver, and allocate 16GB of bigpage memory; All modules are ready.

[0095] Step 2: Business Access and Demand Awareness: The upper-layer business system calls ut_init() to initialize the interface; the business feature collection unit collects the transmission data length and flags low latency + high throughput; the demand classification unit determines it as type A business; the policy output unit outputs RoCE priority allocation of resources, reserves 2 TCP / IP sockets for backup, and uses the kernel bypass policy.

[0096] Step 3: Resource allocation, including primary transmission link selection: The resource allocation unit allocates RoCE resources (port 1, hardware queues 1-2) to the upper-layer services and reserves TCP / IP resources (port 2, socket handles 100-101); the resource scheduling unit calculates the RoCE link and the TCP / IP link, and selects the RoCE link; the service establishes a RoCE connection with the storage node through ut_connect().

[0097] Step 4: Data transmission and performance optimization: The upper-layer business calls ut_send() to transmit data, and the interface adaptation unit maps it to ibv_post_send call. The data is transmitted through the RoCE hardware queue; the kernel bypasses the optimization module to enable DMA transmission and hardware checksum, and the performance monitoring unit collects latency and throughput, keeping kernel bypass enabled.

[0098] Step 5: Main transmission link fault detection, mainly targeting RoCE link fault detection: The fault perception unit detects that the packet loss rate of the RoCE link exceeds the threshold, records the error counter of port 1, and records the fault event (link fault, time T1); When a fault is detected, the fault decision unit performs fault risk assessment, calculates the parameters, and if the switching conditions are met, outputs "Switch immediately"; The switching execution unit freezes RoCE queues 1-2, allocates TCP / IP Socket handles 100-101, the protocol conversion unit converts the RoCE data into TCP / IP format, re-establishes the connection through the Socket interface, and sends the switch completion message to the upper-layer service at time T2.

[0099] Step 6: Main transmission link recovery detection. If recovered, proceed with RoCE recovery migration: At time T3, the recovery migration unit detects that the RoCE link status has recovered. It periodically sends probe packets to the target node. If a certain number of response packets are received consecutively, and the response time is within the normal threshold range, the physical link connection is preliminarily determined to have recovered. Check the control plane status of the RoCE protocol, such as whether the QP state has recovered from the error state to the normal working state. Simultaneously verify whether the PFC and ECN mechanisms can interact normally, ensuring no congestion signals are continuously triggered. Monitor key indicators such as packet loss rate and throughput. When the packet loss rate is lower than the preset threshold and the throughput recovers to the level before the fault... The above indicates that transmission performance has returned to normal. For links carrying services, simulate low-volume service data transmission to verify the connection status of the application layer protocol and whether the service response meets expectations. If the service can send and receive data normally without any abnormal errors, the recovery is finally confirmed. Obtain the off-peak window from the upper-layer service (by real-time monitoring of network traffic, device load, and service request frequency data indicators, combined with historical data statistics and machine learning algorithms, to identify time periods with lower network load and relatively abundant bandwidth resources); at time T4, rebuild RoCE queues 1-2 and migrate... Traffic to the RoCE link was monitored for 200ms with no anomalies detected; the migration ratio was then gradually increased. When increasing the migration ratio, the ratio can be dynamically adjusted each time based on real-time network load, remaining resource pool capacity, etc., and is not fixed to the same value. After each migration operation is completed, it is necessary to continuously monitor for 200ms to confirm that the system is running normally, thereby ensuring the stability and reliability of RoCE fault switching and network card converged transmission; release Socket handles 100-101 and send a migration completion message to the service.

[0100] Step 7: Transmission End and System Reset: The upper-layer service completes data transmission and calls ut_close() to close the connection; the transmission execution and control module releases RoCE queues 1-2, and the resource pool management module reclaims resources to the resource pool; each module updates its status to standby, waiting for the next service access.

[0101] In summary, this application constructs a RoCE and TCP / IP converged resource pool, achieving dynamic scheduling of both types of resources through a unified resource pool management module. Combined with a weighted fair scheduling algorithm (including link scheduling weight formulas), resource utilization is improved, avoiding single-network idleness. Based on multi-dimensional fault detection (link / port / queue / protocol stack) and risk assessment (including fault risk value formulas), rapid service switching during RoCE failures and a gradual migration strategy after fault recovery are implemented to ensure service continuity. A unified transmission interface, `unified_transport_api`, is designed to encapsulate Verbs and Socket interfaces. Combined with a kernel bypass optimization module (based on DPDK technology), it bypasses the TCP kernel protocol stack and utilizes RoCE hardware to accelerate TCP transmission, reducing latency. The system architecture includes six core modules, such as network resource pool management, RoCE fault detection and switching, with particular emphasis on protecting the resource scheduling weight calculation, the quantitative formula for fault risk assessment, and the coordination mechanism between modules. A seven-step converged transmission method, from system initialization, service access, resource allocation to fault switching, recovery migration, and system reset, focuses on protecting the service K-means classification strategy, kernel bypass switch control logic, and the gradual traffic migration process. The design of the unified interface unified_transport_api, the bidirectional conversion mechanism between RoCE and TCP / IP protocols, and the data encapsulation transmission control header and producer-consumer interaction model.

[0102] Please refer to Figure 4 The embodiments of this application also provide a network resource scheduling device, including: a policy determination module 11, used to determine the main transmission protocol type of a service based on the transmission requirement characteristics of the service in response to a service access request; a resource allocation module 12, used to allocate main link resources corresponding to the main transmission protocol type to the service in a converged resource pool, wherein the converged resource pool stores link resources of at least two transmission protocol types; and a transmission execution module 13, used to establish a network connection based on the allocated main link resources to execute the transmission of service data.

[0103] In an exemplary embodiment, determining the primary transmission protocol type of a service based on its transmission demand characteristics includes: collecting the transmission demand characteristics of the service; classifying the service based on the transmission demand characteristics to obtain a service classification result; and determining the primary transmission protocol type of the service based on the service classification result.

[0104] In one exemplary embodiment, transmission requirement characteristics include data volume characteristics, performance requirement characteristics, and connection characteristics; classifying services includes classifying services based on data volume characteristics, performance requirement characteristics, and connection characteristics.

[0105] In an exemplary embodiment, the service classification result includes a first type, a second type, and a third type; wherein, the performance requirement feature of the transmission requirement feature corresponding to the first type satisfies a first preset condition, the connection feature of the transmission requirement feature corresponding to the second type satisfies a second preset condition, and the transmission requirement feature of the third type does not satisfy the first preset condition and the second preset condition.

[0106] In an exemplary embodiment, determining the primary transmission protocol type of a service based on the service classification result includes: if the service classification result is a first type or a second type, determining the primary transmission protocol type of the service according to a preset priority rule; if the service classification result is a third type, determining the primary transmission protocol type of the service by comparing the scheduling weights of link resources of different transmission protocol types.

[0107] In an exemplary embodiment, determining the primary transmission protocol type of a service according to a preset priority rule includes: when the service is of a first type, determining the primary transmission protocol type of the service as a first protocol type; when the service is of a second type, determining the primary transmission protocol type of the service as a second protocol type.

[0108] In an exemplary embodiment, establishing a network connection based on the allocated primary link resources to perform the transmission of service data includes: when the primary transport protocol type of the service is a second protocol type, performing the data transmission of the service through a kernel bypass mode; wherein, the kernel bypass mode includes: bypassing the network protocol stack of the operating system kernel, implementing network protocol processing in user space, and using the hardware acceleration function of the network card to transmit data.

[0109] In one exemplary embodiment, data transmission using the hardware acceleration function of the network interface card (NIC) includes: directly accessing the user-mode data buffer through the NIC's direct memory access engine; enabling the NIC's hardware checksum and offload function; and having the NIC's hardware perform data fragmentation and reassembly operations.

[0110] In an exemplary embodiment, the network resource scheduling device is further configured to: obtain performance indicators of the kernel bypass mode, including transmission efficiency improvement rate and CPU utilization rate; and control the kernel bypass mode to be started or stopped based on the performance indicators.

[0111] In one exemplary embodiment, obtaining the transmission efficiency improvement rate of the kernel bypass mode includes: obtaining the transmission efficiency improvement rate of the kernel bypass mode using a first relation, wherein the first relation is: ,in, To improve transmission efficiency, For the transmission latency of the kernel protocol stack, To bypass the transmission delay after the operating system kernel.

[0112] In one exemplary embodiment, controlling the activation or deactivation of the kernel bypass mode based on performance metrics includes: deactivating the kernel bypass mode if the transmission efficiency improvement rate is lower than a first performance threshold or the CPU utilization rate is higher than a first resource threshold; and keeping the kernel bypass mode activated if the transmission efficiency improvement rate is not lower than a second performance threshold and the CPU utilization rate is lower than a second resource threshold; wherein the first performance threshold is less than or equal to the second performance threshold, and the first resource threshold is greater than or equal to the second resource threshold.

[0113] In one exemplary embodiment, the primary transmission protocol type of a service is determined by comparing the scheduling weights of link resources of different transmission protocol types, including: obtaining the real-time status of link resources corresponding to multiple transmission protocol types in the converged resource pool; calculating the scheduling weight of each transmission protocol type based on the real-time status of the link resources of the transmission protocol type; and determining the transmission protocol type with the highest scheduling weight as the primary transmission protocol type of the service.

[0114] In one exemplary embodiment, calculating the scheduling weight of a transmission protocol type based on the real-time status of link resources of that transmission protocol type includes: calculating the scheduling weight of the transmission protocol type using a second relation, wherein the second relation is... ,in, For link scheduling weights, For the available bandwidth of the link, This represents the total bandwidth of the link. For link load rate, The average link delay, , , All are weighting coefficients. .

[0115] In an exemplary embodiment, during the process of establishing a network connection based on the allocated primary link resources to perform the transmission of service data, the network resource scheduling device is further configured to: detect faults in the primary link resources; and when a fault is detected and the switching conditions are met, trigger a switching operation to forward the service data to the backup link resources of the backup transmission protocol type allocated for the service for transmission.

[0116] In one exemplary embodiment, when a fault is detected and the switching conditions are met, a switching operation is triggered, including: when a fault is detected, calculating a fault risk value using a third relation; when the fault risk value is greater than a preset switching threshold, triggering the switching operation; the third relation is... ;in, As the first risk factor, As the second risk factor, The third risk factor, The probability of failure occurring. To estimate the fault recovery time, This is the business timeout threshold. For the amount of business data affected, Total business data volume This represents the fault risk value.

[0117] In one exemplary embodiment, triggering a handover operation includes: pausing the transmission of service data through the primary link resource; converting the service data to be sent from the protocol format of the primary transmission protocol type to the protocol format of the backup transmission protocol type; and sending the protocol-converted service data through the established backup link resource of the backup transmission protocol type.

[0118] In an exemplary embodiment, the network resource scheduling device is further configured to: after triggering a handover operation, continuously monitor the recovery status of the primary link resource of the primary transport protocol type; and when it is confirmed that the primary link resource has recovered stably, gradually migrate service data from the backup link resource of the backup transport protocol type back to the primary link resource of the primary transport protocol type.

[0119] In an exemplary embodiment, the network resource scheduling device is further configured to: release the primary link resources and backup link resources allocated for the service when the service data transmission ends; and mark the status of the released link resources in the converged resource pool as idle.

[0120] For a description of the features in the embodiments corresponding to the network resource scheduling device, please refer to the relevant descriptions in the embodiments corresponding to the network resource scheduling method, which will not be repeated here.

[0121] Embodiments of this application also provide a network device, including: a communication unit for receiving service access requests and performing service data transmission based on an established network connection; a storage unit for storing a fused resource pool containing link resources of at least two transmission protocol types; and a processing unit connected to the communication unit and the storage unit, respectively, for: responding to a service access request received by the communication unit, determining the main transmission protocol type of the service based on the transmission requirement characteristics of the service; allocating main link resources corresponding to the main transmission protocol type for the service from the fused resource pool of the storage unit; and controlling the communication unit to establish a network connection based on the allocated main link resources to perform service data transmission.

[0122] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to perform the steps in any of the network resource scheduling method embodiments described above.

[0123] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described network resource scheduling method embodiments when it runs.

[0124] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0125] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described network resource scheduling method embodiments.

[0126] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described network resource scheduling method embodiments.

[0127] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0128] The foregoing has provided a detailed description of a network resource scheduling method, apparatus, and device provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are merely for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A network resource scheduling method, characterized in that, include: In response to a service access request, the system collects the service's transmission requirement characteristics, classifies the service based on these characteristics, obtains the service classification results, and determines the main transmission protocol type of the service based on these classification results. The transmission requirements include data volume characteristics, performance requirements characteristics, and connection characteristics, and the main transmission protocol type is RoCE protocol or TCP / IP protocol; In the converged resource pool, primary link resources corresponding to the primary transport protocol type are allocated to the service. The converged resource pool stores link resources of at least two transport protocol types. The link resources include at least one of RoCE hardware queues, ports, TCP / IP sockets, and connection handles. A network connection is established based on the allocated primary link resources to perform the transmission of business data; Based on the service classification results, the primary transport protocol type of the service is determined, including: If the service classification result is a first type or a second type, the main transmission protocol type of the service is determined according to the preset priority rules; If the service classification result is the third type, obtain the real-time status of the link resources corresponding to multiple transmission protocol types in the fusion resource pool; For each of the aforementioned transport protocol types, the scheduling weight of that transport protocol type is calculated using a second relation, wherein the second relation is: ,in, For link scheduling weights, For the available bandwidth of the link, This represents the total bandwidth of the link. For link load rate, This represents the average link latency. , , All are weighting coefficients. , It is a positive number; The transmission protocol type with the highest scheduling weight is determined as the main transmission protocol type of the service; In the converged resource pool, primary link resources corresponding to the primary transport protocol type are allocated to the service, including: In the fusion resource pool, the link resource with the highest scheduling weight among the multiple link resources under the main transmission protocol type is determined as the main link resource; In the process of establishing a network connection based on the allocated primary link resources to perform the transmission of service data, the network resource scheduling method further includes: Perform fault detection on the main link resources; When a fault is detected and the switching conditions are met, a switching operation is triggered to forward the service data to a backup link resource of the backup transmission protocol type allocated for the service for transmission; the fault includes at least one of link fault, port fault, queue fault, and protocol fault; When a fault is detected and the switching conditions are met, a switching operation is triggered, including: When a fault is detected, the fault risk value is calculated using the third relational formula. When the fault risk value exceeds a preset switching threshold, a switching operation is triggered. The third relation is: ; in, As the first risk factor, As the second risk factor, The third risk factor. The probability of failure occurring. To estimate the fault recovery time, This is the business timeout threshold. For the amount of business data affected, Total business data volume The fault risk value is defined as follows.

2. The network resource scheduling method according to claim 1, characterized in that, The services are categorized as follows: The services are classified based on the data volume characteristics, performance requirement characteristics, and connectivity characteristics.

3. The network resource scheduling method according to claim 2, characterized in that, The business classification results include a first type, a second type, and a third type; Wherein, the performance requirement feature in the transmission requirement feature corresponding to the first type meets the first preset condition, the connection feature in the transmission requirement feature corresponding to the second type meets the second preset condition, and the transmission requirement feature of the third type does not meet the first preset condition and the second preset condition.

4. The network resource scheduling method according to claim 1, characterized in that, The primary transport protocol type of the service is determined according to a preset priority rule, including: When the service is of the first type, the main transmission protocol type of the service is determined to be the first protocol type; When the service is of the second type, the main transport protocol type of the service is determined to be the second protocol type.

5. The network resource scheduling method according to claim 4, characterized in that, Establish a network connection based on the allocated primary link resources to perform the transmission of service data, including: When the primary transport protocol type of the service is the second protocol type, the data transmission of the service is performed through kernel bypass mode; The kernel bypass mode includes: bypassing the network protocol stack of the operating system kernel, implementing network protocol processing in user space, and using the hardware acceleration function of the network card to transmit data.

6. The network resource scheduling method according to claim 5, characterized in that, Data transmission utilizes the hardware acceleration capabilities of the network interface card, including: The user-space data buffer is directly accessed through the network card's direct memory access engine; Enable the hardware verification and uninstallation function of the network card; The network interface card (NIC) performs data fragmentation and reassembly operations via hardware.

7. The network resource scheduling method according to claim 5, characterized in that, The network resource scheduling method further includes: Obtain the performance metrics of the kernel bypass mode, including the transmission efficiency improvement rate and the CPU utilization rate; Based on the aforementioned performance metrics, the kernel bypass mode is controlled to be enabled or disabled.

8. The network resource scheduling method according to claim 7, characterized in that, Obtaining the transmission efficiency improvement rate of the kernel bypass mode includes: The transmission efficiency improvement rate of the kernel bypass mode is obtained using the first relation, which is: ,in, To improve transmission efficiency, For the transmission latency of the kernel protocol stack, To bypass the transmission delay after the operating system kernel.

9. The network resource scheduling method according to claim 7, characterized in that, Based on the aforementioned performance metrics, controlling the activation or deactivation of the kernel bypass mode includes: If the transmission efficiency improvement rate is lower than the first performance threshold, or the central processing unit utilization rate is higher than the first resource threshold, the kernel bypass mode is turned off. If the transmission efficiency improvement rate is not lower than the second performance threshold and the central processing unit utilization rate is lower than the second resource threshold, the kernel bypass mode remains enabled. Wherein, the first performance threshold is less than or equal to the second performance threshold, and the first resource threshold is greater than or equal to the second resource threshold.

10. The network resource scheduling method according to claim 1, characterized in that, Triggering the switching operation includes: Suspend the transmission of service data through the main link resources; The service data to be sent is converted from the protocol format of the primary transmission protocol type to the protocol format of the backup transmission protocol type; The service data after protocol conversion is sent through the established backup link resources of the backup transmission protocol type.

11. The network resource scheduling method according to claim 10, characterized in that, The network resource scheduling method further includes: After the switching operation is triggered, the recovery status of the primary link resources of the primary transport protocol type is continuously monitored; When it is confirmed that the primary link resource has recovered to a stable state, the service data will be gradually migrated from the backup link resource of the backup transmission protocol type back to the primary link resource of the primary transmission protocol type.

12. The network resource scheduling method according to any one of claims 1-11, characterized in that, The network resource scheduling method further includes: When the service data transmission ends, the primary link resources and backup link resources allocated for the service are released; The released link resources are marked as idle in the merged resource pool.

13. A network resource scheduling device, characterized in that, include: The strategy determination module is used to respond to a service access request, collect the transmission requirement characteristics of the service, classify the service based on the transmission requirement characteristics, obtain the service classification result, and determine the main transmission protocol type of the service based on the service classification result; the transmission requirement characteristics include data volume characteristics, performance requirement characteristics, and connection characteristics, and the main transmission protocol type is RoCE protocol or TCP / IP protocol. The resource allocation module is used to allocate main link resources corresponding to the main transmission protocol type to the service in the converged resource pool. The converged resource pool stores link resources of at least two transmission protocol types. The link resources include at least one of RoCE hardware queues, ports and TCP / IP sockets, and connection handles. The transmission execution module is used to establish a network connection based on the allocated main link resources in order to execute the transmission of business data; Based on the service classification results, the primary transport protocol type of the service is determined, including: If the service classification result is a first type or a second type, the main transmission protocol type of the service is determined according to the preset priority rules; If the service classification result is the third type, obtain the real-time status of the link resources corresponding to multiple transmission protocol types in the fusion resource pool; For each of the aforementioned transport protocol types, the scheduling weight of that transport protocol type is calculated using a second relation, wherein the second relation is: ,in, For link scheduling weights, For the available bandwidth of the link, This represents the total bandwidth of the link. For link load rate, This represents the average link latency. , , All are weighting coefficients. , It is a positive number; The transmission protocol type with the highest scheduling weight is determined as the main transmission protocol type of the service; In the converged resource pool, primary link resources corresponding to the primary transport protocol type are allocated to the service, including: In the fusion resource pool, the link resource with the highest scheduling weight among the multiple link resources under the main transmission protocol type is determined as the main link resource; The network resource scheduling device is also used for: During the process of establishing a network connection based on the allocated main link resources to perform the transmission of business data, fault detection is performed on the main link resources. When a fault is detected and the switching conditions are met, a switching operation is triggered to forward the service data to a backup link resource of the backup transmission protocol type allocated for the service for transmission; the fault includes at least one of link fault, port fault, queue fault, and protocol fault; When a fault is detected and the switching conditions are met, a switching operation is triggered, including: When a fault is detected, the fault risk value is calculated using the third relational formula. When the fault risk value exceeds a preset switching threshold, a switching operation is triggered. The third relation is: ; in, As the first risk factor, As the second risk factor, The third risk factor. The probability of failure occurring. To estimate the fault recovery time, This is the business timeout threshold. For the amount of business data affected, Total business data volume The fault risk value is defined as follows.

14. A network device, characterized in that, include: The communication unit is used to receive service access requests and to perform service data transmission based on the established network connection. A storage unit is used to store a converged resource pool containing link resources of at least two transport protocol types; the link resources include at least one of RoCE hardware queues, ports and TCP / IP sockets, and connection handles; The processing unit, connected to both the communication unit and the storage unit, is used for: In response to a service access request received by the communication unit, the transmission requirement characteristics of the service are collected, the service is classified based on the transmission requirement characteristics to obtain a service classification result, and the main transmission protocol type of the service is determined based on the service classification result. The transmission requirements include data volume characteristics, performance requirements characteristics, and connection characteristics, and the main transmission protocol type is RoCE protocol or TCP / IP protocol; From the converged resource pool of the storage unit, allocate main link resources corresponding to the main transport protocol type for the service; The control unit establishes a network connection based on the allocated main link resources to perform the transmission of service data; Based on the service classification results, the primary transport protocol type of the service is determined, including: If the service classification result is a first type or a second type, the main transmission protocol type of the service is determined according to the preset priority rules; If the service classification result is the third type, obtain the real-time status of the link resources corresponding to multiple transmission protocol types in the fusion resource pool; For each of the aforementioned transport protocol types, the scheduling weight of that transport protocol type is calculated using a second relation, wherein the second relation is: ,in, For link scheduling weights, For the available bandwidth of the link, This represents the total bandwidth of the link. For link load rate, This represents the average link latency. , , All are weighting coefficients. , It is a positive number; The transmission protocol type with the highest scheduling weight is determined as the main transmission protocol type of the service; The network device is also used for: During the process of establishing a network connection based on the allocated main link resources to perform the transmission of business data, fault detection is performed on the main link resources. When a fault is detected and the switching conditions are met, a switching operation is triggered to forward the service data to a backup link resource of the backup transmission protocol type allocated for the service for transmission; the fault includes at least one of link fault, port fault, queue fault, and protocol fault; When a fault is detected and the switching conditions are met, a switching operation is triggered, including: When a fault is detected, the fault risk value is calculated using the third relational formula. When the fault risk value exceeds a preset switching threshold, a switching operation is triggered. The third relation is: ; in, As the first risk factor, As the second risk factor, The third risk factor. The probability of failure occurring. To estimate the fault recovery time, This is the business timeout threshold. For the amount of business data affected, Total business data volume The fault risk value is defined as follows.

Citation Information

Patent Citations

  • Flow congestion control method and device, computer readable medium and electronic equipment

    CN114884823A

  • High and low orbit satellite communication resource scheduling method and system

    CN120979520A