Fault locating method and system, device across ib and roce network

By standardizing and unifying the category mapping of the original metrics of IB and RoCE networks, and combining static and dynamic analysis, the problems of low efficiency and poor accuracy in cross-network fault location in existing technologies are solved, and rapid and accurate fault source identification is achieved.

CN121690994BActive Publication Date: 2026-07-03BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
Filing Date
2025-12-17
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing monitoring systems cannot provide a global health view across IB and RoCE networks, resulting in long fault analysis cycles, poor accuracy, and difficulty in quickly locating the source of the fault.

Method used

By acquiring the original metrics of the IB and RoCE networks, standardizing them, and then performing unified category mapping, combined with static and dynamic analysis, static fault sources and trend anomalies are identified, and the fault source is determined comprehensively.

Benefits of technology

It enables rapid and accurate location of network faults, improves the efficiency and accuracy of fault location, and ensures the stable operation of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121690994B_ABST
    Figure CN121690994B_ABST
Patent Text Reader

Abstract

This application discloses a method, system, and apparatus for fault location across IB and RoCE networks. The method includes: acquiring original indicators of IB and RoCE network classes in the IB and RoCE network environments; determining unified physical network category information based on fault location analysis requirements and the target scenario; mapping the standardized original indicators of IB and RoCE network classes to the unified category information to obtain IB-dimensional and RoCE-dimensional indicator mapping information; filtering out unsatisfactory original indicators using preset conditions and determining static fault sources for the indicator items using the IB / RoCE-dimensional mapping; analyzing indicators within a preset period to identify trend-based abnormal indicator information and determine the corresponding dynamic fault sources for the indicator items; and combining static and dynamic fault sources to finally determine the comprehensive fault source for the indicator items. This method enables comprehensive and accurate fault location across IB and RoCE networks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of network technology, and in particular to a fault location method, system, and apparatus for cross-IB and RoCE networks. Background Technology

[0002] With the rapid development of artificial intelligence, especially large model training and inference, AI cloud platforms have become the core infrastructure for carrying computing power. In order to meet the communication requirements of distributed training for high bandwidth, low latency and high reliability, current AI platforms generally deploy two types of high-performance networks at the same time: IB (InfiniBand) and RoCE (RDMA over Converged Ethernet).

[0003] Among them, IB provides port and link metrics, such as port data transmission and reception volume, link errors, and latency statistics, through the subnet manager and performance management interface; RoCE, on the other hand, obtains information including traffic statistics, packet loss rate, and PFC / ECN congestion control information based on the monitoring interface provided by the Ethernet switch and network card driver.

[0004] Although both types of networks can output performance data, they differ significantly in terms of indicator systems, naming conventions, statistical granularity, and interface protocols. Existing monitoring systems can only display these separately and cannot provide a global health view. When network bottlenecks or failures occur, maintenance personnel can only analyze IB and RoCE indicators separately, leading to problems such as long analysis cycles and inaccurate fault location. Summary of the Invention

[0005] In view of this, the present disclosure provides a fault location method, system, and apparatus for cross-IB and RoCE networks, which can solve the problems of low efficiency, poor accuracy, and time-consuming and labor-intensive fault information location when an AI platform with both IB and RoCE high-performance networks is deployed simultaneously fails.

[0006] In a first aspect, embodiments of this disclosure provide a fault location method across IB and RoCE networks, comprising:

[0007] Obtain the raw metrics of the IB network class in the IB network environment and the raw metrics of the RoCE network class in the RoCE network environment.

[0008] Based on the fault location analysis requirements and target scenario, a unified category information for the physical network is determined. The unified category information for the physical network includes several indicator types and several indicator items corresponding to each of the indicator types.

[0009] The standardized original indicators of the IB network class and the original indicators of the RoCE network class are mapped to the unified category information of the physical network to obtain the IB dimension indicator item mapping information and the RoCE dimension indicator item mapping information.

[0010] Acquire the instantaneous real-time collected original indicators of the IB network class and all indicators in the original indicators of the RoCE network class that do not meet the preset conditions, and determine the static fault source of the indicator item based on the mapping information of the IB dimension indicator item and the mapping information of the RoCE dimension indicator item.

[0011] Analyze the original indicators of the IB network class and the RoCE network class within the preset period of collection to obtain trend anomaly indicator information;

[0012] Based on the IB dimension indicator item mapping information and the RoCE dimension indicator item mapping information, determine the dynamic fault source of the indicator item corresponding to the trend anomaly indicator information;

[0013] Based on the static fault sources and dynamic fault sources of the aforementioned indicator items, the comprehensive fault source of the indicator item is determined.

[0014] Secondly, embodiments of this disclosure also provide a fault location system across IB and RoCE networks, comprising:

[0015] The raw metrics acquisition module is used to acquire raw metrics of IB network classes in the IB network environment and raw metrics of RoCE network classes in the RoCE network environment.

[0016] A unified module is used to determine unified category information of physical networks based on fault location analysis requirements and target scenarios. The unified category information of physical networks includes several indicator types and several indicator items corresponding to each indicator type.

[0017] The mapping module is used to map the standardized original indicators of the IB network class and the original indicators of the RoCE network class to the unified category information of the physical network, respectively, to obtain IB dimension indicator item mapping information and RoCE dimension indicator item mapping information.

[0018] The indicator item static fault source acquisition module is used to acquire all indicators in the IB network class original indicators and the RoCE network class original indicators that do not meet the preset conditions in real time, and determine the indicator item static fault source according to the IB dimension indicator item mapping information and the RoCE dimension indicator item mapping information.

[0019] The indicator item dynamic fault source acquisition module is used to analyze the original indicators of the IB network class and the original indicators of the RoCE network class within a preset period to obtain trend anomaly indicator information; and to determine the indicator item dynamic fault source corresponding to the trend anomaly indicator information based on the IB dimension indicator item mapping information and the RoCE dimension indicator item mapping information.

[0020] The fault source location module is used to determine the comprehensive fault source of the indicator item based on the static fault source and the dynamic fault source of the indicator item.

[0021] Thirdly, this disclosure also provides a computer device, which adopts the following technical solution:

[0022] The computer device includes:

[0023] At least one processor; and,

[0024] A memory communicatively connected to the at least one processor; wherein,

[0025] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the fault location methods described above across IB and RoCE networks.

[0026] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing computer instructions for causing a computer to perform any of the fault location methods across IB and RoCE networks described above.

[0027] Fifthly, embodiments of this disclosure also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.

[0028] The fault location method proposed in this application for cross-IB and RoCE networks first determines the switch port indicators of the IB network and the network monitoring indicators of the RoCE network, and performs standardization processing on them. Then, it determines the unified category information of the physical network and maps the two types of indicators to it. Next, it uses static analysis to obtain instantaneous abnormal indicators to identify static fault sources, and uses time-series analysis to obtain trend-based abnormal information within a preset period to identify dynamic fault sources. Finally, it combines both methods to determine the comprehensive fault source. This solution solves the problems of existing monitoring systems not being able to provide a global health view, having long analysis cycles, and inaccurate fault source location. By eliminating the differences between the two types of network indicators through standardization processing and indicator mapping, and combining static and dynamic analysis, it comprehensively and accurately locates faults, improving the efficiency and accuracy of network fault location and ensuring the stable operation of the network.

[0029] The above description is merely an overview of the technical solution disclosed herein. In order to better understand the technical means of this disclosure and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0030] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 This is a flowchart illustrating the fault location method across IB and RoCE networks provided in an embodiment of this disclosure.

[0032] Figure 2 This is a flowchart illustrating the method for obtaining trend anomaly indicator information provided in an embodiment of this disclosure.

[0033] Figure 3 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. Detailed Implementation

[0034] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0035] It should be understood that the following specific examples illustrate the implementation of this disclosure, and those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0036] Reference Figure 1 This application discloses a fault location method for cross-IB and RoCE networks, specifically a fault location method for network systems (such as AI cloud platforms, financial trading systems, scientific computing clusters, etc.) that simultaneously deploy both IB (InfiniBand) and RoCE (RDMA over Converged Ethernet) high-performance networks. The method includes:

[0037] S100: Obtain the raw metrics of the IB network class in the IB network environment and the raw metrics of the RoCE network class in the RoCE network environment.

[0038] This step ensures that all potentially relevant raw data is collected from both the IB and RoCE high-performance network environments.

[0039] S200 determines the unified category information of the physical network based on the fault location analysis requirements and the target scenario. The unified category information of the physical network includes several indicator types and several indicator items corresponding to each indicator type.

[0040] Among them, the physical network unified category corresponds to the commonalities of IB and RoCE in the underlying network behavior.

[0041] S300 maps the standardized IB network class original indicators and RoCE network class original indicators to the physical network unified category information to obtain IB dimension indicator item mapping information and RoCE dimension indicator item mapping information.

[0042] This step gives the original, potentially ambiguous, metrics more specific and domain-relevant meanings as "metric items," creating dimensional mapping information for both IB and RoCE. This ensures that key metrics can be uniformly identified and processed even across different devices and versions of the IB / RoCE protocol, improving the comparability of the analysis.

[0043] S400 acquires all indicators in the instantaneous real-time collected IB network category raw indicators and RoCE network category raw indicators that do not meet the preset conditions, and determines the static fault source of the indicator item based on the IB dimension indicator item mapping information and the RoCE dimension indicator item mapping information.

[0044] This step can quickly capture explicit, immediate anomalies. By setting preset conditions (thresholds, ranges), abnormal raw indicators can be identified immediately. Raw indicators exceeding the threshold are directly converted into static fault sources through mapping, representing a clear, static problem state in the current network. This serves as a preliminary fault indication, avoiding complex time-series analysis of all indicators. Instead, it directly filters out "problematic" indicators, focusing on subsequent root cause analysis.

[0045] The S500 analyzes the raw indicators of IB network type and RoCE network type within a preset period to obtain information on trend anomalies.

[0046] This step focuses on the evolution of indicators over a period of time, rather than the instantaneous value at a single point in time; static fault sources only see the "effect", while static abnormal indicator information has the opportunity to see the evolution of the "cause", providing richer contextual information.

[0047] S600 determines the dynamic fault source of the indicator item corresponding to the trend anomaly indicator information based on the IB dimension indicator item mapping information and the RoCE dimension indicator item mapping information.

[0048] Dynamic fault sources are an indispensable part of determining the root cause, as they help distinguish between transient problems and persistent, structural problems.

[0049] S700 determines the comprehensive fault source of the indicator item based on the static fault source and the dynamic fault source of the indicator item.

[0050] This step integrates the previously identified information from two dimensions: "what is obviously bad" (static fault source) and "why it is bad / how it is deteriorating" (dynamic fault source). Static fault sources are usually symptoms, dynamic fault sources describe the evolution of symptoms, and the combined fault source attempts to infer the root cause of all this. Relying solely on static fault sources may lead to misdiagnosis (for example, a brief peak may not represent the underlying problem), while a single dynamic fault source may be difficult to delimit. Combining the two can improve the accuracy of fault location and the robustness of diagnosis.

[0051] The method disclosed in this application, through layered and step-by-step processing, covers both IB and RoCE networks, from raw data to the final comprehensive fault source, taking into account both static explicit faults and dynamic trend problems. Through mapping and subsequent analysis, it transforms raw indicators into meaningful "fault source" information; combining static and dynamic information reduces misjudgments and gets closer to the root cause of the fault; it quickly locates explicit faults and provides early warnings and analysis of trend anomalies; the output "comprehensive fault source" directly guides maintenance personnel in repair; and the design of unified category information for physical networks gives the method a certain degree of scenario adaptability and cross-IB / RoCE universality. This application constructs a complete closed loop from data collection, feature extraction, fault mode identification, and root cause inference, representing an excellent framework for handling complex high-performance network faults.

[0052] For S100, this specifically includes: 1) collecting raw IB network metrics through the Subnet Manager and Performance Management Agent (PMA) interfaces; 2) collecting raw RoCE network metrics through the Ethernet switch's Telemetry / SNMP, interfaces, RoCE, network card drivers (such as rdma-core, ethtool), or the operating system performance counters on the GPU node side.

[0053] Among them, the original indicators of IB network category include switch port indicators, which include one or more of the following: port transmit / receive bytes, port transmit / receive packets, link error count, and latency statistics.

[0054] RoCE network-related raw metrics include network monitoring metrics, which include one or more of the following: traffic-related metrics, error and packet loss metrics, congestion control metrics, and resource utilization metrics.

[0055] The method for S200 to "determine the unified category information of the physical network based on the fault location analysis requirements" specifically includes:

[0056] S210, Obtain all collectable field indicator types in the scenario of the fault to be analyzed;

[0057] S220 compares the fault location analysis requirements with the field indicator types. If the field indicator types include the fault location analysis requirements, the unified category information of the physical network is determined based on the fault location analysis requirements.

[0058] If the on-site indicator types do not all include the fault location analysis requirements, determine the unified category information of the physical network based on the on-site indicator types.

[0059] The unified category information for physical networks includes several indicator types and several indicator items corresponding to each indicator type; the indicator types include throughput, error, latency and link status, and the indicator items for different types are different.

[0060] In this embodiment, the unified category information for physical networks establishes a common system for IB and RoCE, two networks that differ in their underlying implementation. The phrase "determined based on fault location analysis requirements and target scenarios" emphasizes that this classification is not rigid but can be adjusted according to specific application scenarios (such as AI training focusing on throughput, or financial transactions focusing on latency), thereby focusing on the most relevant indicator types and items. This unified category information is the key basis for mapping in S300. It defines "what indicator attributes need attention," providing a framework for the subsequent generation and classification of the concept of "indicator items."

[0061] Specifically, the throughput category includes metrics such as: number of bytes sent by the port, number of bytes received by the port, number of packets sent by the port, number of packets received by the port, and link bandwidth capacity. The error category includes metrics such as: number of symbol errors, CRC check errors, number of receive errors, number of transmit errors, number of link outages, and link integrity errors. The latency category includes metrics such as: average latency. The link status category includes metrics such as: link rate.

[0062] The original metrics for IB networks include:

[0063] PortXmitData, PortRcvData, PortXmitPkts, PortRcvPkts, LinkSpeed×LinkWidth, SymbolErrorCounter, PortRcvRemotePhysicalErrors, PortRcvErrors, PortXmitConstraintErrors, LinkDownedCounter, LocalLinkIntegrityErrors, LatencyCounters, LinkSpeed.

[0064] RoCE network raw metrics include: TxBytes, RxBytes, TxPackets, RxPackets, LinkSpeed, Symbol Errors, CRC Errors, InErrors / InDiscards, OutErrors / OutDiscards, LinkDown Events / Port Resets, Alignment Errors, ECN RTT Estimation, and LinkSpeed.

[0065] Table 1 Mapping Information of IB Dimension Indicator Items

[0066]

[0067] Referring to Table 1, the original IB network metrics for throughput include: PortXmitData, PortRcvData, PortXmitPkts, PortRcvPkts, and LinkSpeed×LinkWidth. After mapping, the unified throughput metrics are: TxBytes (number of bytes sent by the port), RxBytes (number of bytes received by the port), TxPackets (number of packets sent by the port), RxPackets (number of packets received by the port), and Throughput (link bandwidth capacity).

[0068] The original IB network error categories include: SymbolErrorCounter, PortRcvRemotePhysicalErrors, PortRcvErrors, PortXmitConstraintErrors, LinkDownedCounter, and LocalLinkIntegrityErrors. After mapping, the unified error category metrics are: SymbolErrors (number of symbol errors), CRCErrors (CRC checksum errors), RxErrors (number of receive errors), TxErrors (number of transmit errors), LinkDownEvents (number of link interruptions), and IntegrityErrors (link integrity errors).

[0069] The original IB network metrics for latency include Latency Counters; after mapping, the unified latency metrics are LatencyAvg (average latency).

[0070] The original IB network class metrics for the link-state class include: LinkSpeed; after mapping, the unified link-state class metrics are: LinkSpeed ​​(link speed).

[0071] Through the mapping process in this embodiment, the original indicators (i.e., underlying indicators) of the IB network class can be abstracted into several specific unified indicator items corresponding to unified categories.

[0072] Table 2 RoCE Dimension Indicator Mapping Information

[0073]

[0074] Referring to Table 2, the original RoCE network metrics for throughput include: TxBytes, RxBytes, TxPackets, RxPackets, and LinkSpeed. After mapping, the unified throughput metrics are: TxBytes (number of bytes sent by the port), RxBytes (number of bytes received by the port), TxPackets (number of packets sent by the port), RxPackets (number of packets received by the port), and Throughput (link bandwidth capacity).

[0075] The original RoCE network error categories include: Symbol Errors, CRC Errors, InErrors / InDiscards, OutErrors / OutDiscards, Link Down Events / Port Resets, and AlignmentErrors. After mapping, the unified error category metrics are: SymbolErrors (number of symbol errors), CRCErrors (CRC checksum errors), RxErrors (number of receive errors), TxErrors (number of transmit errors), LinkDownEvents (number of link interruptions), and IntegrityErrors (link integrity errors).

[0076] The original metrics for RoCE networks in the latency category include: ECN RTT Estimation; after mapping, the unified latency category metrics are: LatencyAvg (average latency).

[0077] The original metrics for the RoCE network class in the link-state category include: LinkSpeed; after mapping, the metrics for the unified link-state category are: LinkSpeed ​​(link rate).

[0078] Through the mapping process in this embodiment, the original indicators (i.e., underlying indicators) of the RoCE network class can be abstracted into several specific unified indicator items corresponding to unified categories.

[0079] The standardization process for both IB and RoCE network raw metrics includes: 1) timestamp alignment (uniform format processing); and 2) unit conversion. Specifically, the timestamp information of both IB and RoCE network raw metrics is obtained, and the obtained timestamp information is uniformly parsed and converted into a standardized universal timestamp format that includes UTC (Coordinated Universal Time). For metrics representing the same physical quantity (e.g., bandwidth, throughput, latency, packet loss rate, etc.) in both IB and RoCE network raw metrics, their original units of measurement are identified. The identified original units of measurement are then converted to preset, comparable universal units of measurement.

[0080] By standardizing the raw metrics of InfiniBand (IB) and RoCE (RDMA over Converged Ethernet) networks, we ensure that all data points are based on the same time base (UTC). This prevents misjudgments caused by time zone differences or varying recording precision across different data sources when analyzing network behavior over a specific time period. For example, when a network congestion event occurs, we can accurately see how various relevant IB and RoCE metrics change simultaneously, thus pinpointing the root cause. Furthermore, we can achieve joint analysis and comparison across network technologies. After standardization, IB and RoCE metrics can be treated and analyzed equally within the same dataset, facilitating the construction of a unified network monitoring and analysis platform. This allows for a comprehensive understanding of the health status of the entire heterogeneous network and direct comparison of IB and RoCE performance (such as latency, throughput, and packet loss rate) under the same load and time period, thereby assessing which technology is superior in a specific scenario or identifying performance bottlenecks. In addition, it can simplify and automate the calculation of advanced metrics. Many advanced and comprehensive network performance metrics (such as application layer throughput, end-to-end latency, link utilization, and congestion indicators) require the fusion calculation of multiple basic metrics from different layers and technologies.

[0081] For S400's "acquiring all indicators that do not meet preset conditions among the instantaneously collected IB network-type raw indicators and RoCE network-type raw indicators, and recording them as static abnormal indicators," specifically, this step refers to using a single-point snapshot method to collect key indicators of all nodes or ports at a certain moment, and comparing them once against normal thresholds or healthy baselines to quickly determine whether there is an explicit fault. At time t, all port indicators are captured and compared with standard thresholds (such as CRCErrors>10 / min, Throughput<80%LinkSpeed, LatencyAvg>50µs). If the threshold is exceeded or an abnormal value appears (such as SymbolErrors>0), the port is immediately determined to have a problem. Then, based on all static abnormal indicators and the mapping information of IB dimension indicator items and RoCE dimension indicator items, the static fault source (explicit fault) of the indicator item is determined, that is, the specific unified indicator item corresponding to the raw indicator (i.e., the underlying indicator) is determined.

[0082] For example, if PortXmitData=0 and PortRcvData=0 in the original IB network metrics, it means that these two original metrics are static abnormal metrics. The corresponding unified metrics are TxBytes=0 and RxBytes=0, which means that the static fault source of the metrics is that the port is not working / the process is not communicating.

[0083] In the RoCE network class raw metrics, TxBytes=0 and RxBytes=0, it is indicated that these two raw metrics are static abnormal metrics. The corresponding unified metrics are TxBytes=0 and RxBytes=0, indicating that the static fault source of the metrics is that the port is not working / the process is not communicating.

[0084] If the SymbolErrorCounter, PortRcvRemotePhysicalErrors, and LocalLinkIntegrityErrors in the IB network class raw indicators are non-zero, it indicates that these three raw indicators are static abnormal indicators. The corresponding unified indicator items, SymbolErrors, CRCErrors, and IntegrityErrors, are non-zero, indicating that the static fault source of the indicator items is an optical module, cable, or physical layer fault.

[0085] If the Symbol Errors, CRC Errors, and Alignment Errors in the RoCE network class raw indicators are non-zero, it indicates that these three raw indicators are static abnormal indicators. The corresponding unified indicator items, SymbolErrors, CRCErrors, and IntegrityErrors, are non-zero, indicating that the static fault source of the indicator items is an optical module, cable, or physical layer fault.

[0086] If SymbolErrors, CRC Errors, and IntegrityErrors are not zero, it indicates a fault in the optical module, cable, or physical layer.

[0087] An increase in LinkDownedCounter in the original IB network metrics indicates that this original metric is a static anomaly metric. The corresponding unified metric item, LinkDownEvents, also increases, indicating that the static fault source of this metric item is a link disconnection or port reset.

[0088] An increase in Link Down Events / Port Resets in the original RoCE network metrics indicates that the original metrics are static anomalies. The corresponding unified metric item, LinkDownEvents, also shows an increase, indicating that the static fault source for the metric item is a link disconnection or port reset.

[0089] The fact that LinkSpeed ​​× LinkWidth ≪ LinkSpeed ​​in the original IB network metrics indicates that these two original metrics are static abnormal metrics. The corresponding unified metric item is Throughput ≪ LinkSpeed, indicating that the static fault source of the metric item is port speed reduction or rate negotiation failure.

[0090] An anomaly in the LinkSpeed ​​indicator of the IB network category indicates that the original indicator is a static anomaly. The corresponding unified indicator item is also an anomaly in LinkSpeed, indicating that the static fault source of the indicator item is an incorrect link rate configuration.

[0091] The LinkSpeed ​​anomaly in the RoCE network class raw metrics indicates that the raw metric is a static anomaly. The corresponding unified metric item is LinkSpeed ​​anomaly, indicating that the static fault source of the metric item is a link rate configuration error.

[0092] If the Latency Counters in the IB network category raw metrics is extremely high (>100µs), it indicates that the raw metrics are static anomalies. The corresponding unified metrics item, LatencyAvg, is extremely high (>100µs), indicating that the static fault source of the metrics item is congestion or routing loops.

[0093] In the RoCE network class raw metrics, ECN RTT Estimation is extremely high (>100µs), indicating that this raw metric is a static anomaly. The corresponding unified metric item, LatencyAvg, is extremely high (>100µs), indicating that the static fault source of the metric item is congestion or routing loop.

[0094] Reference Figure 2 The method for S500 to "analyze the original indicators of IB network type and RoCE network type within a preset period to obtain trend anomaly indicator information" specifically includes the following methods for obtaining trend anomaly indicator information:

[0095] S510, determine the preset cycle;

[0096] S520 divides the preset period into several sub-segments;

[0097] S530: Obtain the sub-segments of the original indicators of the same IB network class within the preset period that deviate from the preset normal range, and record them as IB network class abnormal segments.

[0098] S540, obtain consecutive IB network class anomaly segments. If the number of consecutive IB network class anomaly segments exceeds the preset number threshold, determine the corresponding IB network class original index as the IB network class trend anomaly index.

[0099] S550: Obtain the sub-segments of the original indicators of the same RoCE network class within the preset period that deviate from the preset normal range, and record them as the RoCE network class abnormal segments.

[0100] S560: Obtain consecutive RoCE network anomaly segments. If the number of consecutive RoCE network anomaly segments exceeds a preset threshold, determine the corresponding RoCE network original index as a RoCE network trend anomaly index.

[0101] Each preset normal range is matched with the corresponding original indicator.

[0102] Specifically, the setting of the preset normal range can be based on either statistical methods or business rules. Statistical methods include calculating the mean μ and standard deviation σ of historical data for each network bandwidth metric. The preset normal range can then be set to [μ−kσ,μ+kσ], where k is a constant (common values ​​are 2 or 3). This range covers most data values ​​under normal conditions. Business rule-based methods involve determining the normal range based on actual business needs and experience. For example, if a system's response latency is required to be no more than 100 milliseconds, the normal range can be set to [0,100] milliseconds.

[0103] In this embodiment, "continuous" means that within a fixed time window, the value of the indicator repeatedly exceeds the normal range, and this exceedance is not a random, single occurrence. Taking the collection of time-series data from the most recent hour as an example, suppose this hour is divided into multiple smaller time segments (e.g., one segment per minute), with each time segment corresponding to the current value of the indicator. If, in multiple consecutive time segments, the indicator value falls outside the normal range, this is called continuous deviation. For example, if network bandwidth data is collected every minute for the last 10 minutes, with a normal range of [100Mbps, 200Mbps], and if the bandwidth data collected for five consecutive minutes is greater than 200Mbps or less than 100Mbps, then the condition for continuous deviation is met.

[0104] When an indicator deviates continuously from the normal range, it indicates that the change in the indicator is not a random fluctuation, but rather a continuous trend that differs from the normal situation. This trend may indicate a potential problem in the system.

[0105] The long-term decline in PortXmitData (port sent data) and PortRcvData (port received data) in the IB network category raw metrics indicates that these two raw metrics are trend anomalies. The corresponding unified metrics, TxBytes (number of bytes sent) and RxBytes (number of bytes received), have also been declining for a long time. This indicates that the overall network traffic is decreasing significantly. The corresponding dynamic fault sources for these metrics include one or more of the following: reduced traffic to upper-layer applications / services, other bottlenecks / faults on the network path, processing capacity issues at the target end (receiving end), congestion at the protocol stack / RDMA protocol layer, network device configuration changes, and performance issues of the node / server itself.

[0106] The long-term decline in TxBytes and RxBytes in the original RoCE network metrics indicates that these two original metrics are trend anomalies. The corresponding unified metrics are also long-term decline in TxBytes and RxBytes, indicating link bandwidth degradation and load imbalance. The corresponding dynamic fault sources include one or more of the following: a significant reduction in activity of upstream applications / data sources, bottlenecks or faults in specific paths in the RoCE network, and configuration errors or faults in RoCE network devices (switches, network cards).

[0107] The continuous increase in PortRcvRemotePhysicalErrors, PortRcvErrors, and PortXmitConstraintErrors in the original IB network metrics indicates that these three original metrics are trend-based anomalies. The continuous increase in the corresponding unified metrics CRCErrors, RxErrors, and TxErrors indicates that the dynamic fault sources of these metrics include one or more of the following: physical layer connection problems / media aging, IB / RoCENIC (network interface card) hardware failure or performance degradation, IB / RoCE switch hardware failure or port problems, and unstable power supply / environmental problems.

[0108] The continuous increase in CRC Errors, InErrors / InDiscards, and OutErrors / OutDiscards in the original RoCE network metrics indicates that these three original metrics are trend-abnormal indicators. The corresponding unified metrics, CRCErrors, RxErrors, and TxErrors, also show continuous increases, indicating optical attenuation / interference bit errors. The corresponding dynamic fault sources include one or more of the following: quality degradation of the fiber optic link (cable, connector, splice), aging or failure of the transceiver, and external electromagnetic / radio frequency interference.

[0109] If the Latency Counters in the IB network category raw indicators shows an upward trend or increased volatility, it indicates that the raw indicator is an abnormal trend indicator. The corresponding unified indicator item is LatencyAvg, which shows an upward trend or increased volatility. The corresponding dynamic fault sources include one or more of the following: network congestion (links, switches, NIC ports), IB / RoCENIC processing bottlenecks (CPU, its own performance), increased processing latency of IB / RoCE switches, and performance problems of RDMA applications or protocol stacks.

[0110] The increasing trend / increased volatility of ECN RTT Estimation in the original RoCE network metrics indicates that this original metric is an anomaly. The corresponding unified metric, LatencyAvg, also shows an increasing trend / increased volatility, indicating that the dynamic fault sources for congestion propagation and routing imbalance include congestion propagation and / or routing imbalance. Congestion propagation includes global or regional congestion caused by cascading effects, the propagation and spread of bursty traffic in the network, and instability in the feedback loop of congestion control mechanisms such as ECN / PFC, leading to significant fluctuations in traffic rates. Routing imbalance includes uneven traffic distribution caused by the ECMP load balancing algorithm, creating "hotspot" links or ports, incorrect static or dynamic routing configurations that direct traffic to inefficient paths, and failure to adjust routes in a timely manner after network topology changes, resulting in traffic concentration.

[0111] If the Latency Counters in the IB network category raw indicators shows an upward trend or increased volatility, it indicates that the raw indicator is an abnormal trend indicator. The corresponding unified indicator item is LatencyAvg, which shows an upward trend or increased volatility. The corresponding indicator item dynamic faults include dynamic congestion propagation and chain reaction, oscillations introduced by the debugging of congestion control mechanisms, uneven traffic distribution and path competition, dynamic inappropriateness in QoS policy execution, fluctuations in internal processing latency of the switch, and one or more of the following: dynamic processing bottlenecks of terminal NIC or host CPU.

[0112] If the problems include frequent changes in LinkSpeed, it indicates that the automatic speed-down / self-healing is unstable; if the multi-node latency variance continues to increase, it indicates that the AllReduce latency is uneven and the training synchronization is abnormal.

[0113] For S700, "determine the comprehensive fault source of the indicator item based on the static fault source and the dynamic fault source of the indicator item", it can include:

[0114] If the static fault source of the indicator is a data transmission symbol error (i.e., SymbolErrors is non-zero) or a data integrity error (i.e., IntegrityErrors is non-zero), and the dynamic fault source of the indicator is a continuous increase in the static fault source of the indicator and a decrease in throughput (i.e., Throughput decreases), then the comprehensive fault source of the indicator is determined to be a physical layer fault, and the corresponding fault type is optical module damage.

[0115] In digital communication, data is transmitted in units of symbols (such as one bit, two bits, etc.). A non-zero SymbolErrors value indicates that the received symbol is inconsistent with the sent symbol during transmission, resulting in a symbol recognition error. This is usually caused by physical layer signal quality issues. A continuous increase means that the number of SymbolErrors or IntegrityErrors increases steadily over time, rather than being a one-off event. This indicates that the problem is ongoing and may even be worsening.

[0116] Computer networks are typically divided into multiple layers, with the physical layer being the lowest layer, responsible for the actual signal transmission (electrical signals, optical signals, etc.). The occurrence of SymbolErrors and IntegrityErrors, along with a decrease in throughput, strongly points to a problem with the quality of the transmitted signal, which is a typical manifestation of a physical layer failure. By combining the phenomena of "errors appearing immediately" and "errors worsening over time and overall transmission efficiency decreasing," the root cause of this abnormal metric lies in the physical layer of network communication. An optical module is hardware that converts electrical signals into optical signals (transmitting) or vice versa. If an optical module is damaged, it will directly cause errors in the conversion between electrical and optical signals, generating SymbolErrors and IntegrityErrors, and severely affecting the transmission rate, leading to a decrease in throughput. This step refers to, after ruling out other possible physical layer causes (such as fiber optic cable damage, connector problems, etc.), or further inferring based on data patterns (e.g., these error patterns resemble characteristics unique to optical modules), that the most likely specific hardware failure causing this physical layer failure is the optical module.

[0117] If the static fault source of the metric is a single LatencyAvg peak, and the dynamic fault source is periodic latency fluctuations and TxErrors accumulation, the overall fault source of the metric is determined to be RoCE congestion path propagation, and the corresponding fault type is congestion propagation. RoCE networks rely on Ethernet infrastructure and typically use mechanisms such as PFC (Priority Flow Control) and DCQCN (Data Center Quantized Congestion Notification) to manage congestion. This example shows that the problem occurs directly at the RoCE protocol level and is reflected in its transmission performance. The problem does not originate from a permanent failure of a single device or link, but rather from congestion signals propagating from one node to another in the network. This can cause periodic increases and fluctuations in latency throughout the entire path or area through feedback loops (such as rate adjustments in DCQCN). For example, congestion on a switch port may cause the NIC of a connected host to stop transmitting via PFC signals, thus affecting other applications connected to that host. This propagation effect causes periodic fluctuations in LatencyAvg, and the accumulation of TxErrors is due to signal quality degradation or packet loss retransmission caused by congestion.

[0118] If the static fault source of the indicator is a significantly low node TxBytes (data transmission volume), and the dynamic fault source of the indicator is an increase in the bandwidth variance of each node, then the comprehensive fault source of the indicator is determined to be GPU communication load imbalance, and the corresponding fault type is node imbalance.

[0119] An increase in variance indicates a widening difference in bandwidth between nodes. In other words, some nodes may still be communicating normally, while the communication volume of other nodes (consistent with low TxBytes in static fault sources) drops sharply or becomes very unstable. A significantly low TxBytes (static) and an increase in bandwidth variance (dynamic, indicating that this imbalance is persistent and is intensifying) clearly point to significant differences in communication activities between nodes.

[0120] If the static fault source of the indicator is no significant static error, and the dynamic fault source of the indicator is a localized persistently high Throughput and an increase in LatencyAvg, then the comprehensive fault source of the indicator is determined to be a local link hotspot, and the corresponding fault type is a topology hotspot.

[0121] In this context, a local link refers to a specific physical or logical link in the network, a switch port, or a switch "backplane" connecting multiple devices. A hotspot refers to an area carrying far more traffic than its design or current operating capacity (especially relative to other links). This situation typically occurs in: 1) Uneven traffic patterns: Traffic generated by a specific task or application is concentrated entirely (or mostly) through one or a few links / ports; 2) Uneven ECMP load balancing: Even with ECMP, the design of hash functions or the characteristics of traffic patterns can lead to abnormally high traffic on some paths; 3) Aggregated connections within switch ports: Excessive traffic pressure on uplinks or downlinks; 4) Bandwidth mismatch between the server NIC and the switch port. Because the bandwidth and buffer of this "hotspot" link are heavily occupied, all packets passing through this area will experience longer queuing times, even if there may be no direct packet loss or errors, resulting in increased LatencyAvg. This situation is not a general congestion problem, but is closely related to the physical / logical topology of the network. The problem lies in specific "nodes" (links / ports / switches) because their location and connection relationships in the overall topology lead to the concentration of traffic.

[0122] If the static fault source of the indicator is a sudden increase in latency, and the dynamic fault source of the indicator is an increase in the average latency of the continuous window, the corresponding fault type is abnormal training latency. The corresponding comprehensive fault sources of the indicator include one or more of the following: persistent RoCE congestion, RDMA protocol stack processing bottleneck, continuous high load caused by training tasks, and insufficient network resource reservation / allocation.

[0123] The fault location method for cross-IB and RoCE networks disclosed in this application further includes: acquiring historical target data within a preset period, wherein the historical target data includes static fault sources and dynamic fault sources of indicator items corresponding to all IB network classes and RoCE network classes;

[0124] Based on the historical target data, obtain the fault association information between the IB network class and the RoCE network class;

[0125] The large model is trained based on the fault association information;

[0126] The newly emerging faults are analyzed based on the trained large model to obtain fault location results.

[0127] Traditional fault diagnosis relies on manually written, fixed rules that are experience-based and potentially static. Large-scale models, however, can learn hidden, non-linear, and subtle patterns of correlation within data, patterns that exceed the intuition or rule-writing abilities of human experts. These models can learn from massive amounts of historical data that, under specific conditions (e.g., a sudden increase in load accompanied by TxErrors), the most common root cause is "RoCE congestion path propagation." They can identify combined characteristics of static and dynamic fault sources, not just single features. Large-scale models can understand how dynamic and static anomalies of different metrics combine to form more complex fault scenarios. Once the model is trained, it can automatically input data for analysis when new faults occur, quickly providing location results and significantly reducing the time spent manually analyzing logs and compiling metrics. Automated analysis enables near real-time fault location, which is crucial for business continuity, especially for latency-sensitive IB / RoCE networks. The entire diagnostic process is standardized, ensuring consistent location results whether performed by experienced experts or new engineers.

[0128] This embodiment enables a leap from "phenomenon" to "root cause." Specifically, static fault sources describe the "symptoms" of the fault (such as a sudden increase in latency), while dynamic fault sources describe how the symptoms "evolve" (such as periodic fluctuations). Historical correlation analysis learns how the "symptoms" and "evolution" jointly point to the "root cause." The large model ultimately maps the combination of "symptoms" and "evolution" precisely to the "comprehensive fault source" (such as RoCE congestion path propagation), thereby locating the root cause of the problem. This achieves highly automated and intelligent fault location, simplifying complex problems into rapid predictions, thus significantly improving fault response speed, reducing reliance on human experts, and demonstrating stronger universality in the face of new problems.

[0129] Furthermore, this application also includes: when an IB network problem occurs, firstly determine whether the problem exists in the stored IB dimension indicator mapping information; if so, determine the corresponding problem fault source; if not, determine the corresponding physical network unified category information, and obtain the corresponding fault source from the RoCE dimension indicator mapping information based on the physical network unified category information as the fault source corresponding to the IB network problem.

[0130] This embodiment introduces a strategy that utilizes RoCE dimension information to assist in diagnosing IB network-related problems, effectively improving the accuracy and coverage of IB network-related problem diagnosis. Specifically, when an IB network-related alarm or diagnostic process is triggered, the system first searches the IB dimension indicator mapping information to see if a mapping relationship between the indicator item and the fault source for this IB-related problem has been established previously. If found, the fault source is directly identified, which is the most direct and ideal situation. If no directly corresponding IB-related problem is found in the IB dimension mapping, then this IB problem is classified into a unified physical network category that we have defined. Based on this unified physical network category, the system searches the RoCE dimension indicator mapping information, and the fault source found in the RoCE dimension mapping that corresponds to the same unified physical network category is taken as the fault source of the current IB network-related problem.

[0131] This embodiment effectively compensates for the lack of knowledge in the IB domain and cleverly utilizes the existing, potentially more comprehensive, diagnostic knowledge base (mapping of indicators to fault sources) in the RoCE domain. When similar problems occur in the IB network, even if the expert experience or data mapping in the IB itself is incomplete, the successful experience of RoCE can be used to infer the problem in the IB. It eliminates the need to independently build mappings for all possible fault scenarios from scratch for both the IB and RoCE dimensions. RoCE mappings can be built first and then used to supplement the IB, maintaining a unified physical network category system. This allows the diagnostic models / rules of IB and RoCE to share some logic. In the absence of IB dimension mappings, directly obtaining the fault source through the RoCE dimension is much faster than in completely uncertain situations, effectively avoiding a lengthy and aimless troubleshooting process, automating analogical diagnosis, and reducing the need for human expert intervention.

[0132] Similarly, when a RoCE network-related problem occurs, first determine whether the problem exists in the stored RoCE dimension indicator mapping information. If it does, determine the corresponding fault source. If not, determine the corresponding physical network unified category information and obtain the corresponding fault source from the IB dimension indicator mapping information based on the physical network unified category information as the fault source corresponding to the RoCE network-related problem.

[0133] Furthermore, the application also includes: building a unified observable interface covering the entire AI cluster on the monitoring platform, using abstract indicators as the core dimension, shielding the underlying network differences, presenting fault source information in the form of a visual graph, providing a unified global view for operations and maintenance personnel, and providing users with a global and standardized health view.

[0134] The proposed solution allows for the comprehensive consideration and analysis of both static and dynamic fault source indicators, providing a more complete understanding of the background and causes of faults. When static indicators reach preset thresholds or exhibit abnormalities, a potential fault is indicated. For example, detecting a hardware device's temperature exceeding the safe range is a static indicator trigger signal, suggesting a possible hardware fault. After a static indicator is triggered, dynamic indicators are further examined. If these dynamic indicators also show abnormal changes related to the fault, the presence and location of the fault source can be more definitively confirmed. For instance, if a hardware device's temperature is excessively high while its CPU usage is abnormally high within the same timeframe, this provides dynamic evidence of a fault, indicating that the problem may not solely stem from the hardware itself but may also be related to the system's operational status.

[0135] By combining static indicator triggering with dynamic indicator corroboration, the source of a fault can be more accurately identified, yielding fault source information that comprehensively considers both static and dynamic factors. Compared to relying on a single type of indicator for fault location, this method reduces the possibility of false positives and false negatives, thereby improving the accuracy of fault location. For example, in a complex industrial control system, by jointly analyzing static equipment parameter indicators and dynamic system operating status indicators, the specific location and cause of the fault can be found more quickly and accurately, allowing for timely corrective measures and reducing system downtime and losses.

[0136] Existing monitoring systems can only display performance data for IB and RoCE networks separately, and cannot provide a global health view. This application standardizes the switch port metrics of the IB network and the network monitoring metrics of the RoCE network, and determines the unified category information of the physical network. This allows the metrics of the two types of networks to be integrated, so that maintenance personnel can view the overall health status of the network from a global perspective, rather than viewing two independent network metrics, which greatly improves their overall understanding of the network status.

[0137] When network bottlenecks or failures occur, the original solution relied on maintenance personnel to analyze IB and RoCE indicators separately, resulting in long analysis cycles. This solution maps and integrates the two types of network indicators, using static analysis and time-series analysis methods to quickly locate the fault source, reducing the workload and time required for manual analysis and significantly shortening the fault analysis cycle, enabling the network to recover to normal operation more quickly. Furthermore, through standardized processing and indicator mapping, this application can more accurately correlate and analyze the two types of network indicators, effectively analyzing both static instantaneous anomalies and dynamic trend anomalies, thus more precisely locating the fault source and avoiding misjudgments and omissions caused by scattered and inconsistent information.

[0138] By mapping the standardized indicators to the unified category information of physical networks, we obtain the IB dimension indicator mapping information and the RoCE dimension indicator mapping information. This mapping relationship enables the indicators of the two types of networks to be correlated and analyzed within a unified framework, which facilitates the discovery of potential connections and mutual influences between different networks, thereby providing a more comprehensive understanding of the network's operational status.

[0139] By acquiring real-time abnormal indicators at a certain moment and combining them with indicator mapping information to determine the source of static faults, this method can quickly capture sudden abnormal situations in the network, such as instantaneous link errors and traffic anomalies, and promptly discover and locate static faults, providing strong support for real-time network monitoring and emergency response.

[0140] By employing time-series analysis to analyze indicators within a preset period, trend anomaly information can be obtained and dynamic fault sources can be identified. This helps to discover potential problems and development trends in the network, such as gradual increases in traffic and slow increases in packet loss rate. By detecting these dynamic anomalies in advance, maintenance personnel can take preventative measures to avoid further deterioration of faults and improve network stability and reliability.

[0141] By identifying comprehensive fault sources based on both static and dynamic fault sources, and combining static and dynamic analysis, network faults can be located more comprehensively and accurately. This approach considers the transient and evolving nature of network faults, avoids the limitations of single analysis methods, and improves the accuracy and effectiveness of fault location. In conclusion, this fault location method for cross-IB and RoCE networks significantly improves the efficiency and accuracy of network fault location by addressing the shortcomings of existing monitoring systems, optimizing indicator processing and integration, and employing effective fault analysis methods, thus providing strong support for stable network operation.

[0142] This invention achieves end-to-end observability across IB and RoCE networks by uniformly abstracting, transforming, and fusing heterogeneous network metrics at the platform layer. Specifically, it maps the underlying metrics of IB and RoCE to a unified abstract category, achieving standardization and comparability of metrics. Regardless of network protocol differences, the platform layer can uniformly manage and display key performance indicators, thus solving the problems of scattered and difficult-to-compare metrics in existing technologies. Metric collection covers switches and GPU nodes, forming a multi-dimensional data source, which is then correlated and fused through the platform layer. This end-to-end, multi-source fusion approach allows operations personnel to obtain the overall network health status without separately analyzing the original IB and RoCE metrics, thereby enhancing cross-platform observation capabilities. Metric abstraction and transformation at the platform layer presents heterogeneous network data under a unified view, while retaining the original metrics for in-depth diagnostics. Operations personnel can directly perform alarms, trend analysis, and fault location based on the unified abstract metrics, reducing the analysis complexity in multi-network environments and significantly improving operational efficiency. The technical solution is designed to support multiple data collection methods (polling or event-driven) and can expand metric categories to adapt to different scales and heterogeneous network environments. This solution maintains consistent observability even when the AI ​​cloud platform rapidly expands or adds new network types, ensuring system scalability and adaptability. Through a unified metric view and real-time alarm mechanism, combined with interfacing upper-layer AIOps or scheduling platforms, intelligent operation and maintenance management is achieved. This not only reflects the network's operational status in real time but also provides data support for predictive maintenance and automated scheduling, significantly improving the platform's overall reliability and operational intelligence. In summary, this invention, through the standardization, abstraction, and transformation of metrics across IB and RoCE networks, achieves unified metrics, end-to-end observability, multi-source fusion, and intelligent operation and maintenance, effectively solving problems such as scattered metrics, difficulties in cross-platform observation, and insufficient operational efficiency in existing technologies.

[0143] A computer device according to embodiments of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.

[0144] The processor may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In one embodiment of this disclosure, the processor is used to execute computer-readable instructions stored in the memory, causing the computer device to perform all or part of the steps of the fault location methods across IB and RoCE networks described in the foregoing embodiments of this disclosure.

[0145] Those skilled in the art will understand that, in order to solve the technical problem of how to achieve a good user experience, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included within the protection scope of this disclosure.

[0146] like Figure 3 This is a schematic diagram of a computer device provided for an embodiment of the present disclosure. It illustrates a structural schematic diagram suitable for implementing the computer device in the embodiments of the present disclosure. Figure 3 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0147] like Figure 3 As shown, a computer device may include a processor (such as a central processing unit, graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or programs loaded from storage devices into random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0148] Typically, the following devices can be connected to the I / O interface: input devices, such as sensors or visual information acquisition devices; output devices, such as displays; storage devices, such as magnetic tapes or hard drives; and communication devices. Communication devices allow the computer device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 3 A computer apparatus with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or included alternatively.

[0149] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from ROM. When the computer program is executed by a processor, all or part of the steps of the fault location method across IB and RoCE networks according to embodiments of this disclosure are performed.

[0150] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0151] A computer-readable storage medium according to embodiments of the present disclosure stores non-transitory computer-readable instructions. When these non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the fault location methods across IB and RoCE networks described in the foregoing embodiments of the present disclosure are performed.

[0152] The aforementioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or portable hard drive), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).

[0153] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0154] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0155] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0156] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A method for fault localization across IB and RoCE networks, the method comprising: include: Obtain the raw metrics of the IB network class in the IB network environment and the raw metrics of the RoCE network class in the RoCE network environment. Based on the fault location analysis requirements and target scenario, determine the unified category information of the physical network, specifically including: obtaining all collectable field indicator types in the target scenario of the fault to be analyzed; The fault location analysis requirements are compared with the field indicator types. If the field indicator types include the fault location analysis requirements, the unified category information of the physical network is determined based on the fault location analysis requirements. If the field indicator types do not all include the fault location analysis requirements, the unified category information of the physical network is determined based on the field indicator types. The unified category information of the physical network includes several indicator types and several indicator items corresponding to each indicator type. The standardized original indicators of the IB network class and the original indicators of the RoCE network class are mapped to the unified category information of the physical network to obtain the IB dimension indicator item mapping information and the RoCE dimension indicator item mapping information. Acquire the instantaneous real-time collected original indicators of the IB network class and all indicators in the original indicators of the RoCE network class that do not meet the preset conditions, and determine the static fault source of the indicator item based on the mapping information of the IB dimension indicator item and the mapping information of the RoCE dimension indicator item. Analyze the original indicators of the IB network class and the RoCE network class within the preset period of collection to obtain trend anomaly indicator information; Based on the IB dimension indicator item mapping information and the RoCE dimension indicator item mapping information, determine the dynamic fault source of the indicator item corresponding to the trend anomaly indicator information; Based on the static fault sources and dynamic fault sources of the aforementioned indicator items, the comprehensive fault source of the indicator item is determined.

2. The fault location method across IB and RoCE networks according to claim 1, characterized in that, The acquisition of raw IB network class metrics in the IB network environment and raw RoCE network class metrics in the RoCE network environment includes: The raw metrics of the IB network class are collected through the subnet manager and performance management agent interfaces; The IB network class raw metrics include switch port metrics, which include one or more of the following: port transmit / receive bytes, port transmit / receive packets, link error count, and latency statistics. The raw metrics of the RoCE network class are collected through the Telemetry / SNMP, interface, RoCE, network card driver, or operating system performance counters on the GPU node side of the Ethernet switch. The original metrics for RoCE networks include network monitoring metrics, which include one or more of the following: traffic-related metrics, error and packet loss metrics, congestion control metrics, and resource utilization metrics.

3. The fault location method across IB and RoCE networks according to claim 1, characterized in that, The aforementioned metric types include throughput, error, latency, and link status, and the metric items are different for each type.

4. The fault location method across IB and RoCE networks according to claim 1, characterized in that, The analysis of the original indicators of the IB network class and the original indicators of the RoCE network class within the preset collection period to obtain trend anomaly indicator information includes: Determine the preset cycle; The preset period is divided into several sub-segments; The sub-segments in which the original indicators of the same IB network class deviate from the preset normal range within the preset period are recorded as IB network class abnormal segments. Obtain consecutive IB network anomaly segments. If the number of consecutive IB network anomaly segments exceeds a preset threshold, determine the corresponding IB network original index as an IB network trend anomaly index. The sub-segments in which the original indicators of the same RoCE network class deviate from the preset normal range within the preset period are recorded as RoCE network class abnormal segments. Obtain consecutive RoCE network anomaly segments. If the number of consecutive RoCE network anomaly segments exceeds a preset threshold, determine the corresponding RoCE network original index as a RoCE network trend anomaly index. Each of the preset normal intervals is matched with the corresponding original indicator.

5. The fault location method across IB and RoCE networks according to claim 1, characterized in that, The step of determining the comprehensive fault source of an indicator item based on its static and dynamic fault sources includes: If the static fault source of the indicator item is a data transmission symbol error or a data integrity error, and the dynamic fault source of the indicator item is the continuous increase of the static fault source of the indicator item and the decrease in throughput, then the comprehensive fault source of the indicator item is determined to be a physical layer fault, and the corresponding fault type is optical module damage. If the static fault source of the indicator item is a single LatencyAvg peak, and the dynamic fault source of the indicator item is the periodic fluctuation of delay and the accumulation of TxErrors, the comprehensive fault source of the indicator item is determined to be RoCE congestion path propagation, and the corresponding fault type is congestion propagation. If the static fault source of the indicator item is that the node TxBytes is significantly low, and the dynamic fault source of the indicator item is that the bandwidth variance of each node increases, the comprehensive fault source of the indicator item is determined to be GPU communication load imbalance, and the corresponding fault type is node imbalance. If the static fault source of the indicator item is no significant static error, and the dynamic fault source of the indicator item is a persistently high local Throughput and an increase in LatencyAvg, the comprehensive fault source of the indicator item is determined to be a local link hotspot, and the corresponding fault type is a topology hotspot.

6. The fault location method across IB and RoCE networks according to claim 5, characterized in that, Also includes: Acquire historical target data within a preset period, including static fault sources and dynamic fault sources of indicator items corresponding to all IB network classes and RoCE network classes; Based on the historical target data, obtain the fault association information between the IB network class and the RoCE network class; The large model is trained based on the fault association information; The newly emerging faults are analyzed based on the trained large model to obtain fault location results.

7. A fault location system across IB and RoCE networks, characterized in that, include: The raw metrics acquisition module is used to acquire raw metrics of IB network classes in the IB network environment and raw metrics of RoCE network classes in the RoCE network environment. A unified module is used to determine unified category information of the physical network based on fault location analysis requirements and target scenarios. Specifically, it includes: acquiring all collectable field indicator types in the target scenario of the fault to be analyzed; comparing the fault location analysis requirements with the field indicator types; if the field indicator types contain the fault location analysis requirements, determining unified category information of the physical network based on the fault location analysis requirements; if the field indicator types do not all contain the fault location analysis requirements, determining unified category information of the physical network based on the field indicator types; the unified category information of the physical network includes several indicator types and several indicator items corresponding to each indicator type. The mapping module is used to map the standardized original indicators of the IB network class and the original indicators of the RoCE network class to the unified category information of the physical network, respectively, to obtain IB dimension indicator item mapping information and RoCE dimension indicator item mapping information. The indicator item static fault source acquisition module is used to acquire all indicators in the IB network class original indicators and the RoCE network class original indicators that do not meet the preset conditions in real time, and determine the indicator item static fault source according to the IB dimension indicator item mapping information and the RoCE dimension indicator item mapping information. The indicator item dynamic fault source acquisition module is used to analyze the original indicators of the IB network class and the original indicators of the RoCE network class within a preset period to obtain trend anomaly indicator information; and to determine the indicator item dynamic fault source corresponding to the trend anomaly indicator information based on the IB dimension indicator item mapping information and the RoCE dimension indicator item mapping information. The fault source location module is used to determine the comprehensive fault source of the indicator item based on the static fault source and the dynamic fault source of the indicator item.

8. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the fault location method across IB and RoCE networks as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the fault location method across IB and RoCE networks as described in any one of claims 1-6.

10. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the steps of the method according to any one of claims 1-6.