Intelligent operation and maintenance method and system based on end-network cooperation

CN122870618APending Publication Date: 2026-10-02CHINA UNITED NETWORK COMM GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611032174.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-10-02

AI Technical Summary

Technical Problem

拥塞控制优化手段不足,无法动态适应AI负载特征

Benefits of technology

[0016]本公开所提供的实施例,分别从计算节点硬件和网络设备两个来源采集状态数据,解决了传统方案中计算域与网络域数据割裂、缺乏统一采集视角导致跨域问题无法诊断的问题。在此基础上对两类数据进行跨域关联分析,当端侧指标与网侧指标共同满足预设规则时生成跨域告警,解决了传统方案中指标孤立、告警粗放、故障诊断分析能力有限而难以快速精准定位复杂故障根因的问题,能够精准识别网络拥塞导致计算性能下降等跨域故障,显著提升了故障定位的准确性和时效性,同时避免了单一指标阈值告警产生的误告或漏告。进一步地,权1通过对采集数据的解析、清洗和格式转换生成标准化数据,解决了多源异构数据格式不统一导致的跨域分析困难,为后续的拓扑校验、RDMA参数优化等智能运维应用提供了可直接使用的标准化数据基础,从而降低了运维人员手动关联分析不同系统告警和数据的工作量,提升了整体运维效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122870618A_ABST
    Figure CN122870618A_ABST
Patent Text Reader

Abstract

The present disclosure provides an intelligent operation and maintenance method and system of a smart computing center based on end-network cooperation. The method comprises: S1, collecting computing state data and network state data from computing node hardware and network equipment respectively. S2, analyzing, cleaning and format converting the computing state data and network state data to generate standardized data. According to the end-side indicators and network-side indicators in the standardized data, correlation analysis is performed to determine whether the preset cross-domain correlation rule is met. When the correlation between the end-side indicators and the network-side indicators meets the cross-domain correlation rule, cross-domain correlation alarm information is generated. The present disclosure can accurately identify cross-domain faults such as network congestion leading to computing performance decline, significantly improve the accuracy and timeliness of fault location of the smart computing center, and at the same time avoid the problem of false alarm or missed alarm of a single indicator, providing a reliable data foundation for subsequent intelligent operation and maintenance decision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of intelligent operation and maintenance technology for intelligent computing centers, and in particular to an intelligent operation and maintenance method and system for intelligent computing centers based on end-to-end network collaboration. Background Technology

[0002] With the rapid development of artificial intelligence (AI), big data, and cloud computing technologies, especially the rise of applications such as large language models (LLM), the demand for computing power has exploded, driving the large-scale construction of intelligent computing centers. Intelligent computing centers typically employ clusters of tens of thousands of GPU servers, interconnected through high-performance lossless networks based on Remote Direct Memory Access (RDMA) technologies (such as RDMA over Converged Ethernet version 2, RoCEv2) to meet the extreme requirements of high bandwidth and low latency for tasks such as AI training. However, the massive scale, heterogeneous equipment, high-speed networks, and complex dynamics of AI workloads in intelligent computing centers present unprecedented challenges to their operations and maintenance (O&M), which traditional O&M methods struggle to address. The main challenges include: effective collection and analysis of massive multidimensional monitoring data, accurate discovery and verification of large-scale complex network topologies, performance optimization of RDMA networks (especially RoCEv2) (particularly congestion control), and rapid location and root cause analysis of complex faults.

[0003] Currently, several network management and operation platforms for data centers or high-performance computing exist in the industry, such as Cisco DNA Center, Huawei iMaster NCE-Fabric, and NVIDIA Mellanox UFM. These systems typically provide functions such as network device monitoring, basic topology display, fault alarms, and automated configuration. Some solutions are also beginning to introduce the concept of Intelligent Operations Interface (AIOps), attempting to use machine learning for anomaly detection or root cause analysis.

[0004] However, most existing solutions manage the computing and network domains separately, lacking a unified perspective and data correlation analysis capabilities. They cannot effectively diagnose cross-domain issues such as performance degradation caused by network problems or network congestion caused by computing node anomalies. Monitoring indicators are limited to their respective domains, making it difficult to form a global status view. Most systems still rely on alarms based on static thresholds and manual maintenance experience for fault handling. There is a lack of effective automation methods for complex parameter tuning of RDMA networks (such as Data Center Quantized Congestion Notification, DCQCN). Fault diagnosis and analysis capabilities are limited, making it difficult to quickly and accurately locate the root causes of complex faults. Topology management fails to deeply verify specific connection requirements that meet AI training performance (such as "co-track topology"). Congestion control optimization methods are insufficient and cannot dynamically adapt to AI load characteristics. The ability to analyze the correlation between core computing resources such as Graphics Processing Units (GPUs) and the network is weak. Furthermore, due to the lack of a unified and intelligent platform, maintenance personnel need to process alarms and data from different systems, manually performing correlation analysis and fault diagnosis, which is inefficient and requires high skill levels. Summary of the Invention

[0005] This disclosure provides an intelligent operation and maintenance method and system for intelligent computing centers based on end-to-end network collaboration.

[0006] Firstly, this disclosure provides an intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration, comprising: S1, collecting computing status data and network status data from computing node hardware and network devices respectively. The computing status data includes at least one of the following: GPU utilization of the computing node, disk SMART data, and Xid errors. The network status data includes network interface card (NIC) data, switch data, and optical module data. The NIC data includes at least one of the following: PFC count, DCQCN status, and traffic; the switch data includes the switch's queue depth; and the optical module data includes the optical module's received optical power. S2, parsing, cleaning, and format conversion of the computing status data and network status data to generate standardized data. Correlation analysis is performed on the end-side and network-side indicators in the standardized data to determine whether preset cross-domain correlation rules are met. When the correlation between the end-side and network-side indicators meets the cross-domain correlation rules, cross-domain correlation alarm information is generated.

[0007] In some embodiments, the cross-domain association rule includes: determining that the association rule for network congestion leading to decreased computing performance is met when the GPU utilization of a compute node is lower than a first preset threshold, the PFC pause frame transmission count of the compute node's network card exceeds a second preset threshold, and the outbound queue depth of the switch port directly connected to the compute node's network card exceeds a third preset threshold. The cross-domain association alarm is used to indicate slow training possibly due to downstream network congestion. The PFC count includes the PFC pause frame transmission count.

[0008] In some embodiments, the intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration further includes: S3, storing the structured information in the standardized data into a relational database, and storing the time-series data into a time-series database.

[0009] In some embodiments, the intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration further includes: S4, predicting the time series of the collected optical module received optical power or disk SMART data, and generating predictive alarm information when the predicted value in a future preset time period is lower than a preset alarm threshold.

[0010] In some embodiments, the intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration further includes: S5, receiving LLDP neighbor information from network devices, and constructing an actual physical topology based on the LLDP neighbor information. A preset planned topology is imported, and the device connection relationships between the actual physical topology and the planned topology are compared for consistency to generate a topology verification result. The planned topology includes device connection relationships and alignment rules, where alignment rules include the correspondence between switch port groups and GPUs within computing nodes. The topology verification result includes connection matching results and port status results.

[0011] In some embodiments, the consistency comparison method includes: S51, for each switch port involved in the same track rule, determining the target compute node connected to the port and the corresponding network interface card (NIC) identifier based on LLDP neighbor information. S52, obtaining the mapping relationship between the NIC identifier and the associated GPU identifier. S53, comparing the GPU identifier in the mapping relationship with the GPU set in the same track rule. If the GPU identifier in the mapping relationship does not belong to the GPU set specified in the same track rule, the same track verification is determined to have failed, and violation information is recorded in the topology verification result.

[0012] In some embodiments, the intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration further includes: S6, receiving RDMA network performance metrics from the network interface card (NIC) of the computing node. Using the RDMA network performance metrics as feedback, running an improved simulated annealing algorithm to dynamically search for the optimal DCQCN parameter combination. Sending the optimal DCQCN parameter combination to the target NIC to complete parameter configuration. The RDMA network performance metrics include at least one of traffic, round-trip time (RTT), and PFC count.

[0013] In some embodiments, the improved simulated annealing algorithm employs at least one of the following mechanisms for optimal parameter search: multi-parameter joint adjustment, variable step size search, and multi-round annealing.

[0014] In some embodiments, the intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration further includes: S7, sending a node marking instruction to the resource scheduling system according to a preset linkage strategy, marking the computing node that generates cross-domain association alarm as unhealthy, so as to prevent new training tasks from being assigned to the node.

[0015] Secondly, this disclosure provides an intelligent operation and maintenance system for a smart computing center based on end-to-end network collaboration, comprising a data access layer and a logic layer. The data access layer, deployed on the computing nodes, is used to: collect computing status data and network status data from the computing node hardware and network devices, respectively. The computing status data includes at least one of the following: GPU utilization, disk SMART data, and Xid errors. The network status data includes network interface card (NIC) data, switch data, and optical module data. NIC data includes at least one of the following: PFC count, DCQCN status, traffic, and packet errors; switch data includes the switch's queue depth; and optical module data includes the optical module's received optical power. The logic layer is used to: parse, clean, and convert the computing status data and network status data to generate standardized data. Based on the correlation analysis between end-side and network-side indicators in the standardized data, it determines whether preset cross-domain correlation rules are met. When the correlation between end-side and network-side indicators meets the cross-domain correlation rules, cross-domain correlation alarm information is generated.

[0016] The embodiments provided in this disclosure collect status data from both computing node hardware and network devices, solving the problem of fragmented data between the computing and network domains and the lack of a unified collection perspective in traditional solutions, which leads to the inability to diagnose cross-domain problems. Based on this, cross-domain correlation analysis is performed on the two types of data. When both end-side and network-side indicators meet preset rules, a cross-domain alarm is generated. This solves the problems of isolated indicators, coarse alarms, and limited fault diagnosis and analysis capabilities in traditional solutions, making it difficult to quickly and accurately locate the root causes of complex faults. It can accurately identify cross-domain faults such as network congestion leading to decreased computing performance, significantly improving the accuracy and timeliness of fault location, while avoiding false alarms or missed alarms caused by single indicator threshold alarms. Furthermore, by parsing, cleaning, and format conversion of the collected data to generate standardized data, this solves the difficulty of cross-domain analysis caused by inconsistent formats of multi-source heterogeneous data. It provides a directly usable standardized data foundation for subsequent intelligent operation and maintenance applications such as topology verification and RDMA parameter optimization, thereby reducing the workload of operation and maintenance personnel in manually correlating and analyzing alarms and data from different systems and improving overall operation and maintenance efficiency.

[0017] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0018] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of exemplary embodiments with reference to the accompanying drawings, in which: Figure 1 A flowchart illustrating an intelligent operation and maintenance method for a smart computing center based on end-to-end network collaboration, provided as an embodiment of this disclosure; Figure 2 A flowchart of another intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration provided in this disclosure embodiment; Figure 3 A block diagram of an intelligent operation and maintenance system for a smart computing center based on end-to-end network collaboration, provided as an embodiment of this disclosure; Figure 4 A system block diagram of an example of an intelligent operation and maintenance system for a smart computing center based on end-to-end network collaboration, provided in an embodiment of this disclosure; Figure 5 This is a connection diagram illustrating a fault scenario of a central optical module in an example of an intelligent operation and maintenance system for a smart computing center based on end-to-end network collaboration, provided in an embodiment of this disclosure. Detailed Implementation

[0019] To enable those skilled in the art to better understand the technical solutions of this disclosure, exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments of this disclosure to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0020] Where there is no conflict, the various embodiments of this disclosure and the features thereof in the embodiments may be combined with each other.

[0021] As used herein, the term “and / or” includes any and all combinations of one or more related enumerated entries.

[0022] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. As used herein, the singular forms “a” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that when the terms “comprising” and / or “made of” are used in this specification, the presence of the stated feature, integral, step, operation, element, and / or component is specified, but the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof is not excluded. Words such as “connected” or “linked” are not limited to physical or mechanical connections but can include electrical connections, whether direct or indirect.

[0023] Unless otherwise specified, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art. It will also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and this disclosure, and will not be interpreted as having an idealized or overly formal meaning, unless expressly so defined herein.

[0024] Figure 1 A flowchart illustrating an intelligent operation and maintenance method for a smart computing center based on end-to-end network collaboration, provided as an embodiment of this disclosure. Figure 1 As shown, the method includes: Step S1: Collect computing status data and network status data from computing node hardware and network devices, respectively.

[0025] In step S1, the computing status data includes at least one of the following: GPU power consumption, GPU temperature, GPU utilization, memory usage, NVIDIA Link (NVLink) traffic, card loss status, and Xid errors. The network status data includes network interface card (NIC) data, switch data, and optical module data. NIC data includes at least one of the following: Priority Flow Control (PFC) count, Explicit Congestion Notification (ECN) count, DCQCN status, traffic, and packet errors. Switch data includes the switch's queue depth and port buffer utilization. Optical module data includes the optical module's transmit and receive power and temperature.

[0026] Step S2 involves parsing, cleaning, and format conversion of the computational status data and network status data to generate standardized data. Correlation analysis is then performed on the endpoint and network-side indicators within the standardized data to determine if they meet preset cross-domain correlation rules. When the correlation between the endpoint and network-side indicators satisfies the cross-domain correlation rules, a cross-domain correlation alarm is generated.

[0027] Understandably, in traditional solutions, the compute and network domains are monitored separately. When cross-domain issues such as "network congestion causing slower computing" occur, operations and maintenance personnel need to check both monitoring systems separately and manually piece together clues. This solution, through correlation analysis, can directly discover cross-domain issues and generate accurate alerts, helping operations and maintenance personnel quickly locate faults.

[0028] Step S1 clarifies the specific data collection objects and indicator types (GPU, PFC, queue depth, optical modules, etc.), filling the gaps in traditional monitoring of "critical status monitoring blind spots" and providing a complete data foundation for subsequent analysis. Step S2, the data cleaning and format conversion steps, solves the problem of inconsistent formats of multi-source heterogeneous data, enabling subsequent analysis to be performed on a standardized data view.

[0029] For example, cross-domain related alarm information can be displayed in the user interface for easy viewing by the user.

[0030] In some embodiments, the cross-domain association rule includes: determining that the association rule for network congestion leading to decreased computing performance is met when the GPU utilization of a compute node is lower than a first preset threshold, the PFC pause frame transmission count of the compute node's network card exceeds a second preset threshold, and the outbound queue depth of the switch port directly connected to the compute node's network card exceeds a third preset threshold. The cross-domain association alarm information is used to indicate slow training possibly due to downstream network congestion. The PFC count includes the PFC pause frame transmission count.

[0031] Understandably, the cross-domain association rule consists of three conditions: Condition 1, the GPU utilization of the computing node is lower than a first preset threshold; Condition 2, the PFC pause frame transmission count of the computing node's network card exceeds a second preset threshold; Condition 3, the outbound queue depth of the switch port directly connected to the network card exceeds a third preset threshold. These three conditions correspond to the edge-side computation result (e.g., low GPU utilization), edge-side network behavior (e.g., the network card continuously sending PFC pause frames), and network-side congestion status (e.g., switch queue accumulation). When all three conditions are met simultaneously, a scenario where network congestion leads to decreased computational performance is identified, and an alarm is generated to indicate "slow training, suspected downstream network congestion."

[0032] Understandably, a single metric (such as low GPU utilization alone) cannot determine whether the problem lies with the compute node itself or the network, easily leading to false alarms or missed alarms. Three conditions form a complete causal chain: queue backlog indicates downstream congestion, PFC pause frames indicate the network card is under backpressure, and low GPU utilization indicates that compute tasks are stalled while waiting for data. This combined approach significantly improves the accuracy of alerts. Furthermore, the alert information directly points to the fault, rather than simply reporting abnormal metrics, providing clear troubleshooting guidance for operations personnel and shortening fault location time.

[0033] In some embodiments, the intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration further includes: step S3, storing the structured information in the standardized data into a relational database, and storing the time-series data into a time-series database.

[0034] Understandingly, storing structured data and time-series data separately leverages the transactional consistency and complex join queries of relational databases, as well as the advantages of time-series databases in high-concurrency writes and fast queries by time range, thereby improving overall data access efficiency. Persistent storage provides the data foundation for predictive alerts, historical backtracking, and trend analysis.

[0035] In some embodiments, the intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration further includes: step S4, predicting the time series of the collected optical module received optical power or disk SMART data, and generating predictive alarm information when the predicted value in a future preset time period is lower than a preset alarm threshold.

[0036] Understandably, time series data from optical module received power or self-monitoring analysis and reporting technology (SMART) are modeled and predicted (e.g., using an autoregressive integral moving average model (ARIMA) or a long short-term memory network (LSTM) model). When the predicted value output by the model within a preset future time period is lower than a preset alarm threshold, an alarm is generated in advance and pushed to the user interface. Unlike traditional alarms that only sound when the threshold is exceeded, step S4 can predict the trend of the indicator before it reaches the threshold.

[0037] Understandably, a decrease in optical module receive power is a typical precursor to fiber optic link aging, and abnormal SMART data on a disk indicates a potential disk failure. This solution predictively identifies potential problems in advance, providing maintenance personnel with a window of opportunity to address issues before a failure actually occurs (e.g., replacing optical modules or migrating data), thus reducing the impact of sudden failures on services. Predictive maintenance shifts the maintenance model from reactive response to proactive prevention, reducing the workload of emergency repairs and the risk of service interruption.

[0038] In some embodiments, the intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration further includes: S5, receiving Link Layer Discovery Protocol (LLDP) neighbor information from network devices, constructing the actual physical topology based on the LLDP neighbor information, importing a preset planned topology, comparing the device connection relationships between the actual physical topology and the planned topology for consistency, and generating a topology verification result.

[0039] In step S5, the topology planning includes device connection relationships and alignment rules. The alignment rules include the correspondence between switch port groups and GPUs within compute nodes. The topology verification results include connection matching results and port status results.

[0040] Understandably, step S5 automates network topology discovery and consistency verification. First, it receives neighbor information reported by network devices via the LLDP protocol (i.e., each device knows who its neighbors are), and constructs the actual physical connection topology accordingly. Simultaneously, the administrator imports a planned topology, which describes the expected connection relationships between devices and alignment rules (i.e., which GPUs correspond to switch port groups). Then, the actual topology is compared with the planned topology: checking whether each actual connection matches the plan (i.e., connection matching), checking whether ports that should be operational in the plan are actually online (i.e., port status verification), and generating and displaying a topology verification report containing matching / mismatch / missing results.

[0041] Understandably, traditional topology tools can only display the actual connection status and cannot determine whether the connections are correct. This solution, through a dual-source comparison of the actual topology and the planned topology, can automatically detect construction or configuration errors such as incorrect or missing connections, thus avoiding the impact of topology issues on AI training performance.

[0042] In addition, the topology verification results can be visualized in the user interface, which can reduce the workload of manual verification, especially in ultra-large-scale clusters with tens of thousands of connections, where manual verification is almost infeasible.

[0043] In some embodiments, such as Figure 2 As shown, the methods for consistency comparison include: Step S51: For each switch port involved in the same track rule, determine the target computing node connected to the port and the corresponding network card identifier based on the LLDP neighbor information.

[0044] Step S52: Obtain the mapping relationship between the network card identifier and the associated GPU identifier.

[0045] Step S53: Compare the GPU identifier in the mapping relationship with the GPU set in the same track rule.

[0046] In step S53, if the GPU identifier in the mapping relationship does not belong to the GPU set specified by the same track rule, the same track verification is determined to be unsuccessful, and the violation information is recorded in the topology verification result.

[0047] Understandably, for each switch port defined in the alignment rules, the actual target compute node and corresponding network interface card (NIC) identifier connected to the port are first found based on LLDP neighbor information. Then, the mapping relationship between the NIC and its associated GPU identifier is obtained (by executing commands such as `lspci` and `nvidia-smi topo -m` on the target compute node). Finally, the GPU identifier in this mapping relationship is compared with the GPU set specified in the alignment rules—if the NIC is connected to the wrong GPU (e.g., the port is planned to connect to GPUs 0-3, but is actually connected to GPU 4), ​​the alignment verification fails and the violation information is recorded.

[0048] Understandably, "same track" refers to ensuring that GPU communication occurs within the same switch port group as much as possible, avoiding additional hops across switches that increase communication latency. If the mapping between the network card and the GPU is incorrect (i.e., misalignment), communication between GPUs will take detours, significantly impacting the efficiency of distributed AI training. This solution can automatically detect these logical mapping errors that LLDP alone cannot detect. The verification process is automated, eliminating the need for manual login and inspection of each server, resulting in high efficiency in large-scale clusters.

[0049] In some embodiments, the intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration further includes: S6, receiving RDMA network performance metrics from the network interface card (NIC) of the computing node; using the RDMA network performance metrics as feedback, running an improved simulated annealing algorithm to dynamically search for the optimal DCQCN parameter combination; and distributing the optimal DCQCN parameter combination to the target NIC to complete the parameter configuration.

[0050] In step S6, the RDMA network performance metrics include at least one of traffic, round-trip time (RTT), and PFC count.

[0051] Understandably, the system receives RDMA network performance metrics from the compute node's network interface card (NIC), uses these real-time metrics as feedback input, runs an improved simulated annealing algorithm to search for the optimal DCQCN parameter combination, and then sends the search results to the target NIC via a data access proxy to complete parameter configuration. The entire process forms a closed loop of data acquisition, analysis, optimization, configuration, and effect feedback, and is executed cyclically to adapt to load changes.

[0052] Understandably, in traditional methods, DCQCN parameters are manually set by administrators based on experience. Once configured, they are fixed and cannot adapt to dynamic changes such as the start and stop of AI training tasks and traffic fluctuations. This solution is based on real-time feedback and automatic optimization, allowing parameters to be dynamically adjusted according to load changes. Compared to manual parameter tuning or traditional grid search, the improved simulated annealing algorithm can find optimal solutions more efficiently in high-dimensional parameter spaces, improving network throughput and reducing congestion risk.

[0053] In some embodiments, the improved simulated annealing algorithm employs at least one of the following mechanisms for optimal parameter search: multi-parameter joint adjustment, variable step size search, and multi-round annealing.

[0054] Understandably, simulated annealing is a stochastic search algorithm inspired by the physical process of solid-state annealing, specifically designed to find the global optimum in a complex parameter space. However, traditional simulated annealing algorithms suffer from two prominent problems in practical applications: first, they converge slowly, requiring a large number of iterations to find a relatively optimal solution; second, they are prone to getting trapped in local optima, especially when the objective function has multiple local minima, where the algorithm may become stuck in a local optimum during the cooling process and be unable to escape.

[0055] Generally, Data Center Quantized Congestion Notification (DCQCN) is an end-to-end rate control mechanism in RoCEv2 networks. Its parameter combination (including multiple interrelated parameters such as congestion threshold, timeout, rate increase factor, and rate decrease factor) directly affects the network's throughput, latency, and congestion control effectiveness. These parameters have complex coupling relationships—adjusting one parameter may improve performance in one aspect, but may simultaneously degrade performance in others. Therefore, they need to be optimized synergistically as a whole, rather than adjusted individually in isolation.

[0056] Therefore, to address the slow convergence of traditional simulated annealing algorithms, this embodiment introduces a variable step-size search mechanism: in the early stages of the algorithm's operation, when the temperature is high, a larger search step size is used, allowing the parameter combination to move quickly across the global scope and rapidly locate the region where the optimal solution may exist; as the temperature gradually decreases, the search step size also decreases, performing a fine-tuning search within the located high-quality region. This strategy of "large-step exploration in the early stage and small-step fine-tuning in the later stage" significantly accelerates the convergence speed of the algorithm. To address the problem of easily getting trapped in local optima, this solution designs a multi-round annealing mechanism—instead of executing a single annealing cycle, it executes multiple independent "heating-cooling" annealing processes. Each round of annealing starts from a different initial parameter combination, independently searching for its own optimal solution, and finally selecting the best one from multiple search results. Multi-round annealing is equivalent to simultaneously exploring the parameter space from multiple different starting points, greatly increasing the probability of escaping local optima and finding the globally optimal parameter combination. Furthermore, this scheme employs a multi-parameter joint adjustment strategy, simultaneously fine-tuning multiple DCQCN parameters in each neighborhood perturbation, rather than adjusting each parameter sequentially, thus fully considering the coupling effect between parameters. Through the combined application of the above three improvement mechanisms, the algorithm can dynamically search for the DCQCN parameter combination that best matches the current network load state within an acceptable time overhead.

[0057] In other words, the standard simulated annealing algorithm suffers from slow convergence and a tendency to get trapped in local optima. Multi-parameter joint adjustment, however, leverages the coupling relationships between parameters to avoid suboptimal solutions encountered during step-by-step optimization. Variable step size search balances exploration efficiency and convergence accuracy. Multiple rounds of annealing increase the chances of escaping local optima. These three mechanisms collectively improve the algorithm's search efficiency and solution quality, and can be used individually or in combination.

[0058] Take, for example, a smart computing center training a large AI model. This center comprises 64 GPU servers interconnected via a RoCEv2 network. All server network cards currently use a default DCQCN parameter configuration (e.g., default congestion threshold, default timeout). After the system initiates the RDMA network parameter adaptive optimization task, it first collects real-time network performance metrics through data access agents deployed on each computing node. These metrics include the traffic rate of each network card, round-trip time (RTT), and PFC pause frame count. These real-time metrics serve as input to the objective function—designed to maximize effective network throughput while maintaining low packet loss and low latency.

[0059] After the optimization process begins, the improved simulated annealing algorithm first randomly generates a set of initial DCQCN parameter combinations as the current optimal solution, and sets the current "annealing temperature" to a relatively high initial value. In the first round of annealing, the algorithm enters the outer iteration: in each inner iteration, the algorithm applies random perturbations to the current parameter combination—for example, adjusting the congestion threshold by ±5% from the current value, adjusting the timeout by ±10%, etc., with multiple parameters changing simultaneously. After the perturbation, a new set of parameter combinations is obtained. The system distributes this new set of parameters to each network interface card (or network interface cards in a test subset) through the data access agent, and re-collects network performance indicators within a short effective period to calculate the objective function value corresponding to the new parameter combination. If the performance of the new parameter combination is better than the current optimal solution, it is unconditionally accepted as the new current optimal solution; if the performance of the new parameter combination is worse, the acceptance probability is calculated according to the Metropolis criterion—the acceptance probability is higher in the high-temperature stage, so even parameter combinations with deteriorating performance are more likely to be accepted, thus allowing the algorithm to jump out of the local optimum region in the early stages of the search to explore other parameter spaces. After each inner layer iteration, the temperature is slightly reduced according to a preset cooling strategy (such as exponential cooling), and the search step size is also reduced accordingly. When the temperature drops to a certain preset low value, the first round of annealing ends, and the optimal parameter combination found in that round is recorded.

[0060] Subsequently, the algorithm initiates a second round of annealing: a different set of initial DCQCN parameter combinations are randomly generated, and the entire annealing process is repeated. After multiple rounds (e.g., 3 to 5 rounds) of independent annealing, the system selects the set with the optimal objective function value from the optimal parameter combinations searched in each round, as the final result of this optimization, and distributes it to all network cards through the data access proxy. The entire optimization process requires no manual intervention and is completely executed automatically by the system. When the state of the AI ​​training task changes (e.g., a batch of training tasks ends, a new task starts, causing a change in network traffic patterns), the system will re-trigger the above optimization process, dynamically searching for DCQCN parameter combinations that match the new load state, thereby achieving continuous adaptive optimization of RDMA network performance.

[0061] In some embodiments, the intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration further includes: S7, sending a node marking instruction to the resource scheduling system according to a preset linkage strategy, marking the computing node that generates cross-domain association alarm information as unhealthy, so as to prevent new training tasks from being assigned to the node.

[0062] For example, after generating the cross-domain association alarm information in step S2, a node marking instruction is sent to the resource scheduling system according to a preset linkage strategy, marking the computing node that generated the alarm as unhealthy. The resource scheduling system will then avoid the marked node when allocating new training tasks subsequently.

[0063] Understandably, if a computing node is already experiencing network congestion leading to slower computation, assigning new tasks to that node will not only exacerbate the congestion but also cause the new tasks to also experience slow training. This solution proactively avoids problematic nodes at the task allocation level through linked tagging, preventing a vicious cycle of accumulating problems. The entire process is completed automatically, requiring no manual intervention from operations personnel, thus shortening response time.

[0064] Understandably, this disclosure provides an intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration. Addressing the operational challenges posed by the ultra-large scale, heterogeneous equipment, high-speed networks, and dynamic AI load of intelligent computing centers, it constructs a three-layer operation and maintenance architecture based on end-to-end network collaboration. This achieves unified collection, associated storage, and collaborative analysis of data from the computing and network domains, breaking down data silos between end-to-end networks for the first time in an intelligent computing center scenario. Through refined monitoring and intelligent alarm methods that link end-to-end networks, it overcomes the limitations of traditional fragmented indicators, accurately identifying cross-domain faults such as performance degradation due to network congestion and providing early warnings, significantly improving the accuracy and timeliness of fault location. The intelligent computing center-specific network topology consistency verification method, particularly the same-track topology verification mechanism, automatically detects irregular configurations in physical connections and logical mappings, ensuring the topology compliance of the AI ​​training cluster. Through an adaptive optimization method for RDMA network DCQCN parameters, based on an improved simulated annealing algorithm, it achieves dynamic closed-loop adjustment of network parameters, effectively improving network throughput and stability. The above methods enable collaborative operation and one-stop operation and maintenance management between modules, thereby improving the overall operation and maintenance efficiency and system stability of the intelligent computing center, reducing reliance on human experience and operation and maintenance costs, and providing strong support for the efficient and stable operation of core businesses such as AI training.

[0065] In addition, this disclosure also provides an intelligent operation and maintenance system for intelligent computing centers based on end-to-end network collaboration, which can be used to implement any of the intelligent operation and maintenance methods for intelligent computing centers based on end-to-end network collaboration provided in this disclosure. The corresponding technical solutions and descriptions are described in the relevant section of the method.

[0066] like Figure 3 As shown, this disclosure also provides an intelligent operation and maintenance system for intelligent computing centers based on end-to-end network collaboration. The intelligent operation and maintenance system 30 for intelligent computing centers based on end-to-end network collaboration includes a data access layer 301 and a logic layer 302.

[0067] The data access layer 301 is deployed on the compute nodes and is used to collect compute status data and network status data from the compute node hardware and network devices, respectively. Computation status data includes at least one of the following: GPU power consumption, GPU temperature, GPU utilization, memory usage, NVLink traffic, card loss status, and Xid errors. Network status data includes network interface card (NIC) data, switch data, and optical module data. NIC data includes at least one of the following: PFC count, ECN count, DCQCN status, traffic, and packet errors. Switch data includes the switch's queue depth and port buffer utilization. Optical module data includes the optical module's transmit and receive power and temperature.

[0068] The logical layer 302 is used to: parse, clean, and convert the computational status data and network status data to generate standardized data. Based on the correlation analysis between endpoint and network-side indicators in the standardized data, it determines whether preset cross-domain correlation rules are met. When the correlation between endpoint and network-side indicators meets the cross-domain correlation rules, a cross-domain correlation alarm message is generated.

[0069] The following example illustrates an intelligent operation and maintenance system for a smart computing center based on end-to-end network collaboration provided in this disclosure.

[0070] The core of this example lies in building a platform that integrates data acquisition, intelligent analysis, collaborative control, and unified presentation to achieve in-depth collaborative management and optimization of computing and network resources within the intelligent computing center.

[0071] like Figure 4 As shown, the system architecture in this example can be divided into a user interface layer, a core logic layer, and a network data access layer.

[0072] User Interface (UI) Layer: Provides the human-computer interaction interface, including a graphical web interface and a command-line interface (CLI). The web interface visually displays system status and analysis results through dashboards, topology diagrams, alarm lists, configuration pages, etc., and receives user operation commands. The CLI interface is geared towards advanced users or automation scripts, providing more flexible query and control capabilities.

[0073] Core Logic Layer (Master): As the brain of the system, it is typically deployed on a management server cluster. It contains several key functional modules: Data Access & Management module: Responsible for communicating with lower-level agents (e.g., via gRPC), receiving, parsing, cleaning, and storing (in time-series databases such as InfluxDB and relational databases such as MySQL) operational data from various sources.

[0074] Status Monitoring & Analysis module: Processes and analyzes the collected data in real time, calculates key performance indicators (KPIs), and identifies abnormal states.

[0075] Intelligent Alerting module: Based on rule engine and / or machine learning model, it judges the analysis results, generates alarm events, and performs alarm suppression, aggregation and correlation analysis.

[0076] Topology Management & Validation module: responsible for building and maintaining the global network topology (physical and logical), and performing consistency checks (including same-track topology checks).

[0077] Intelligent Optimization & Control module: runs optimization algorithms (such as AI-DCQCN), calculates optimization strategies (such as optimal DCQCN parameters), and sends control commands to relevant devices.

[0078] Task scheduling and workflow engine: Automates tasks within the system, such as periodic inspections, topology verification, and parameter tuning.

[0079] Northbound interface module: Provides APIs (such as RESTful APIs) for the user interface layer or other external systems to call.

[0080] Network Data Access Layer (Agent): Acting as the system's sensor and executor, it is deployed as a lightweight agent process on managed computing nodes (such as GPU servers) and interacts with network devices (switches) via standard protocols (such as SNMP, gRPC, SSH, Telemetry, and NETCONF). Its main functions include: Data Acquisition: Based on the instructions of the core logic layer, actively collect or receive status information, performance counts, configuration data, events / logs (such as Syslog, Traps) from local nodes (CPU, memory, disk, GPU, network card driver / firmware) and network devices.

[0081] For example, in addition to SNMP, gRPC, and Telemetry, data acquisition protocols such as NETCONF / YANG can also be used for data acquisition and configuration of network devices. Other message passing protocols such as MQTT can also be used between the Agent and the Master.

[0082] Instruction execution: Executes instructions issued by the core logic layer, such as running diagnostic scripts, modifying local configurations (e.g., network card DCQCN parameters), and performing performance tests (e.g., NCCL-Test).

[0083] Local preprocessing: Perform preliminary filtering, aggregation, or formatting on the collected data to reduce the amount of data transmitted to the core logic layer.

[0084] The main modules are described below: I. Refined Status Monitoring and Intelligent Alarm Module: This module aims to achieve comprehensive, in-depth, and high-frequency monitoring of the status of the "end" and "network" of the intelligent computing center, and to issue alarms through intelligent means.

[0085] 1. Multi-dimensional data collection: (1) Compute Node (Agent): It collects data in seconds or even sub-milliseconds through operating system interfaces (such as / proc, / sys), specific hardware driver interfaces (such as NVIDIA SMI / NVML to obtain GPU power consumption, temperature, utilization, video memory, NVLink traffic, card drop status, Xid error, etc.), and network card driver interfaces (to obtain PFC / ECN count, DCQCN status, traffic, error packets, etc.).

[0086] (2) Network devices (Master / Agent via Protocols): Obtain basic device status and interface counts via SNMP Get / Trap; Subscribe to high-frequency performance metrics (such as queue depth, port buffer occupancy, and precise traffic count) via Streaming Telemetry (gRPC / Dial-out); Obtain configuration information and execute diagnostic commands via SSH / NETCONF; Receive device logs via Syslog.

[0087] (3) Optical module (Master / Agent via Protocols): The optical module’s transmit and receive power, temperature, voltage, bias current and other information are periodically polled or subscribed to through the interface provided by the switch (such as DOM - Digital Optical Monitoring).

[0088] 2. Data Processing and Storage: The collected data is aggregated into the core logic layer, where it is timestamped, normalized, and structured information is stored in MySQL. Massive amounts of time-series data (such as GPU power consumption, port traffic, PFC counts, etc.) are stored in InfluxDB.

[0089] 3. Intelligent alarm generation: (1) Basic alarm: Alarms are generated based on static / dynamic thresholds (such as GPU temperature > 85℃, optical module received optical power < -7dBm, port CRC error rate > threshold).

[0090] (2) Correlation Alarms: Establish a correlation rule engine. For example, Rule 1: IF (Server A GPU utilization < 10% for 60s) AND (Server A NIC X PFC Pause send count > 1000 / s) AND (Switch port Y directly connected to NIC X outbound queue depth > 80%) THEN Generate "Server A training slow, suspected downstream network congestion alarm". Rule 2: IF (Switch port Z optical module Rx Power < -8dBm) AND (Port Z CRC ErrorRate > 1e-6) THEN Generate "Switch port Z physical link quality degradation alarm".

[0091] (3) Predictive Alarm: Time series prediction of key indicators (such as optical module Rx Power, disk SMART data) is performed (such as using ARIMA, LSTM models). When the predicted value may fall below the threshold in the future, a predictive alarm is generated.

[0092] (4) Alarm Suppression / Aggregation: Suppress and aggregate alarm storms caused by the same root cause (such as a switch power failure causing a large number of port down alarms) to present the root cause alarms.

[0093] 4. Alarm Notification and Presentation: Alarm events are stored in the database and displayed through an alarm list on a web interface. High-priority alarms can be pushed to operations and maintenance personnel via webhooks, email, SMS, or integration into enterprise communication tools (such as Lark or WeChat Work robots).

[0094] Example 1: Optical Module Fault Early Warning and Correlation Diagnosis

[0095] like Figure 5 As shown, (a) the Switch1 port P1 is connected to the ServerX network card eth0; (b) the optical module OM1 is installed on the Switch1 port P1.

[0096] process: 1. The system periodically (e.g., every minute) collects the DOM information of all optical modules on Switch1 via SNMP or Telemetry, including the received optical power (Rx Power) of OM1.

[0097] 2. The monitoring module detected that the Rx Power of OM1 continued to decrease, for example, from -5dBm to -7.5dBm, which is close to or below the preset alarm threshold (such as -7dBm).

[0098] 3. At the same time, the system detected an abnormal increase in the CRC error packet count on Switch1 port P1, and the received packet loss count on the ServerX network card eth0 was also increasing.

[0099] 4. The intelligent alarm module determines that these events are highly correlated based on the preset association rules (similar to rule 2 above, but combining port packet errors and server packet loss). The root cause is likely the performance degradation of the optical module OM1 or the quality problem of the optical fiber link.

[0100] 5. The system generates a single high-priority alarm: "Warning: Switch1 port P1 link quality degraded, which may affect ServerX communication", instead of multiple independent low-priority alarms.

[0101] 6. Alarms are displayed through the operations and maintenance platform interface and may be pushed to the engineers responsible for network maintenance via Lark robot.

[0102] II. Large-scale network topology discovery and consistency verification module: This module aims to automatically construct the physical network topology of the intelligent computing center and verify whether its actual connections conform to the planning and design, especially the specific connection rules that meet the needs of AI training and optimization.

[0103] 1. Topology data acquisition: LLDP discovery: The system collects LLDP neighbor information (peer device ID, port ID, system name, etc.) by using an agent or by directly querying network devices.

[0104] Topology planning import: Administrators can import the planned topology information (device list, port list, expected connection relationships, device roles such as Leaf / Spine, and "same track" rules, etc.) into the system database via files or API interfaces.

[0105] 2. Topology Construction: The topology management module of the core logic layer parses the LLDP data and combines it with the device list to construct the actual physical connection topology data structure.

[0106] 3. Consistency check logic: Step a (Connection Matching): Traverse every connection in the planned topology (e.g., device A-port X should connect to device B-port Y). Search for the LLDP neighbor information of device A-port X in the actual topology data and determine if it is device B-port Y. Record the matching, non-matching (misconnection), and missing (no actual connection found) cases.

[0107] Step b (Port Status Verification): For ports that should be in the Up state according to the plan, query their actual operating status (obtained via SNMP or Telemetry) and record the ports whose status does not match.

[0108] Step c (“Same Track Topology” Verification): This is a special verification for intelligent computing centers.

[0109] Prerequisite: The topology planning must define the same track rule, for example: "All server network cards connected to the Leaf1 switch port group PG1 (including Port1-4) should be mapped to GPU 0-3 in the corresponding server (or belong to the same NUMA node)".

[0110] Verification Process: i. For each Leaf switch port involved in the rule (e.g., Leaf1-Port1), determine the connected server network card (e.g., ServerA-eth0) based on LLDP. ii. The Agent executes commands (e.g., combinations of lspci, ibdev2netdev, nvidia-smi topo -m, etc.) on server ServerA to obtain the mapping relationship between network card eth0 and GPU ID (or NUMA node). iii. Compare the obtained GPU ID (or NUMA node) with the corresponding track rule (which should be mapped to GPU 0-3) of port group PG1 to which the port belongs. iv. If there is a mismatch (e.g., eth0 is mapped to GPU 4), ​​the track verification is deemed to have failed, and the violating connection and mapping relationship are recorded.

[0111] 4. Result Presentation and Visualization: Verification results (match list, error list, missing list, and same-track violation list) are stored in the database and can be queried through a web interface. In the topology visualization interface, links or devices that failed verification are highlighted with different colors or markers, and detailed error information is provided.

[0112] III. Collaborative Scheduling and Unified Management: This aims to manage the collaborative work between various modules through a unified platform.

[0113] Coordinated scheduling: The results from the analysis module can trigger actions from the control module. For example: If a topology check detects a misalignment, an alarm notification can be automatically triggered, and (depending on the policy configuration) the resource scheduling system may be activated to temporarily mark the node as "unhealthy" to prevent it from being assigned a new training task.

[0114] After AI-DCQCN completes optimization, the optimal parameters are automatically distributed to the application.

[0115] A persistent high PFC count and low GPU utilization associated with an alarm may trigger the system to automatically run more in-depth network diagnostic scripts (such as iperf, NCCL-Test) to collect more evidence.

[0116] Unified management platform: Provides a centralized management portal.

[0117] Visualization: Unified display of global topology (highlighting anomalies), real-time / historical monitoring data of each node and link, alarm list, task execution status (such as topology verification progress, AI-DCQCN iteration process), etc.

[0118] Configuration Management: Unifies the system's own configuration, access information of managed devices, alarm rules, inspection strategies, track-based rules, AI-DCQCN tasks (target node, start / stop), etc.

[0119] Operation entry: Provides operation interfaces for starting / stopping diagnostic tasks, manually triggering verification or optimization, confirming / handling alarms, and viewing logs.

[0120] Through the above architecture and key modules, this example achieves closed-loop intelligent operation and maintenance with end-to-end network collaboration.

[0121] Understandably, this disclosure addresses the operational challenges posed by the ultra-large scale, heterogeneous equipment, high-speed networks, and dynamic AI load of intelligent computing centers. It overcomes the pain points of traditional operations and maintenance, such as "end-to-network fragmentation, insufficient intelligence, and lack of scenario adaptation," focusing on three dimensions: end-to-network collaborative architecture design, cross-domain intelligent analysis, and scenario-based operations and maintenance optimization. Specifically, it includes the following aspects: I. Architecture Level: Constructing an Intelligent Operation and Maintenance System Architecture for Intelligent Computing Centers with End-to-End Network Collaboration Breaking through the limitations of traditional operation and maintenance systems that separate the computing domain from the network domain, this design employs a three-layer collaborative architecture: user interface layer, core logic layer, and network data access layer. This forms a closed-loop operation and maintenance system covering both the computing node and the network (RDMA network), providing a unified management framework for intelligent computing centers. 1. Precisely Adapted Layered Responsibilities: The user interface layer provides global visualization and operation entry; the core logic layer acts as the brain, integrating core modules such as data management, monitoring and analysis, alarms, topology verification, and intelligent optimization to achieve decision-making and control; the network data access layer (Agent) acts as the sensor and executor, deployed on computing nodes such as GPU servers, collecting end-network data through multiple protocols (SNMP, gRPC, Telemetry, etc.) and executing control commands, solving the problems of incomplete data collection and untimely command issuance in traditional architectures.

[0122] 2. Unified data flow between endpoints and networks: The architecture supports standardized access, cleaning, and associated storage of data from computing nodes (GPUs, network cards) and network devices (switches, optical modules) (structured data is stored in MySQL, and time-series data is stored in InfluxDB). For the first time, it realizes the same source, same analysis, and same use of data between endpoints and networks in the intelligent computing center, laying the foundation for cross-domain operation and maintenance analysis.

[0123] II. Monitoring and Alarming Level: Refined Monitoring and Intelligent Alarming Methods for End-to-End Network Correlation

[0124] To address the problems of "isolated indicators and inadequate alarms" in traditional operations and maintenance, a multi-dimensional and cross-domain intelligent monitoring and alarm solution is proposed to achieve accurate identification and early warning of faults in intelligent computing centers: 1. High-frequency, multi-dimensional data acquisition method: Breaking through the limitations of traditional threshold-based monitoring, it achieves second-level data acquisition through operating system interfaces, hardware drivers (NVIDIA SMI / NVML, network card drivers), and device protocols (Streaming Telemetry, DOM), covering core indicators of intelligent computing centers such as GPU (power consumption, temperature, video memory, NVLink traffic, card drop status), network cards (PFC / ECN count, DCQCN status), switches (queue depth, port buffer occupancy), and optical modules (receiver power, temperature), solving the problem of blind spots in critical status monitoring.

[0125] 2. Terminal-Network Correlation Alarm Mechanism: Innovative design of correlation rule engine and prediction model to break down the separation between terminal and network indicators: Related alarms: Cross-domain alarms are generated based on rules (such as low GPU utilization + high NIC PFC count + full switch queue) to accurately locate cross-domain issues such as network congestion leading to decreased computing performance; Alarm aggregation: Suppress alarm storms caused by the same root cause (such as multiple port alarms caused by switch power failure), and only display the root cause alarms to reduce maintenance interference.

[0126] III. Topology Management Level: A Dedicated Network Topology Consistency Verification Method for Intelligent Computing Centers

[0127] To address the shortcomings of traditional topology tools that can only display connections and cannot verify scenario adaptability, a full-process topology management solution including same-track topology verification is proposed to ensure that AI training meets the stringent requirements of network topology. 1. Dual-source acquisition of topology data: Combining the LLDP protocol to automatically discover the actual physical topology and the planned topology imported by the administrator (including device roles, port groups, and parallel rules), solving the problem that traditional topologies can only reflect the actual status and cannot be compared with the plan.

[0128] 2. Three-layer verification logic: Breaking through the limitations of traditional methods that only verify port connections, it implements three-layer verification: connection correctness, port status, and parallel compliance. Connection matching: Compare the actual and planned device-port connection relationships to identify incorrect or missing connections; Port status verification: Check the actual operating status of ports that should be up in the plan, and locate abnormal ports; Same-track topology verification: By executing commands such as lspci and nvidia-smi topo -m through the Agent, the mapping relationship between network cards and GPU / NUMA nodes is obtained. This is compared with the rule in the plan that "Leaf switch port groups correspond to specific GPUs" to identify "misaligned" connections (such as network cards mapped to non-specified GPUs) and avoid the impact of topology issues on AI training efficiency.

[0129] IV. Network Optimization Level: Adaptive Optimization Method for DCQCN Parameters in RDMA Networks (RoCE)

[0130] To address the challenge that traditional RDMA network static parameter configurations cannot adapt to the dynamic nature of AI loads, a closed-loop optimization scheme based on an improved simulated annealing algorithm is proposed to enhance network throughput and stability. 1. Real-time feedback driven: Based on real-time performance indicators of the RDMA network (traffic, RTT, PFC count), it replaces traditional manual experience judgment to ensure that the optimization direction is in line with the actual load.

[0131] 2. Improved Simulated Annealing Algorithm: Innovative design of multi-parameter adjustment strategy, variable step size, and multi-round annealing mechanism to solve the problems of slow convergence and easy getting trapped in local optima in traditional algorithms. It can dynamically search for the optimal DCQCN parameter combination (such as threshold and timeout).

[0132] 3. Closed-loop control: Optimization parameters are automatically sent to the network card through the core logic layer to realize a closed loop of data acquisition, algorithm optimization, parameter configuration and effect feedback without manual intervention, and adapt to the dynamic changes of AI load (such as training task start and stop, traffic fluctuations).

[0133] V. New at the System Integration Level: A Collaborative Operation and Maintenance System for End-to-End Networks Integrating Core Methods

[0134] By integrating the aforementioned end-to-end network collaborative architecture, associated alarms, topology verification, RDMA optimization, and other methods into a unified system, the entire process of monitoring, analysis, alarming, optimization, and control can be automated. 1. Module Collaboration and Linkage: Status monitoring results trigger intelligent alarms, topology verification anomalies link with the resource scheduling system to mark unhealthy nodes, and RDMA optimization results automatically execute configurations, solving the problem of independent modules and lack of linkage in traditional systems; 2. Unified Management Portal: The web interface provides a unified view of the global topology (highlighting anomalies), real-time metrics of the terminal network, alarm lists, and optimization task progress. It also provides a configuration management and operation portal, reducing operational complexity and enabling one-stop intelligent computing center operation and maintenance.

[0135] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in connection with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in connection with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of this disclosure as set forth by the appended claims.

Claims

1. A method for intelligent operation and maintenance of intelligent computing centers based on end-to-end network collaboration, characterized in that, include: S1, collect computing status data and network status data from computing node hardware and network devices respectively; the computing status data includes at least one of the following: GPU utilization of computing node, disk SMART data, and Xid error; the network status data includes network interface card (NIC) data, switch data, and optical module data; the NIC data includes at least one of the following: NIC PFC count, DCQCN status, and traffic; the switch data includes the switch queue depth; and the optical module data includes the optical module's received optical power. S2, the computational status data and the network status data are parsed, cleaned, and format-converted to generate standardized data; the correlation analysis of the terminal-side indicators and network-side indicators in the standardized data is performed to determine whether the preset cross-domain correlation rules are met; when the correlation between the terminal-side indicators and the network-side indicators meets the cross-domain correlation rules, cross-domain correlation alarm information is generated.

2. The intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration according to claim 1, characterized in that, The cross-domain association rules include: when the GPU utilization of the computing node is lower than a first preset threshold, the PFC pause frame transmission count of the computing node's network card exceeds a second preset threshold, and the outbound queue depth of the switch port directly connected to the computing node's network card exceeds a third preset threshold, it is determined that the association rule of network congestion leading to a decrease in computing performance is met; the cross-domain association alarm information is used to indicate that slow training is suspected to be due to downstream network congestion; the PFC count includes the PFC pause frame transmission count.

3. The intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration according to claim 1, characterized in that, Also includes: S3, store the structured information in the standardized data into a relational database, and store the time series data into a time series database.

4. The intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration according to claim 1, characterized in that, Also includes: S4 predicts the time series of the collected optical module received power or disk SMART data, and generates predictive alarm information when the predicted value in the future preset time period is lower than the preset alarm threshold.

5. The intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration according to claim 1, characterized in that, Also includes: S5, Receive LLDP neighbor information from network devices, and construct the actual physical topology based on the LLDP neighbor information; Import the preset planned topology, compare the device connection relationship between the actual physical topology and the planned topology for consistency, and generate a topology verification result; The planned topology includes device connection relationships and track alignment rules, wherein the track alignment rules include the correspondence between switch port groups and GPUs within computing nodes; the topology verification results include connection matching results and port status results.

6. The intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration according to claim 5, characterized in that, The consistency comparison method includes: S51, for each switch port involved in the same track rule, determine the target computing node connected to the port and the corresponding network card identifier according to the LLDP neighbor information; S52, obtain the mapping relationship between the network card identifier and the associated GPU identifier; S53, compare the GPU identifier in the mapping relationship with the GPU set in the same track rule; if the GPU identifier in the mapping relationship does not belong to the GPU set specified by the same track rule, the same track verification is determined to fail, and the violation information is recorded in the topology verification result.

7. The intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration according to claim 1, characterized in that, Also includes: S6, Receive RDMA network performance metrics from the network interface card of the computing node; Using the RDMA network performance metrics as feedback, an improved simulated annealing algorithm is run to dynamically search for the optimal DCQCN parameter combination; the optimal DCQCN parameter combination is then sent to the target network card to complete the parameter configuration. The RDMA network performance metrics include at least one of traffic, round-trip time (RTT), and PFC count.

8. The intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration according to claim 7, characterized in that, The improved simulated annealing algorithm employs at least one of the following mechanisms for optimal parameter search: multi-parameter joint adjustment, variable step size search, and multi-round annealing.

9. The intelligent operation and maintenance method for intelligent computing centers based on end-to-end network collaboration according to claim 1, characterized in that, Also includes: S7. According to the preset linkage strategy, a node marking instruction is sent to the resource scheduling system to mark the computing node that generated the cross-domain association alarm information as unhealthy, so as to prevent new training tasks from being assigned to the node.

10. An intelligent operation and maintenance system for a smart computing center based on end-to-end network collaboration, characterized in that, include: The data access layer, deployed on the computing nodes, is used to: collect computing status data and network status data from the computing node hardware and network devices, respectively; the computing status data includes the GPU utilization of the computing node; the network status data includes network interface card (NIC) data, switch data, and optical module data; the NIC data includes at least one of the following: NIC PFC count, DCQCN status, and traffic; the switch data includes the switch queue depth; and the optical module data includes the optical module's received optical power. The logic layer is used to: parse, clean, and convert the computational state data and the network state data to generate standardized data; The correlation analysis is performed on the terminal-side indicators and network-side indicators in the standardized data to determine whether the preset cross-domain correlation rules are met; when the correlation between the terminal-side indicators and the network-side indicators meets the cross-domain correlation rules, cross-domain correlation alarm information is generated.