RDMA network monitoring method and system in cloud native artificial intelligence system
By deploying switches and cluster monitoring modules on cloud-native platforms and combining resource identity information for RDMA network monitoring, the problem of traditional monitoring not being able to identify container-level faults is solved, fine-grained network monitoring and fault location are achieved, and the stability and security of AI tasks are improved.
Patent Information
- Application Number
- CN202510898720.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional RDMA network monitoring cannot penetrate into the container or even AI task level, resulting in the failure cause being unable to accurately identify when AI task performance is poor or network communication is abnormal. The existing technology stack is very different from cloud-native platforms and cannot be directly applied.
It provides an RDMA network monitoring method in cloud-native artificial intelligence system. It collects traffic data through the switch monitoring module and cluster monitoring module, combines resource identity information for identity mapping, and converts it into cloud-native platform network index format for analysis, realizing fine-grained network monitoring.
Build a comprehensive and traceable security observation system to ensure the security and controllability of AI tasks network communication, support multi-tenant and microservice architectures, have good flexibility and compatibility, and are suitable for large-scale AI cloud platform deployment.
Smart Images

Figure CN120498880A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of cloud native technology, and in particular to an RDMA network monitoring method and system in a cloud native artificial intelligence system. Background Art
[0002] Cloud Native is a modern software architecture approach that fully leverages the advantages of cloud computing, while the Cloud Native AI System fully containerizes, automates, and visualizes AI training, reasoning, and service processes, and relies on cloud native platforms such as Kubernetes to achieve efficient, secure, and scalable AI lifecycle management.
[0003] Typically, cloud-native AI systems require high-performance network and storage support capabilities, such as improving model training efficiency by supporting the Remote Direct Memory Access (RDMA) protocol.
[0004] However, traditional RDMA network monitoring observes physical devices or virtual networks as units, and can often only see traffic behavior at the device level or port level, but cannot go deep into the container or even AI task level. This leads to poor AI task performance or the inability to accurately identify the cause of the fault when AI network communication anomalies occur.
[0005] In other existing technologies, the IO performance monitoring capability of RDMA applications is provided by improving the user-mode driver library provided by RDMA device manufacturers. Specifically, performance data collection logic is inserted into the key interfaces of the user-mode driver library (sending work requests WR to the send queue SQ / receive queue RQ, polling work completion WC), and then these performance indicators are exposed to the outside through the user-mode performance statistics service. However, the providers of this solution are often hardware solution providers. They often only focus on the performance indicators of RDMA applications at the network level and focus on performance collection within a single type and device. The RDMA network on the cloud-native platform may involve multiple types of RDMA hardware devices. The cost of upgrading the driver library for multiple types of RDMA devices is too high. In addition, the technology stack for improving the RDMA device driver is also very different from the technology stack of the cloud-native platform itself, and cannot be directly applied to RDMA network monitoring of cloud-native artificial intelligence systems.
[0006] Therefore, it is necessary to provide an improved technical solution to the above-mentioned deficiencies in the prior art. Summary of the Invention
[0007] The purpose of this application is to provide an RDMA network monitoring method and system in a cloud-native artificial intelligence system to solve or alleviate the problems existing in the above-mentioned prior art.
[0008] In order to achieve the above objectives, this application provides the following technical solutions:
[0009] In a first aspect, the present application provides an RDMA network monitoring method in a cloud-native artificial intelligence system. The method is performed by a network monitoring component, which is containerized and deployed on each node of the cloud-native platform. The network monitoring component includes: a switch monitoring module and a cluster monitoring module. The method includes:
[0010] The switch monitoring module collects traffic data of each RDMA switch to obtain first traffic data; the cluster monitoring module collects RDMA traffic data on each node of the cloud native platform to obtain second traffic data;
[0011] The network monitoring component obtains resource identity information of the cloud native platform, and identifies the flow identities of the first flow data and the second flow data based on the resource identity information, so as to establish an identity mapping relationship between the first flow data, the second flow data and each resource of the cloud native platform;
[0012] Based on the identity mapping relationship, the first traffic data and the second traffic data are converted into the data format of the cloud native platform network indicator, so as to analyze and process the first traffic data and the second traffic data to obtain the monitoring result of the RDMA network.
[0013] In some possible implementations, the switch monitoring module collects traffic data of each RDMA switch to obtain first traffic data, including:
[0014] The switch monitoring module is configured with the IP address of each RDMA switch, and the switch monitoring module is connected to all RDMA switches based on the IP address;
[0015] The switch monitoring module collects the flow data of each RDMA switch itself based on the sFlow protocol to obtain third flow data; and / or,
[0016] The switch monitoring module collects port flow data of each RDMA switch based on the gNMI protocol to obtain fourth flow data;
[0017] The third flow data and the fourth flow data are collectively referred to as first flow data.
[0018] In some possible implementations, the cluster monitoring module collects traffic data of RDMA devices on each node of the cloud native platform to obtain second traffic data, including:
[0019] The cluster monitoring module obtains RDMA traffic data from the host and container network cards on each node through periodic polling, which is recorded as fifth traffic data; and / or,
[0020] The cluster monitoring module uses eBPF technology to monitor RDMA read and write events on each host and obtains corresponding process information, which is recorded as the sixth data;
[0021] The fifth flow rate data and the sixth data are collectively referred to as second flow rate data.
[0022] In a possible implementation, the network monitoring component further includes: a GUI module, and the method further includes:
[0023] The switch monitoring module obtains network neighbor topology information of each RDMA switch based on the LLDP protocol, converts the information into a data format of a cloud native platform network indicator, records the information as first topology information, and sends the first topology information to the API server of the cloud native platform;
[0024] The cluster monitoring module collects network neighbor information of all network cards on each node of the cloud native platform based on the LLDP protocol, converts the information into a data format of cloud native platform network indicators, records the information as second topology information, and sends the second topology information to the API server of the cloud native platform;
[0025] The GUI module reads the first topology information and the second topology information from the API server of the cloud native platform, and draws the interconnection relationship between all network devices based on the first topology information and the second topology information to obtain a network topology map.
[0026] In some possible implementations, the first traffic data and the second traffic data include the following five-tuple information: source IP address, destination IP address, protocol type, source port, and destination port; the method also includes: the GUI module, based on the five-tuple information of the first traffic data and the second traffic data, combined with the identity mapping relationship, draws all forwarding paths of the communication flows corresponding to the first traffic data and the second traffic data, and then draws a communication path diagram of any AI task.
[0027] In one possible implementation, the first traffic data and the second traffic data are analyzed and processed to obtain monitoring results of the RDMA network, including: using the network topology diagram and the communication path diagram, combined with the first traffic data and the second traffic data shown in the diagram, to identify communication hotspots and abnormal ports in the network, track the entire network link of AI tasks, identify illegal traffic and process identities, monitor the congestion of the RDMA network, monitor and locate the health status of all network devices, and the above results are collectively referred to as the monitoring results of the RDMA network.
[0028] In one possible implementation, the network monitoring component further includes: an alarm module, which implements an alarm in a cloud native platform manner when the monitoring result of the RDMA network reaches a preset alarm condition.
[0029] In a second aspect, this embodiment provides an RDMA network monitoring system in a cloud-native artificial intelligence system. The system includes a network monitoring component, which is containerized and deployed on each node of the cloud-native platform. The network monitoring component includes a switch monitoring module and a cluster monitoring module. The system includes:
[0030] The collection unit is configured to collect the flow data of each RDMA switch by the switch monitoring module to obtain first flow data; the cluster monitoring module collects the flow data of the RDMA device on each node of the cloud native platform to obtain second flow data;
[0031] A mapping unit is configured to obtain resource identity information of the cloud native platform from the network monitoring component, and identify the flow identities of the first flow data and the second flow data based on the resource identity information to establish an identity mapping relationship between the first flow data, the second flow data and the AI workload;
[0032] The analysis unit is configured to convert the first traffic data and the second traffic data into the data format of the cloud native platform network indicator based on the identity mapping relationship, so as to analyze and process the first traffic data and the second traffic data to obtain the monitoring results of the RDMA network.
[0033] In a third aspect, this embodiment provides a computer device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of any of the above methods.
[0034] In a fourth aspect, this embodiment provides a computer-readable storage medium having a computer program / instruction stored thereon, characterized in that the computer program / instruction implements the steps of any of the above-described methods when executed by a processor.
[0035] The technical solution of the embodiment of the present application has the following beneficial effects:
[0036] By deploying network monitoring components in a containerized manner on each node of the cloud native platform and combining traffic collection from two dimensions, namely RDMA switch traffic and cluster RDMA traffic, it is possible to collect RDMA traffic data (secondary traffic data) for specific containers, Pods, and even AI tasks on each node. Simultaneously, combined with macro-traffic data (primary traffic data) at the RDMA switch level, through multi-dimensional RDMA network and container data collection and analysis, a comprehensive, fine-grained, and traceable security observation system is constructed, starting from AI tasks and spanning the cloud native platform and physical network devices. This effectively ensures the security and controllability of AI tasks during network communication and data interaction, helping administrators fully understand the operating status of RDMA networks and RDMA devices in the cloud native platform and quickly identify and resolve potential issues. Furthermore, the containerized deployment of network monitoring components is naturally compatible with orchestration systems such as Kubernetes, and can dynamically adjust the monitoring scope as nodes automatically scale in and out. It supports complex AI application deployment scenarios under multi-tenant and microservice architectures, and has good elasticity, compatibility, and maintainability, making it suitable for large-scale AI cloud platform deployment. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A flowchart of an RDMA network monitoring method in a cloud-native artificial intelligence system provided according to some embodiments of the present application.
[0038] Figure 2 Schematic diagram of the network architecture of the cloud-native artificial intelligence system.
[0039] Figure 3 This is an example of visualizing basic network indicator data in a network topology diagram.
[0040] Figure 4 An example of network congestion shown in a network topology diagram.
[0041] Figure 5 Design a schematic diagram of the functional modules of the network monitoring component.
[0042] Figure 6 An example of a communication path graph for a set of AI tasks.
[0043] Figure 7 An example of a functional architecture block diagram is shown in Figure 1. DETAILED DESCRIPTION
[0044] Cloud-native technologies typically include containerization, microservices, and declarative APIs. Containerization uses containers to package applications and their dependencies to achieve environmental consistency. Microservices split monolithic applications into multiple independent, loosely coupled small services for rapid iteration and deployment. Declarative APIs and immutable infrastructure define application configurations declaratively, allowing the system to automatically maintain the desired state. Infrastructure (such as containers) is immutable, and errors are resolved by replacing it directly rather than repairing it.
[0045] A typical implementation of cloud-native technology is Kubernetes (abbreviated as K8s), an open source container orchestration platform. Containers run in Kubernetes in the form of Pods. Pods are the smallest deployment unit in Kubernetes and usually contain one or more containers.
[0046] Cloud Native AI System: refers to an AI platform or system built on a cloud-native architecture that distributes, schedules, and manages AI models / applications (such as deep learning training, online inference, and model services) as cloud-native workloads. Cloud-native AI systems centrally manage AI container tasks through cloud-native platforms (such as Kubernetes), supporting the scheduling of heterogeneous computing resources such as GPUs and TPUs. This containerizes AI workloads, meaning that AI training tasks, inference services, data preprocessing, and other modules are all deployed in containers. This facilitates version control, isolation, and expansion, and facilitates elastic scaling and dynamic resource allocation for AI applications. For example, they can automatically scale capacity based on load and automatically request more GPU resources during peak training periods.
[0047] For ease of description, this embodiment will be described using Kubernetes as a typical cloud-native platform. Kubernetes is an open-source container orchestration platform that automates the deployment, management, and scaling of containerized applications. It manages containers across multiple hosts, providing high availability, load balancing, and resource management. Container orchestration refers to the automated management, scheduling, and coordination of the deployment and operation of multiple containers in large-scale environments, ensuring high availability, scalability, and fault recovery for containerized applications.
[0048] The Remote Direct Memory Access (RDMA) protocol, with its kernel-bypassing nature, allows direct data reading and writing between applications and network cards. It can transfer data directly from the memory of one computer to another without the intervention of either operating system, circumventing the limitations of the Transmission Control Protocol / Internet Protocol (TCP / IP). It can effectively reduce communication latency and CPU utilization, and is widely used in in-memory databases, distributed storage, high-performance computing, and other fields. The current industry best practice is to run large-scale AI models on RDMA networks.
[0049] Among them, the RDMA network includes smart network cards (RNICs) that support the RDMA protocol and switches that support the RDMA protocol. Among them, RDMA switches are further divided into switches based on RoCE technology (referred to as RoCE switches) and switches based on the InfiniBand protocol (referred to as InfiniBand switches).
[0050] An AI big model refers to an artificial intelligence model with a massive number of parameters, typically using deep learning networks to handle complex tasks. An AI application refers to a complete, business-oriented AI system or service, typically composed of multiple modules, including data processing, model training, and inference services. It is a macro-level concept that represents the functionality and purpose of the entire AI system and can include multiple AI workloads. For example, an AI application could be an intelligent customer service system, an image recognition platform, or a recommendation system. An AI workload refers to the type of computing task carried out by a specific stage or functional module in an AI application. It is the basic unit of resource consumption and execution logic. Common AI workloads include model training and online / offline inference. An AI task is a specific instance of an AI workload, representing a specific execution operation or process. It has time and status characteristics (such as started, running, completed, or failed). It can typically be scheduled by a scheduler (such as Kubernetes) to different nodes and executed by containers. AI container refers to packaging AI tasks in containers (such as Docker containers) for deployment and operation. It is a key technical means to realize cloud-native AI. It is a runtime encapsulation form that contains the code, dependency libraries, configuration files, etc. required for AI tasks. It supports rapid deployment, isolation, version control, elastic scaling, and can run on container orchestration platforms such as Kubernetes. AI training refers to the training of artificial intelligence systems through large amounts of data and algorithm models to enable them to recognize patterns, make predictions or make decisions. It usually involves deep learning and machine learning. AI reasoning refers to the process of using existing models to predict or make decisions on new input data after training. In actual applications, it is usually used for real-time or batch prediction tasks.
[0051] The emergence of large models has enabled AI to demonstrate enhanced understanding and generative capabilities in multiple fields. Training these large models requires extensive computing resources and data, typically utilizing distributed training and parallel computing techniques. However, training large models faces challenges such as high computational costs, long training cycles, and energy consumption. During the training and inference of large-scale AI models running on cloud-native AI systems, network stability has a crucial impact on overall efficiency and cost control. For example, training the open-source AI model Llama3 405B (a large language model released by Meta) took 45 days. Failures in network switches and cables caused 35 training interruptions, accounting for 8.4% of these interruptions, each resulting in significant costly losses. Traditional RDMA network monitoring observes traffic behavior at the device or port level, often only at the device or port level. This makes it difficult to accurately identify the root cause of these issues, resulting in a delay in quickly and effectively troubleshooting them, severely impacting the normal operation of AI applications. Therefore, a comprehensive and secure monitoring capability, focused on AI applications, is needed for cloud-native AI systems to improve the quality of service for AI applications on cloud-native platforms.
[0052] Existing cloud-native platforms, such as Kubernetes, typically offer a comprehensive suite of monitoring products (also known as native network observability products, such as the Prometheus component). These products provide fine-grained network connection tracking, helping operations personnel visualize traffic paths within Kubernetes clusters, analyze network interactions between microservices, and detect abnormal connections, network jitter, and packet loss. However, current cloud-native network observability products primarily focus on the TCP / IP protocol, using eBPF (Extended Berkeley Packet Filter) technology to directly collect TCP / IP-related data in kernel mode. However, RDMA operates in user mode and bypasses the kernel for communication, making it difficult for traditional eBPF-based data collection methods to directly capture RDMA traffic data. Consequently, monitoring products offered by existing cloud-native platforms offer limited support for RDMA. Monitoring RDMA networks relies more on hardware (such as InfiniBand switches dedicated to traffic monitoring), RDMA statistics tools, or vendor-specific solutions.
[0053] As a multi-tenant platform, Kubernetes supports different tenants and applications running in the same cluster. However, container IP addresses are dynamically assigned and frequently change, making it difficult for traditional network monitoring methods to accurately identify container-level traffic attribution. Furthermore, as previously mentioned, in RDMA networks, since RDMA communication typically uses RoCE (RDMA over Converged Ethernet) or InfiniBand protocols, data transmission bypasses the kernel, making it impossible to determine which container (pod) a particular traffic belongs to. Furthermore, traditional RDMA switch monitoring capabilities primarily focus on device- or port-level traffic statistics and congestion control metrics (such as ECN marking, packet loss rate, and bandwidth utilization), lacking the service topology information required by the Kubernetes ecosystem. Consequently, RDMA switch monitoring software lacks deep integration with Kubernetes and container environments, making it impossible to associate RDMA traffic with specific applications, tenants, or pods. This makes it difficult for operations personnel to track the specific source and destination of traffic when troubleshooting RDMA-related network issues, limiting observability in cloud-native environments.
[0054] Currently, there are many reasons for AI training / inference interruptions, among which network equipment failure is one of the main reasons. Especially in large-scale AI distributed training, since distributed AI training usually relies on high-speed network equipment with RDMA protocol for communication between nodes, if the network fails (such as switch failure, bandwidth congestion or link packet loss), it may cause data transmission interruption or delay during training, thereby affecting the synchronous update of the model and computing efficiency. Therefore, monitoring and ensuring the stability of network equipment is crucial to avoid AI training interruptions and improve training efficiency. Research has found that in cloud-native AI systems, RDMA network monitoring with AI applications faces the following challenges:
[0055] (1) Unsafe network traffic may be generated in the RDMA network. It may come from illegal external injections or from some illegal containers in the cloud native platform (such as Kubernetes). On the one hand, this illegal traffic will cause congestion in the RDMA network, which will affect the performance of the AI workload and even cause it to fail. On the other hand, illegal traffic can be disguised as legitimate traffic, affecting the reasoning and training results of AI during distributed AI training and inference, thus causing data security issues. Therefore, it is an important security issue to accurately identify the identity of each data flow in the network and associate it with the identity of the host, container, and process in the Kubernetes platform to prevent illegal traffic from interrupting AI training.
[0056] (2) When the AI workload itself encounters model or communication library issues, some traffic anomalies will occur during network communication. This can cause the AI workload to fail at best, or even affect the normal network communication of other AI tasks. Therefore, identifying network communication anomalies of AI tasks and preventing AI training interruptions caused by the AI workload itself is also an important security issue.
[0057] In other words, factors affecting network stability may originate from the RDMA network or from the AI workload itself. The lack of traffic monitoring in any of these aspects makes it difficult to support the ability to reshape AI communication paths. Currently, the relevant network monitoring products available in the industry come from two main aspects: on the one hand, products from RDMA switch / network card manufacturers do not associate RDMA traffic in the network with the identity information of AI workloads in Kubernetes, and do not organically integrate switches and Kubernetes platforms. On the other hand, monitoring products from cloud-native platforms do not provide RDMA network communication monitoring capabilities, cannot obtain data from the RDMA switch side, and cannot perform data correlation. They lack complete security monitoring capabilities from the perspective of AI applications, cannot achieve more fine-grained dangerous traffic identification, and cannot determine whether illegal containers are participating in RDMA network communication, which in turn affects the data security and network communication security of AI tasks. Therefore, how to use cloud-native methods to achieve comprehensive monitoring of RDMA networks while providing more fine-grained (such as container-level) and more comprehensive traffic tracking for AI tasks has become an urgent problem to be solved. In view of this, this application provides a technical solution for RDMA network monitoring in cloud-native artificial intelligence systems. Through independently developed network monitoring components, it provides cloud-native artificial intelligence systems with complete security monitoring capabilities from the perspective of AI applications, and can be deeply integrated with cloud-native platforms.
[0058] The terms "first", "second", "third" and "fourth" in the specification, claims and drawings of this application are used to distinguish different objects rather than to describe a specific order.
[0059] The embodiments of the present application are described below with reference to the accompanying drawings.
[0060] Figure 2 The network architecture of the cloud-native AI system is shown, which is also the best practice of hardware topology in the current industry. Figure 2As shown in the figure, this architecture is a spine-leaf topology consisting of spine switches and leaf switches. Spine switches serve as the core backbone, forwarding traffic across leaf switches. Leaf switches are directly connected to hosts or GPU nodes, providing high-density access capabilities. There are multiple spine switches, for example, 64, numbered spine switch 1 to spine switch 64. Each spine switch connects to multiple leaf switches, such as leaf switch 1 to leaf switch 8. Each switch (spine and leaf switches) contains multiple ports and a corresponding IP subnet (such as 172.17.1.0 / 24). The cloud-native platform is managed by the Kubernetes system. As an efficient operating platform for artificial intelligence, the hardware infrastructure of the Kubernetes cluster is designed with the needs of high-performance computing and large-scale data processing in mind. The cluster consists of multiple nodes, each equipped with 8 GPUs (used to accelerate the calculation of artificial intelligence (AI) tasks) and 8 network interface cards (NICs). To meet the needs of large-scale AI distributed training, GPUs and NICs can be grouped into different logical units, called blocks (such as block 1 and block 2). Each block includes several nodes, and the nodes are numbered from node1 to node64. Therefore, a block includes a total of 64 nodes, providing the computing power of 512 GPUs, and each node runs an AI workload. At the same time, multiple storage servers provide distributed storage (such as NFS or Ceph) services for AI workloads, such as storage server 1 and storage server 2. The storage server communicates with each node through a storage switch (such as a dedicated RDMA storage switch).
[0061] To enable the cloud-native platform to provide complete security monitoring capabilities from the perspective of AI applications, this embodiment provides an RDMA network monitoring method in a cloud-native artificial intelligence system. The method is executed by a network monitoring component, which is containerized and deployed on each node of the cloud-native platform. The network monitoring component includes a switch monitoring module (SwitchMonitor) and a cluster monitoring module (Cluster Monitor). The method includes:
[0062] Step S1: The switch monitoring module collects traffic data from each RDMA switch to obtain first traffic data; the cluster monitoring module collects RDMA traffic data on each node of the cloud native platform to obtain second traffic data.
[0063] The method provided in this embodiment is executed by a network monitoring component, which is a component running on a cloud-native platform independently developed to achieve the technical goals proposed in this application. It is a collection of various software modules, services or applications that implement RDMA network monitoring.
[0064] In this embodiment, the network monitoring component is deployed in a containerized manner on each node of the cloud native platform. Containerized deployment refers to the use of container technology (such as Docker) to package software modules, applications, and their dependencies (running environment) together in a container so that they can run consistently on any host that supports container operation.
[0065] It should be noted that a cloud native platform (such as Kubernetes) is usually a cluster composed of multiple nodes. The nodes of the Kubernetes cluster are divided into two categories according to the different roles they assume: control nodes (Master Node) and worker nodes (Work Node). In this embodiment, in order to complete AI-level RDMA network monitoring, the network monitoring component needs to be deployed on the above two types of nodes in the cluster, that is, it needs to be deployed on each node. Therefore, this embodiment does not distinguish between the types of nodes and uses the term "node" to refer to them uniformly. The method of deploying the network monitoring component to each node enables the network monitoring component to monitor the RDMA network traffic of each node and container in the Kubernetes cluster, and at the same time monitor the network traffic of the RDMA switch outside the cluster, thereby realizing AI task-level RDMA network traffic monitoring.
[0066] Specifically, Figure 5 Design a schematic diagram for the functional modules of the network monitoring component. Figure 5 As shown in FIG, network monitoring components are deployed based on the hardware infrastructure.
[0067] The hardware infrastructure used in this embodiment includes an RDMA switch (RDMASwitch) and Kubernetes nodes (Kubernetes Nodes).
[0068] In this embodiment, RDMA switches include the aforementioned Spine and Leaf switches, as well as RDMA storage switches. The Spine and Leaf switches are part of the RDMA computing network for GPU communication, while the RDMA storage switch is part of the RDMA storage network used to load AI model data. These two types of RDMA networks are used in cloud-native platforms.
[0069] In terms of switch types, the RDMA switches in this embodiment include both RDMA switches based on the InfiniBand (IB) protocol and RDMA switches based on the RoCE protocol. InfiniBand switches and RoCE switches are currently two mainstream network protocol devices.
[0070] RoCE switches are based on standard Ethernet hardware but support RDMA communication. To support congestion monitoring, RoCE switches typically need to support flow control mechanisms such as PFC (Priority Flow Control) and ECN (Explicit Congestion Notification) to ensure low latency and high throughput for RDMA communication. Common RoCE switches are provided by manufacturers such as NVIDIA (Mellanox), Arista, Cisco, and Broadcom, and are primarily used for data center RDMA applications such as GPU server clusters and distributed storage.
[0071] InfiniBand switches are high-performance network devices designed specifically for RDMA. They utilize the InfiniBand protocol, offering lower latency, higher bandwidth, and greater scalability than traditional Ethernet. InfiniBand switches typically support high-speed links such as HDR (200Gbps) and NDR (400Gbps), and rely on IB routing protocols (such as LID routing and adaptive routing) for efficient traffic scheduling. Because InfiniBand uses proprietary protocols and hardware, it is more closed than the RoCE solution. Primarily provided by NVIDIA (Mellanox), it is widely used in supercomputing centers (HPC), AI training clusters, and high-frequency financial trading systems.
[0072] Containers and processes run on Kubernetes nodes. Nodes and containers can be configured with smart network cards (RDMA network cards, or NICs) that support the RDMA protocol.
[0073] Because factors affecting network stability may originate from both the RDMA network and the AI workload (Pod, Service, API) itself, in this embodiment, the network monitoring component includes two functional modules: a switch monitoring module and a cluster monitoring module. In step S1, the switch monitoring module is responsible for collecting traffic data from each RDMA switch to obtain first traffic data; the cluster monitoring module is responsible for collecting RDMA traffic data from each node on the cloud native platform to obtain second traffic data.
[0074] Traffic data (such as network packets or connection metadata) usually contains: source IP / destination IP, source port / destination port, protocol (RDMA / TCP / UDP / HTTP), timestamp, and other information. Among them, the first traffic data comes from RDMA switches (such as InfiniBand, RoCEv2 switches). Therefore, this type of traffic data usually only has L2 / L3 information (such as MAC / IP / Port) and no Kubernetes context information. In addition, the IP addresses of different containers (Pods) running the same application change dynamically, and traditional RDMA traffic cannot perceive these changes. The second traffic data can be collected by cloud-native monitoring of cluster nodes. This type of traffic data can carry Kubernetes context information, such as container process (PID), Pod IP, ServiceAccount, Namespace, etc.
[0075] In this embodiment, the switch monitoring module (Switch Monitor) and the cluster monitoring module (ClusterMonitor) are both functional modules independently developed by this embodiment. Among them, the Switch Monitor is deployed based on a set of Deployment components of the Kubernetes platform. It is mainly used to manage all RDMA switches in the network. Through the network management protocols supported by various switches, it obtains various traffic data of the switches (i.e., first traffic data, including metrics and logs), and implements data flow monitoring of the RDMA network, including: traffic of each port on the switch, health indicators of the RDMA switch device itself, etc. The Cluster Monitor is deployed based on a set of DaemonSet components of the Kubernetes platform. The core task of this module is to collect read and write indicators of the RDMA device of each node in the cluster, and at the same time collect RDMA activity events of all processes, and store this information (i.e., second traffic data) in the corresponding storage component for subsequent analysis by other components. The Switch Monitoring Module and the Cluster Monitoring Module collect first traffic data and second traffic data respectively, comprehensively capturing all RDMA traffic of the RDMA network, each node, and the AI workload on the node, laying the data foundation for reshaping the AI communication path and providing complete security monitoring capabilities from an AI perspective.
[0076] Specifically, Cluster Monitor can be used to monitor RDMA events / activity (including RDMA read and write events) of containers and processes, and monitor the RDMA network interface cards (NICs) of nodes / containers to generate RDMA traffic metrics.
[0077] It's important to note that DaemonSet and Deployment are two common workload resource types in Kubernetes, used to manage and run containerized applications. DaemonSet deployment ensures that a Cluster Monitor module runs on every node, ensuring comprehensive and complete collection of RDMA traffic data on each node. Deployment deployment ensures that the Switch Monitor module is scheduled to the appropriate node through Kubernetes' scheduling mechanism.
[0078] Step S2: The network monitoring component obtains the resource identity information of the cloud native platform, and identifies the traffic identities of the first traffic data and the second traffic data based on the resource identity information to establish an identity mapping relationship between the first traffic data, the second traffic data and the AI workload.
[0079] In this embodiment, the network monitoring component can obtain resource identity information through the control center of the cloud native platform, such as calling the relevant interface of the API Server component of Kubernetes.
[0080] It should be noted that in a Kubernetes cluster, resources refer to objects that can be defined, managed, and scheduled in the cluster. Resources are represented by Kubernetes API objects (APIObject), which are the basic units of cluster management and are defined through YAML or JSON files.
[0081] In this embodiment, resources include workload resources (such as Pods, Deployments, DaemonSets, and Jobs), as well as scheduling and node resources (such as Nodes). In AI systems, Pods are the basic unit for running AI tasks and, therefore, are also the resource form for executing AI workloads.
[0082] In this embodiment, resource identity information includes a resource unique identifier (resource ID) and a resource IP address. The resource ID can be, for example, a unique identifier of a resource generated by a cloud native platform (such as a PodID), or other fields that can uniquely identify a resource, such as a node name, a container group name (Pod name), etc. In addition, resource identity information can also be a combination of multiple fields. The combined fields can uniquely identify a resource object, such as the node ID + Pod IP address to identify a Pod running on a node. Optionally, resource identity information can also include information such as the namespace to which the resource belongs, the node where it is located, and labels (such as Labels: app=web).
[0083] Based on the resource identity information, the flow identities of the first flow data and the second flow data are identified. For example, the source / destination IP address (flow feature) can be extracted from the flow data and the corresponding resource identity (resource ID, resource name, etc.) can be found according to the IP address to establish an identity mapping relationship between the first flow data, the second flow data and the resource. The above process can also be described as identifying the source identity (such as the identity of the flow sender / resource) and the target identity (the identity of the flow receiver / resource) of the first flow data and the second flow data based on the resource identity information, so as to associate the flow identity with the resource identity, and at the same time, associate the flow identity with the AI workload according to the AI workload executed by the resource.
[0084] Step S3: Based on the identity mapping relationship, the first flow data and the second flow data are converted into the data format of the cloud native platform network indicator to analyze and process the first flow data and the second flow data to obtain the monitoring results of the RDMA network.
[0085] The data format of cloud-native platform network metrics can be, for example, the Prometheus-style metrics format provided by the Kubernetes cluster, or a standard format such as OpenTelemetry.
[0086] Since the data formats of different flows are different, in this embodiment, based on the identity mapping relationship, the first flow data and the second flow data are converted into the data format of the cloud native platform network indicator, which can be implemented as follows: first, it is necessary to parse the data structure of the first flow data and the second flow data, and extract their key information, such as devices, ports, indicators, values, etc., and then format the first flow data and the second flow data into a Prometheus-style indicator format or a standard format such as OpenTelemetry according to the identity of each flow.
[0087] After formatting the first and second traffic data into Prometheus-style metrics, these metrics can be exposed through cloud-native methods (such as Exporter or Gateway). Cloud-native network observability products (such as Prometheus and Grafana) can then be used to analyze and process them, yielding RDMA network monitoring results. These metrics can also be stored in the cloud-native platform's control center (such as an API server) for other platform components to access and use through APIs.
[0088] Among them, the monitoring results of the RDMA network include network performance indicators. These network performance indicators are obtained by extracting, aggregating and analyzing the formatted first flow data and second flow data, including container (corresponding to AI tasks) level indicators, switch level indicators (Switch Metrics), link level indicators (Link Metrics), and node level indicators (Node Metrics). Among them, container-level indicators are used to monitor the network performance and security identification of a group of Pods running AI tasks. For example, they may include the time nodes when each AI task communicates with each other, the amount of traffic generated by each AI task communication, the switch path through which the AI task communication data packet passes, etc.; the remaining indicators at all levels are used to monitor and securely identify the RDMA traffic of the corresponding devices, and the specific content will be explained later.
[0089] In summary, in this embodiment, the network monitoring components are containerized and deployed on each node of the cloud native platform. The cluster monitoring module can collect the RDMA traffic data (second traffic data) of specific containers, Pods, and even AI tasks on each node. At the same time, the switch monitoring module collects the macro traffic data (first traffic data) at the RDMA switch level, thereby realizing full-link fine-grained monitoring from the hardware layer to the application layer, and ultimately realizing fine-grained monitoring and fault location of the RDMA network at the container level and even the AI task level.
[0090] The network monitoring component obtains resource identity information (such as Pod name, namespace, AI task ID, etc.), and uses this information to tag the collected traffic (first traffic data, second traffic data) with identity tags, forming an identity mapping relationship between traffic and resources, and AI workloads, and clarifying which AI task or container each RDMA traffic segment belongs to, facilitating subsequent analysis and problem tracing. The traffic data from switches and clusters is uniformly converted into the data format of cloud-native platform network indicators (such as Prometheus, OpenTelemetry and other standard formats), supporting subsequent unified analysis, alarms, and visualization, improving data processing efficiency, facilitating integration into existing observability systems, and improving the level of operation and maintenance automation. Based on identity mapping and fine-grained traffic data, the performance of a specific AI task or container in network communication can be tracked, and the specific sources of problems such as high communication latency, insufficient bandwidth, congestion points, and high packet loss rates can be identified, significantly improving troubleshooting efficiency and ensuring the stability and performance of AI training / inference tasks. In addition, the containerized deployment of network monitoring components is naturally compatible with orchestration systems such as Kubernetes. It can dynamically adjust the monitoring scope as nodes automatically scale up and down, support complex AI application deployment scenarios under multi-tenant and microservice architectures, and has good elasticity, compatibility, and maintainability, making it suitable for large-scale AI cloud platform deployment.
[0091] In other words, in this embodiment, by capturing the network indicator data of the cloud native platform (such as Kubernetes) network and RDMA switch, based on the IP information of the message, the messages in all network devices are associated with the container identity, thereby implementing full-link message observability at the AI container level. Ultimately, in the cloud native artificial intelligence platform, in-depth monitoring of the GPU computing power network and the model data storage network is implemented.
[0092] In some embodiments, the switch monitoring module collects traffic data of each RDMA switch to obtain first traffic data, including:
[0093] Step S11a: The switch monitoring module is configured with the IP address of each RDMA switch, and the switch monitoring module connects with all RDMA switches based on the IP address.
[0094] The switch monitoring module maintains a list of all switches in the RDMA network by configuring the IP addresses of each RDMA switch. Therefore, based on the configured IP addresses, it can initiate connections to these switches through different protocols on a scheduled or real-time basis to obtain corresponding traffic data.
[0095] Step S11b: The switch monitoring module collects the flow data of each RDMA switch itself based on the sFlow (Sampling Flow) protocol to obtain third flow data.
[0096] Among them, the sFlow protocol is a technology used for network traffic monitoring and statistical analysis. The implementation of the sFlow protocol includes the coordinated execution of the sFlow Agent and the sFlow Collector. RDMA switch devices are usually embedded with the sFlowAgent to collect network traffic data. The agent uses a passive collection mechanism of flow sampling (Flow Sampling) and counter sampling (Counter Sampling). The collected data is encapsulated into sFlow packets via the UDP protocol and sent to the remote sFlow Collector for analysis. Currently, mainstream switch manufacturers (such as NVID IAMellanox, Arista, Cisco, Juniper, and Broadcom) all support the sFlow Agent and provide corresponding management interfaces to configure the sampling rate and data collection strategy. In this embodiment, the sFlow Collector function is implemented by writing code in the switch monitoring module. The sFlow Collector converts the collected traffic data (third traffic data) into the Metrics format in Kubernetes and stores it in the Prometheus component. In other words, in RDMA switches, sFlow can be used to monitor RoCE traffic and analyze congestion control-related indicators such as ECN marking rate and PFC trigger times. Then, based on the sFlow data collected from the switch, the five-tuple information of the data packet on each port (source IP address, destination IP address, protocol type, source port, destination port) is obtained. At the same time, the IP information of all resources in Kubenetes (mainly containers, i.e. Pods) is monitored. Thus, the container identity association is implemented for the five-tuple information of the switch, the IP address is associated with the container name, and finally the metrics data in Kubenetes is output.
[0097] The following is an example of using the sFlow protocol to collect network traffic:
[0098] sflow_tx_packets{group="gpu",ip="10.193.77.204",port="Ethernet184",switch="gpu-leaf-switch-
[0099] 4",source_ip="172.16.1.200",dest_ip="172.16.1.300",source_port="3000",dest_port="30001",protocol="udp",container_sour ce_name="ai-job1",container_source_interface="net1",container_dest_name="ai-job1",container_dest_interface="net1"}100
[0100] Extract the packet sending statistics value 100 of a set of five-tuple communication flows (source_ip="172.16.1.200",dest_ip="172.16.1.300",source_port="3000",dest_port="30001",protocol="udp") in the port Ethernet184 of the RDMA switch (switch=gpu-leaf-switch-4) given in the above example (where source_ip is the source IP address, source_port is the source port, dest_ip is the destination IP address, and dest_port is the protocol). Using the source IP address and resource identity information from the Kubernetes cluster, we can identify the sender of this communication flow as the network interface card (NIC) container_source_interface = "net1" for the container with name "container_source_name = "ai-job1" and the receiver as the network interface card (NIC) container_dest_interface = "net1" for the container with name "container_dest_name = "ai-job1"). This allows us to map the traffic data to the identities of resources (i.e., containers) in the cloud native platform. This data association allows us to connect all network forwarding paths for this communication flow, creating a communication path map for a specific AI task. This also allows us to identify any illegal packets in the switch network.
[0101] And / or, in step S11c, the switch monitoring module collects port flow data of each RDMA switch based on the gNMI protocol to obtain fourth flow data.
[0102] The gNMI protocol (gRPC Network Management Interface) is a network management protocol based on the gRPC framework, providing efficient device status acquisition, configuration, and monitoring capabilities. Built on gRPC (Google Remote Procedure Call), gNMI leverages HTTP / 2 for efficient network communication and uses Protocol Buffers as a data serialization mechanism to ensure low latency and efficient data transmission.
[0103] In this embodiment, the switch monitoring module uses the gNMI protocol to obtain port indicators of switches and routers, and stores them in the Prometheus component in the form of metrics in Kubernetes. At the same time, the basic information of the device is stored in the Kubernetes API Server through Custom Resource Definition (CRD) for further analysis by other components. In this way, network devices can be incorporated into the Kubernetes resource system through CRD to support scheduling / policy control, etc.
[0104] Step S11d: The third flow data and the fourth flow data are collectively referred to as first flow data.
[0105] In the above embodiment, by simultaneously supporting gNMI (active pull) + sFlow (passive reception), the coverage of RDMA switch traffic data is improved.
[0106] For example, based on the gNMI protocol, the RDMA switch can report key port statistics in real time, such as throughput, packet loss rate, latency, error packet count, PFC (priority flow control) status, and ECN (explicit congestion notification) marking, and convert them into metrics data in the metrics format in Kubernetes. The following is a set of output examples (lines starting with # indicate comments, the same below):
[0107] #The following data shows that 163815 packets have been sent out of port Ethernet184 on switch gpu-leaf-switch-4;
[0108] unifabric_tx_packets{group="gpu",ip="10.193.77.204",port="Ethernet184",switch="gpu-leaf-switch-4"}163815
[0109] #The following data shows that 35289 packets have been received on port Ethernet184 of switch gpu-leaf-switch-4;
[0110] unifabric_rx_packets{group="gpu",ip="10.193.77.204",port="Ethernet184",switch="gpu-leaf-switch-4"}35289
[0111] #The following data shows that 0 packets were discarded on port Ethernet184 of switch gpu-leaf-switch-4;
[0112] port_drop_packets{group="gpu",ip="10.193.77.204",port="Ethernet184",switch="gpu-leaf-switch-4"}0
[0113] The following data shows that the number of RDMA ECN congestion packets monitored on port Ethernet184 of switch gpu-leaf-switch-4 is 0.
[0114] wred_ecn_marked_packets{group="gpu",ip="10.193.77.204",port="Ethernet184",switch="gpu-leaf-switch-4"}0
[0115] In some embodiments, the cluster monitoring module collects traffic data of RDMA devices on each node of the cloud native platform to obtain second traffic data, including:
[0116] Step S12a: The cluster monitoring module obtains RDMA traffic data from the host and container network cards on each node through periodic polling, which is recorded as fifth traffic data.
[0117] On every node in a Kubenetes cluster, each container either shares the host's network namespace or has its own independent network namespace. Containers require RDMA network cards, which can be used to communicate by sharing the host's network namespace and network devices, or by using their own network namespace and SR-IOV monitoring. Therefore, the cluster monitoring module of the network monitoring component collects read and write metrics (indicating information related to read and write events) for RDMA devices (host RDMA network card + container RDMA network card) in all network namespaces on the host. These metrics are then converted into metrics data by associating them with the corresponding container's identity, IP address, and network card name.
[0118] The cluster monitoring module actively queries the RDMA device status on each node at fixed time intervals (such as every second or every few seconds). The polling objects include: host level: RDMA network cards of each node, and container level: RDMA resources used by AI tasks running in containers. The collection content includes: key performance indicators such as send / receive traffic, bandwidth utilization, latency, and packet loss. The collection results are recorded as the fifth traffic data. Through periodic collection and continuous tracking of the operating status of RDMA devices, problems such as network congestion, bandwidth saturation, and increased latency can be discovered in a timely manner.
[0119] And / or, in step S12b, the cluster monitoring module uses the eBPF technology to monitor the RDMA read and write events on each host, and obtains corresponding process information, which is recorded as sixth data.
[0120] It's important to note that RDMA allows hosts to directly access each other's memory, avoiding the overhead of traditional network protocol stacks. RDMA read and write events primarily refer to operations such as RDMAWrite (a local host writes data to a remote host's memory) and RDMARead (a local host reads data from a remote host's memory). These RDMA read and write events are triggered indirectly through the uverbs interface in the RDMA stack (e.g., ibv_post_send() is indirectly triggered by uverbs_ioctl()). ibv_post_send() is a designated function in user-space applications. Therefore, the eBPF kprobe (hook) can monitor the execution of the ibv_post_send() function. When an application calls the RDMA API, the eBPF kprobe captures all RDMA send / write / read operations, obtains information such as the function's parameters, return value, and the calling process's PID, and passes this data to a kernel buffer (e.g., BPF_PERF_OUTPUT). Ultimately, the data is sent to userspace for further analysis via the perf_event mechanism. Every time an RDMA operation event occurs, the detailed context information of the process that initiated the operation can be further identified and extracted to achieve tracking of the person responsible for the operation and more fine-grained network behavior analysis. The obtained process information can include the following fields: PID (Process ID (ProcessID) that initiated the RDMA operation), container_id (ID of the container to which the process belongs (such as in Kubernetes)), cgroup (cgroup path (often used to determine the namespace)), namespace (Kubernetes Namespace to which it belongs). In addition, the name of the Pod to which the process belongs can be inferred through the container label. After the above RDMA events are collected, they will be recorded and stored in elasticsearch as logs, and other components will perform security analysis later.
[0121] Here is an example of a set of read and write metrics for an RDMA device:
[0122] #The following data shows that the RDMA network card ifname="mlx5_12" of the container pod_name="mx-rdma-test-rdma-tools-7ghbv" on the host node_name="sh-cube-worker-1" received 60 packets.
[0123] rdma_rx_vport_unicast_bytes_total{ifname="mlx5_12",is_root="false",net_dev_name="net3",node_guid="ec:a7:ad:fe:ff:21:4 8:d2",node_name="sh-cube-worker-1",otel_scope_name="spiderpool-agent",otel_scope_version="1.24.0",owner_api_version=" apps / v1",owner_kind="DaemonSet",owner_name="mx-rdma-test-rdma-tools",owner_namespace="rdma",pod_name="mx-rdma-test-rd ma-tools-7ghbv",pod_namespace="rdma",port="1",rdma_parent_name="ens841np0",sys_image_guid="64:19:be:00:03:e1:a2:58"}60
[0124] #The following data shows that the host node_name="sh-cube-worker-1" network card net_dev_name="ens1np0" generated 0 RDMA congestion control packets (CNPs), indicating that RDMA network communication is normal and there is no congestion.
[0125] rdma_np_cnp_sent_total{ifname="mlx5_0",is_root="true",net_dev_name="ens1np0",nod e_guid="6c:20:be:00:03:e1:a2:58",node_name="sh-cube-worker-1",otel_scope_name="spiderpool-agent",otel _scope_version="1.24.0",port="1",rdma_parent_name="ens1np0",sys_image_guid="6c:20:be:00:03:e1:a2:58"}0
[0126] The following data shows that the host node_name="sh-cube-worker-1" and the network card ifname="mlx5_12" generated 0 packets out of order, indicating that the RDMA network communication is normal and there is no out of order.
[0127] rdma_out_of_sequence_total{ifname="mlx5_12",is_root="false",net_dev_name="net3",node_guid="ec:a7:ad:fe:ff:21:48:d2" ,node_name="sh-cube-worker-1",otel_scope_name="spiderpool-agent",otel_scope_version="1.24.0",owner_api_version="app s / v1",owner_kind="DaemonSet",owner_name="mx-rdma-test-rdma-tools",owner_namespace="rdma",pod_name="mx-rdma-test-rdm a-tools-7ghbv",pod_namespace="rdma",port="1",rdma_parent_name="ens841np0",sys_image_guid="64:19:be:00:03:e1:a2:58"}0
[0128] Container-specific RDMA activity / event detection: Use eBPF technology to detect container RDMA read and write events, then associate the calling process's PID with the container it belongs to, output metrics data, and store it persistently in Elasticsearch. An example is as follows:
[0129] #The following sample data shows that 100 RDMA events (api_name="ibv_post_send") were detected through eBPF. Through Kubernetes container association, it is shown that the event occurred on the network card (rdma_dev="mlx5_2") of the container (container_name="ai_job5") on the node (node="worker4").
[0130] rdmacall_counter{node="worker4",container_name="ai_job5",api_name="ibv_post_send",return_code="100",rdma_dev="mlx5_2",interface="net5"}=100
[0131] Step S12c: The fifth flow data and the sixth data are collectively referred to as second flow data.
[0132] The above steps not only collect host-level metrics but also cover the container level, identifying the specific pod or AI task causing the network bottleneck, providing the foundation for complete monitoring and end-to-end tracking from an AI perspective. Furthermore, using this detection data, administrators can define detailed monitoring and alerting rules, such as alerting for events with excessively high RDMA API call frequencies or for unexpectedly excessive RDMA calls from containers.
[0133] In some embodiments, the network monitoring component further includes: a GUI module, and the method further includes:
[0134] The switch monitoring module obtains the network neighbor topology information of each RDMA switch based on the LLDP protocol, converts it into the data format of the cloud native platform network indicator, records it as the first topology information, and sends the first topology information to the API server of the cloud native platform;
[0135] The cluster monitoring module collects network neighbor information of all network cards on each node of the cloud native platform based on the LLDP protocol, converts it into the data format of cloud native platform network indicators, records it as the second topology information, and sends the second topology information to the API server of the cloud native platform;
[0136] The GUI module reads the first topology information and the second topology information from the API server of the cloud native platform, and draws the interconnection relationship between all network devices based on the first topology information and the second topology information to obtain a network topology map.
[0137] Furthermore, in some embodiments, the first traffic data and the second traffic data include the following five-tuple information: source IP address, destination IP address, protocol type, source port, and destination port; the method also includes: a GUI module, based on the five-tuple information of the first traffic data and the second traffic data, combined with the identity mapping relationship, draws all forwarding paths of the communication flows corresponding to the first traffic data and the second traffic data, and then draws a communication path diagram of any AI task.
[0138] In this embodiment, the GUI module is a key component independently developed, which is mainly responsible for generating a network topology diagram and performing traffic analysis and aggregation of AI task containers to generate a communication path diagram for any AI task, thereby providing users with a browser-accessible interface to facilitate monitoring and analysis of network status.
[0139] It should be noted that LLDP (Link Layer Discovery Protocol) is a standard protocol for network device self-discovery. Through LLDP Data Units (LLDPDUs), network devices can send local device information (such as device ID, interface ID, system name, etc.) to directly connected neighboring devices.
[0140] During the data collection phase, each RDMA switch starts the LLDP service. Then, based on the LLDP protocol, the switch monitoring module obtains the network neighbor topology information of each RDMA switch through gNMI or SSH, including the neighbor information of each port, and converts it into the data format of the cloud native platform network indicator (that is, the metrics format data in Kubernetes), records it as the first topology information, and sends the first topology information to the API server of the cloud native platform.
[0141] For example, the switch monitoring module collects information such as the peer node name, peer port identifier, local device port, and physical connection relationship from each RDMA switch through the LLDP protocol. The sample collection results are as follows:
[0142] #The following data shows that port port = "Ethernet184" of switch switch = "gpu-leaf-switch-4" is connected to the network card peer_port = "ens842v6" of node peer_name = "master-1" in the Kubenetes cluster.
[0143] lldp_switch_neighbor{group="gpu",host="10.193.77.204",peer_name="master-1",peer_port="ens842v6",port="Ethernet184",switch="gpu-leaf-switch-4"}1
[0144] Convert the above LLDP topology information into standard resource objects or monitoring indicators supported by the platform, such as the Prometheus-style indicator format:
[0145] rdma_link_status{switch="gpu-leaf-switch-4",port="ens842v6",peer="master-1",peer_port="Ethernet184"}1
[0146] Enable the LLDP service on each Kubernetes node. The cluster monitoring module detects the network neighbor topology information on the Kubernetes node in real time based on the LLDP protocol to form the second topology information. The example is as follows:
[0147] #The following sample data shows that the network card port = "eth8" of the node node = "worker4" is connected to the port peer_port = "Ethernet200" of the switch peer_name = "switch1".
[0148] lldp_node_neighbor{host="10.1.55.2",peer_name="switch1",peer_port="Ethernet200",port="eth8",node="worker4"}1
[0149] The first and second topology information obtained from the real-time detection are stored in the Kubernetes API server. The GUI module dynamically generates an overall network topology diagram based on the network neighbor topology information (i.e., the first and second topology information) between the switches and hosts (nodes) stored in the Kubernetes API server.
[0150] By aggregating the first and second topology information, a complete network topology diagram is generated, which can show the connection relationship between each switch and host in the cluster for visitors to browse and analyze. Based on this visualized topology relationship, potential problems or bottlenecks in the network can be discovered in a timely manner.
[0151] Furthermore, the first and second traffic data include the following five-tuple information: source IP address, destination IP address, protocol type, source port, and destination port. Because each flow is associated with an identity, the GUI module can draw a communication path diagram for any AI task based on the five-tuple information.
[0152] Specifically, the communication path diagram is presented as a sequence diagram, using the AI workload for each AI training / inference group as the precision. It monitors the network traffic for a group of AI tasks, including switch forwarding paths, traffic timing, traffic volume, and container identities running in the Kubernetes cluster. This helps identify data anomalies and assesses whether each AI task encounters network communication issues. Ultimately, with the help of network graphical visualization, this feature presents this information in a user-friendly manner. Furthermore, on the communication path diagram, each task (including the node / pod where the task resides) is drawn as a sequence diagram object within a rectangular box. Each rectangular box has a lifeline drawn vertically below it, indicating the object's lifetime during the interaction process. Messages (communication packets), the basic unit of interaction between objects, are represented by arrows pointing from the task sending the packet to the task receiving it. Network metrics such as packet forwarding metrics, bandwidth utilization, source / destination IP addresses, source / destination ports, switches traversed, and port link status are annotated next to the arrows. Furthermore, when a task is processing data (such as training or inference), a vertical rectangular box representing the data processing period is drawn on the lifeline for that period.
[0153] Figure 6 This is an example of a communication path graph for a group of AI tasks, used to demonstrate the RDMA monitoring effect at the container level, that is, the visualization effect of the AI task communication path graph. Figure 6As shown, a set of AI tasks consists of four containers. The diagram illustrates how these four containers communicate with each other at specific points in time, the amount of traffic generated by each communication, the switch paths that packets traverse, and more. In the diagram, the AI task (Job on node 1) is drawn as a yellow rectangle as an object in the sequence diagram. As you can understand, each node can run one or more AI tasks, deployed via containers (pods), involving distributed training / inference (e.g., parameter synchronization and data parallelism). Blue arrows indicate the direction of packet transmission, labeled with the port (eno1->eth1), switch (switch leaf 1), and bandwidth (10Gb). Multiple switches can exist, such as switchleaf 1, switchleaf 2, switchleaf 4, etc., connecting nodes and providing high-bandwidth communication. A lifeline is drawn vertically below each task rectangle. The blue rectangle on the lifeline represents the data processing phase of the task, and the amount of data processed is also labeled (e.g., 10Gb). On this basis, we first provide a complete perspective of the AI task, including the number of pods involved, the amount of switch data, whether congestion occurs, the total amount of data written / read, the communication time (Duration), and the number of suspicious tasks / processes (Suspicious Job / Pro). We then summarize each pod separately, counting the amount of data read / written (Write / Read), communication time (Duration), the number of RDMA network cards (NICs), the number of processes (Progresses), etc., and plot each communication time point / time period in a sequence diagram (as shown on the left side of the figure).
[0154] In this embodiment, the GUI module is developed and customized based on visualization frameworks such as Grafana and Kibana, specifically to implement the drawing of network topology diagrams, communication path diagrams, and the visualization of network performance indicators. Among them, Grafana is an open source data visualization platform that allows users to display and analyze data source information from Prometheus through charts and dashboards. In this embodiment, the GUI module realizes the visualization of RDMA indicator data of switches and containers in Prometheus by calling Grafana to customize the visualization panel. Kibana is a visualization platform designed to work with Elasticsearch, allowing users to search, view and interact with data stored in the Elasticsearch index through a graphical interface, and display data in a variety of ways such as charts, tables and maps, thereby realizing advanced data analysis and visualization functions. In this embodiment, the GUI module realizes the visualization of process RDMA activities / events stored in Elasticsearch by calling Kibana to customize the visualization panel.
[0155] On the basis of automatically constructing the RDMA network topology map and the communication path map of the AI task, in some embodiments, the first traffic data and the second traffic data are analyzed and processed to obtain the monitoring results of the RDMA network, including: using the network topology map and the communication path map, combined with the first traffic data and the second traffic data shown in the map, to identify the communication hotspots and abnormal ports in the network, track the entire network link of the AI task, identify illegal traffic and process identity, monitor the congestion of the RDMA network, monitor and locate the health status of all network devices, and the above results are collectively referred to as the monitoring results of the RDMA network.
[0156] (1) Visualization of network basic indicator data (network performance indicators): Use network topology diagrams and communication path diagrams to intuitively display network basic indicators and achieve visualization of network basic indicator data. By monitoring the RDMA indicator data of Kubernetes hosts and containers, monitoring the port indicator data of switches, and using cloud native methods, the network topology diagram effectively displays information on each network link, including real-time RDMA throughput, RDMA congestion mechanism message statistics, link health status, etc., allowing administrators to quickly understand the overall health status of the network and understand cluster network problems.
[0157] Figure 3 An example of realizing visualization of network basic indicator data in a network topology diagram is shown. Figure 3As shown, basic network metrics data is visualized, specifically including visualization of link-level metrics (Link Metrics), node-level metrics (NodeMetrics), and switch-level metrics (Switch Metrics). Link-level metrics primarily display the transmission rate / bandwidth on the link. It should be noted that links here include not only those between spine switches and leaf switches, but also those between leaf switches and nodes (NICs on nodes), and between nodes and storage switches. Node-level metrics include: the number of RDMA network cards (RDMANics) equipped on each node, the number of AI tasks currently using RDMA network communication (Jobs in RDMAactivity), the number of processes currently using RDMA network communication (Progresses inRDMAactivity), and the number of abnormal tasks or processes (Suspicious Jobs / Processes). Switch-level indicators include: the number of ports in normal operation (Up Ports), the number of faulty or closed ports (Down Ports), CPU utilization (Cpu), the write bandwidth of each port (xxx Write bps), and the number of packet drops on each port (xxx Drop Packets). xxx represents a port, such as Eth1 port or Eth2 port. Figure 3 , administrators can understand the interconnection relationship between network devices, understand the packet sending and receiving indicators on each network link, and obtain the following information: calculate the overall network bandwidth utilization, conduct equipment cost analysis and decision-making; dynamically build accurate computing network and storage network topology, and understand topology changes and status. Therefore, the network topology map not only identifies the connection relationship between each switch and the host, but also displays the key indicators on each link. At the same time, for each node and container in the Kubernetes cluster, the system can monitor its RDMA network card-related performance indicators in real time, including packet forwarding indicators, port link status, etc. Based on these indicators, administrators can easily locate communication hotspot nodes in the cluster, optimize traffic distribution within the cluster, improve overall operating efficiency, and facilitate network rationality analysis of container scheduling; observe the RDMA indicators of each physical link; and monitor the indicators of each RDMA network card of the container and host.
[0158] In the above embodiment, the GUI module uses Grafana's visualization capabilities to display various network performance indicators, including the switch's RDMA traffic metrics (such as the write bandwidth of each port), RDMA congestion metrics (such as PFC, ECN, and packet loss), and RDMA network card data for nodes and pods in the Kubernetes cluster. These metrics include, but are not limited to, packet forwarding metrics, port link status, and bandwidth utilization. Using these metrics, administrators can gain a detailed understanding of the network status of each node in the cluster, identify communication hotspots, and assess whether there are problems such as insufficient bandwidth or network overload.
[0159] (2) Full-link network tracking of AI tasks: Draw the network communication structure between AI task containers to form a communication path diagram, intuitively display the communication path and RDMA communication topology of each container, conduct data aggregation analysis on the container group of each AI task, comprehensively observe its communication behavior, identify data anomalies in the network and task communication (through traffic size and anomaly detection), identify the traffic distribution of communication between containers, evaluate whether each group of AI tasks has network communication problems, and quickly find nodes or paths with abnormal traffic. Historical data comparison: By comparing current communication behavior with historical data, analyze whether the current task has network communication anomalies or performance deviations. Finally, with the help of network graphical visualization effects, the information of various traffic flows is presented in a humanized way to help administrators gain a deep understanding of the communication characteristics of AI tasks and identify whether there are security risks.
[0160] Specifically, in this embodiment, the network monitoring component aggregates data for a group of AI task containers, observing their RDMA communication topology and traffic changes, and identifying issues such as abnormal traffic volume. By comparing this data with the historical data for the group of tasks, it analyzes the network communication behavior of the current task and promptly identifies any abnormal traffic patterns or performance bottlenecks. This is crucial for AI workloads that require high throughput and low latency, ensuring efficient utilization of network resources and optimizing task scheduling.
[0161] For example, in Figure 6 In the figure, network congestion occurred when the AI task on node 4 sent data to the AI task on node 1. Specifically, congestion occurred on the communication path from the eth11 port of switch spine1 to the eth11 port of switch leaf 5. The amount of data involved (10Gb) and the time of occurrence (10:30AM to 10:50AM) are shown.
[0162] (3) Identify illegal traffic and process identities: By associating traffic information in the RDMA network with the identity information of resources (such as containers) in the Kubernetes cluster, illegal traffic can be identified and marked. This includes: (a) illegal traffic generated outside the cluster, such as traffic from IP addresses outside the cluster and source addresses outside the trusted network; (b) illegal traffic generated by processes on the node host on the RDMA network, that is, legitimate nodes but illegal processes initiate RDMA operations; (c) illegal traffic generated by illegal containers on the RDMA network, such as containers bypassing security policies to directly access RDMA devices. By monitoring the RDMA read and write events / activities of all processes on the host, suspicious RDMA communications of containers or host applications can be identified and alarms can be issued, and the identification results can be used for correlation analysis of security issues.
[0163] (4) Monitoring the congestion of the RDMA network: The congestion control mechanism of RDMA is designed to optimize data transmission efficiency, reduce packet loss and network jitter, and ensure low latency and high throughput.
[0164] Congestion management in RoCE networks primarily relies on Priority Flow Control (PFC) and Explicit Congestion Notification (ECN). PFC prevents queue overflow by sending pause frames, achieving lossless transmission, but this can lead to global congestion. ECN, on the other hand, alleviates congestion by marking congestion status in packet headers, allowing the receiver to notify the sender to reduce the rate.
[0165] In InfiniBand networks, congestion control uses a dynamic traffic adjustment mechanism based on the IB routing protocol, including rate-based adjustment and tag-based feedback strategies, allowing data flows to bypass congested paths and improve network throughput. In comparison, RoCE relies on Ethernet's congestion control mechanism and is more compatible with traditional networks, while InfiniBand has a more comprehensive flow control solution suitable for extreme high-performance scenarios such as supercomputing and AI training. Therefore, this embodiment monitors PFC and ECN protocol packets in network devices to understand network congestion.
[0166] Figure 4 An example of network congestion is shown in the network topology diagram. Figure 4As shown, for each link (including all links between the spine switch, leaf switch, node RDMA network card, and storage switch), detailed indicators such as the number of priority flow control packets (PFC packets), explicit congestion notification triggers (ECN), link packet drops (Packet Drop), and the number of congestion notification protocol packets (xxx ECN CNP) received by each port can be further monitored. By monitoring the port status of each switch in detail and displaying PFC, ECN, and packet loss statistics that reflect RDMA congestion in the network, critical congested links can be identified, allowing administrators to accurately locate communication hotspot switches and abnormal ports, providing a reliable basis for network optimization and troubleshooting, thereby optimizing network performance and improving AI communication efficiency.
[0167] (5) Monitor and locate the health status of all network devices: Monitor the health status of all network devices, including RDMA switches, optical fibers and links, and RDMA network cards of nodes / Pods. Through relevant indicator data, identify faulty devices and areas, guide container scheduling, locate faulty devices, solve the root cause, and quickly resume training.
[0168] In the above example, the identity of RDMA traffic in Kubernetes hosts / pods and switches is identified based on the identity information of resources (such as AI workloads) in the Kubernetes platform. This allows us to determine whether there is network congestion caused by illegal traffic interference, or whether there is illegal traffic disguised as legitimate identity communication, affecting the data security of the AI system, and implement network security with traceable traffic identity. eBPF technology is used to monitor the RDMA activities / events of all containers in the Kubernetes cluster host. When host programs and containers exhibit RDMA network activity that does not conform to expectations, alarms are issued using the common methods of the Kubernetes platform. Based on the security rules pre-specified by the administrator, the illegal activities of the container can even be blocked, thus achieving RDMA activity monitoring and security for the container.
[0169] In some embodiments, the network monitoring component further includes: an alarm module, which implements an alarm in a cloud native platform manner when the monitoring result of the RDMA network reaches a preset alarm condition.
[0170] Specifically, alarm conditions include: detecting unsafe network traffic (such as illegal traffic from outside the cluster) from monitored RDMA indicator data, and RDMA network traffic exceeding the traffic threshold. When the RDMA network monitoring results indicate that the above alarm conditions are met, the alarm module will implement the alarm in a cloud-native platform manner.
[0171] It should be noted that implementing alerts in a cloud-native platform manner means using the alert components provided by the cloud-native platform (such as Alertmanager and ElastAlert) to issue alerts.
[0172] Among them, Alertmanager is an open source component in the Prometheus monitoring system, which is mainly responsible for receiving, processing and distributing alarm information. In this embodiment, Prometheus will perform security monitoring based on real-time RDMA switch indicator data (first traffic data) and container RDMA indicator data (second traffic data). If unsafe network traffic is found, the Prometheus component will distribute the alarm to Alertmanager. In this embodiment, the alarm module integrates Alertmanager to classify and group the alarms of the Prometheus component, and send them to designated recipients according to predefined routing rules, such as email, Slack or PagerDuty.
[0173] ElastAlert is an open source alerting framework designed for use with Elasticsearch. It is designed to monitor data from Elasticsearch and issue alerts. It does this by periodically querying Elasticsearch and detecting anomalies, spikes, or other patterns of interest based on predefined rules. When the alert conditions are met, ElastAlert triggers the corresponding alert and supports multiple notification methods such as email, Slack, and HTTP POST. In this embodiment, the alert module integrates ElastAlert to alert on illegal RDMA activity events of processes in Elasticsearch. The ElastAlert framework has flexible rule configuration and high availability features, and can automatically restore the state when Elasticsearch is unavailable, thereby ensuring continuous monitoring and timely response to potential problems.
[0174] In addition, the alarm information generated by the alarm module is combined with the network topology map and communication path map drawn by the GUI module to display network security alarms in a visual manner.
[0175] On the one hand, a visualization panel customized based on Grafana displays information on legal and illegal traffic in the RDMA network, including but not limited to potential security risks such as data flow events from outside the cluster and cluster traffic events of non-AI tasks. On the other hand, a visualization panel customized based on Kibana displays legal and illegal RDMA activity events of host processes, including the time of these events, the host where they are located, process details, and the container to which they belong. The network monitoring component identifies the traffic security of the RDMA network and host processes in real time, and displays the above security information in real time through the GUI module. Once abnormal traffic or potential attack behavior is detected, an alarm is immediately triggered and displayed in the interface to help administrators respond in a timely manner.
[0176] As an example, the method provided in this embodiment supports running in a standard 400G RDMAroce network environment, implements network security detection for GPU computing power networks, implements network security detection for storage networks, supports at least 3,000 GPU card clusters, using a 400Gbps RDMA network, and if each switch has 64 ports, can support the management of at least 50 RoCE switches.
[0177] In summary, the solution provided in this embodiment addresses the problem that existing switch vendors' network monitoring is conducted from the perspective of switch hardware and cannot meet the observation requirements of AI containers and Kubernetes. Traditional cloud-native observation vendors focus on TCP / IP networks and do not involve RDMA networks and switches, which cannot meet the needs of RDMA network observation. Therefore, this embodiment proposes an RDMA network monitoring method for cloud-native AI systems, which has the following technical effects:
[0178] 1) Overall correlation and fusion: This system combines observation data from multiple aspects, including the cloud native platform (Kubernetes container platform), AI tasks (a group of pods), and RDMA switches, to form data correlation analysis, enabling deep fusion observation, improving operation and maintenance efficiency, and avoiding monitoring blind spots.
[0179] 2) Ensure AI data communication security: Combine legitimate container IP identities and switch message statistics to identify communication relationships and security in the RDMA network, preventing AI tasks from being injected with malicious data. Monitor RDMA AI calls across all containers and processes on the container platform to optimize performance and conduct root cause analysis.
[0180] 3) Deep network tuning for AI models: Tracking switch traffic forwarding paths, packet metrics, and network congestion at the AI task level to support deep network performance tuning for different models.
[0181] 4) Network hardware cost analysis: By observing the network resource utilization and RDMA switch forwarding path of each AI training session, it supports hardware cost analysis of network devices to optimize resource allocation.
[0182] The method provided in this embodiment can provide cloud-native platforms with complete security monitoring capabilities from the perspective of AI applications, achieving full-stack visual security monitoring for the entire process of artificial intelligence (AI) application operation. This method can fully perceive and securely control key aspects of AI tasks in cloud-native platforms, such as RDMA network communication, computing resources, and data flow, based on the task logic (process) and operating environment (Pod) of AI applications. Unlike traditional network monitoring, which observes physical devices or virtual networks as units, this monitoring solution takes the AI application itself as the core object of monitoring, focusing on the network behavior and bandwidth usage of a specific AI task (such as model training and inference) in the entire system. It achieves full-link observability of AI tasks, fine-grained security identification, the ability to integrate multi-source data, and ensures the data and communication security of AI tasks. By building a comprehensive, fine-grained, and traceable security observation system starting from AI containers (AI tasks, AI workloads) and running through the Kubernetes platform and physical network devices, it effectively ensures the security and controllability of AI tasks during network communication and data interaction.
[0183] Based on the same inventive concept, this embodiment provides an RDMA network monitoring system in a cloud-native artificial intelligence system. The system includes a network monitoring component that is containerized and deployed on each node of the cloud-native platform. The network monitoring component includes a switch monitoring module and a cluster monitoring module. The system includes:
[0184] The collection unit is configured as a switch monitoring module to collect traffic data of each RDMA switch to obtain first traffic data; the cluster monitoring module collects traffic data of the RDMA device on each node of the cloud native platform to obtain second traffic data;
[0185] A mapping unit is configured as a network monitoring component to obtain resource identity information of the cloud native platform, and based on the resource identity information, identifies the flow identities of the first flow data and the second flow data, so as to establish an identity mapping relationship between the first flow data, the second flow data and the AI workload;
[0186] The analysis unit is configured to convert the first flow data and the second flow data into the data format of the cloud native platform network indicator based on the identity mapping relationship, so as to analyze and process the first flow data and the second flow data to obtain the monitoring results of the RDMA network.
[0187] The following combination Figure 7The conceptual structure of the RDMA network monitoring system in the cloud native artificial intelligence system provided in this embodiment is described in detail. Figure 7 As shown in the figure, the conceptual design of the RDMA network monitoring system in a cloud-native AI system consists of three layers: infrastructure, core features, and operations. The infrastructure layer primarily includes hardware devices (such as RDMA switches and InfiniBand switches) and software platforms (such as Kubernetes). The RDMASwitch for Storage is a switch dedicated to storage communication (such as connecting to a Ceph cluster); the RDMASwitch for GPU is a switch dedicated to inter-GPU node communication. In terms of switch protocol types, both RoCE and InfiniBand switches are supported. The core features include: drawing an RDMA network topology diagram and displaying network performance metrics and congestion monitoring; full-link AI task tracking, namely tracking the full-link performance of distributed training tasks and identifying bottlenecks; identifying and alerting illegal traffic and processes; monitoring network device health, assisting in pod network optimization, and monitoring RDMA activity and events. The operations layer includes a GUI module and an alarm module, providing visualization of alarms, logs, and network performance metrics.
[0188] It should be noted that the RDMA network monitoring system in a cloud-native artificial intelligence system provided in this embodiment can implement the steps and processes of the RDMA network monitoring method in a cloud-native artificial intelligence system provided in any of the above embodiments and achieve the same technical effects, so they will not be described one by one here.
[0189] This embodiment further provides a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described in any one of the above embodiments.
[0190] This embodiment provides a computer-readable storage medium having a computer program / instruction stored thereon, wherein the computer program / instruction, when executed by a processor, implements the steps of the method described in any of the above embodiments.
[0191] The foregoing description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are readily apparent to those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A RDMA network monitoring method in a cloud native artificial intelligence system, characterized in that: The method is executed by a network monitoring component, which is containerized and deployed on each node of the cloud native platform. The network monitoring component includes: a switch monitoring module and a cluster monitoring module. The method includes: The switch monitoring module collects traffic data of each RDMA switch to obtain first traffic data; the cluster monitoring module collects RDMA traffic data on each node of the cloud native platform to obtain second traffic data; The network monitoring component obtains resource identity information of the cloud native platform, and identifies the flow identities of the first flow data and the second flow data based on the resource identity information, so as to establish an identity mapping relationship between the first flow data, the second flow data and each resource of the cloud native platform; Based on the identity mapping relationship, the first traffic data and the second traffic data are converted into the data format of the cloud native platform network indicator, so as to analyze and process the first traffic data and the second traffic data to obtain the monitoring result of the RDMA network.
2. The method according to claim 1, characterized in that The switch monitoring module collects traffic data of each RDMA switch to obtain first traffic data, including: The switch monitoring module is configured with the IP address of each RDMA switch, and the switch monitoring module is connected to all RDMA switches based on the IP address; The switch monitoring module collects the flow data of each RDMA switch itself based on the sFlow protocol to obtain third flow data; and / or, The switch monitoring module collects port flow data of each RDMA switch based on the gNMI protocol to obtain fourth flow data; The third flow data and the fourth flow data are collectively referred to as first flow data.
3. The method according to claim 1, characterized in that The cluster monitoring module collects traffic data of RDMA devices on each node of the cloud native platform to obtain second traffic data, including: The cluster monitoring module obtains RDMA traffic data from the host and container network cards on each node through periodic polling, which is recorded as fifth traffic data; and / or, The cluster monitoring module uses eBPF technology to monitor RDMA read and write events on each host and obtains corresponding process information, which is recorded as the sixth data; The fifth flow rate data and the sixth data are collectively referred to as second flow rate data.
4. The method according to claim 1, wherein The network monitoring component further includes: a GUI module, and the method further includes: The switch monitoring module obtains network neighbor topology information of each RDMA switch based on the LLDP protocol, converts the information into a data format of a cloud native platform network indicator, records the information as first topology information, and sends the first topology information to the API server of the cloud native platform; The cluster monitoring module collects network neighbor information of all network cards on each node of the cloud native platform based on the LLDP protocol, converts the information into a data format of cloud native platform network indicators, records the information as second topology information, and sends the second topology information to the API server of the cloud native platform; The GUI module reads the first topology information and the second topology information from the API server of the cloud native platform, and draws the interconnection relationship between all network devices based on the first topology information and the second topology information to obtain a network topology map.
5. The method according to claim 4, characterized in that The first flow data and the second flow data include the following five-tuple information: source IP address, destination IP address, protocol type, source port, and destination port; the method further includes: The GUI module draws all forwarding paths of the communication flows corresponding to the first traffic data and the second traffic data based on the five-tuple information of the first traffic data and the second traffic data, combined with the identity mapping relationship, and then draws a communication path diagram of any AI task.
6. The method according to claim 5, characterized in that Analyzing and processing the first flow data and the second flow data to obtain monitoring results of the RDMA network includes: Using the network topology diagram and the communication path diagram, combined with the first traffic data and the second traffic data shown in the diagram, communication hotspots and abnormal ports in the network are identified, the entire network link of the AI task is tracked, illegal traffic and process identities are identified, the congestion of the RDMA network is monitored, and the health status of all network devices is monitored and located. The above results are collectively referred to as the monitoring results of the RDMA network.
7. The method according to claim 6, characterized in that The network monitoring component further includes: an alarm module, which implements an alarm in a cloud native platform manner when the monitoring result of the RDMA network reaches a preset alarm condition.
8. An RDMA network monitoring system in a cloud native artificial intelligence system, characterized in that: The system includes a network monitoring component, which is containerized and deployed on each node of the cloud native platform. The network monitoring component includes a switch monitoring module and a cluster monitoring module. The system includes: The collection unit is configured to collect the flow data of each RDMA switch by the switch monitoring module to obtain first flow data; the cluster monitoring module collects the flow data of the RDMA device on each node of the cloud native platform to obtain second flow data; A mapping unit is configured to obtain resource identity information of the cloud native platform from the network monitoring component, and identify the flow identities of the first flow data and the second flow data based on the resource identity information to establish an identity mapping relationship between the first flow data, the second flow data and the AI workload; The analysis unit is configured to convert the first traffic data and the second traffic data into the data format of the cloud native platform network indicator based on the identity mapping relationship, so as to analyze and process the first traffic data and the second traffic data to obtain the monitoring results of the RDMA network.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instruction is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Distributed system monitoring framework construction method and device, equipment and storage medium
CN121209982A
Fault positioning method, system and device across IB and RoCE networks
CN121690994A