Cluster monitoring method and related device for large language model training
Through the independently developed TPU cluster monitoring system, hardware information is obtained and the monitoring panel is updated in real time, which supports custom configuration and unified analysis. This solves the scalability and real-time issues of the TPU cluster monitoring system and improves monitoring efficiency and system stability.
Patent Information
- Application Number
- CN202411828164.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2044-12-12
AI Technical Summary
In existing technologies, the monitoring system of TPU clusters cannot meet the high computing density requirements. The data update and processing speed is slow, the interface is unfriendly, the expansion capability is insufficient, the alarm function is not timely, and the data from different tools is scattered and difficult to integrate and analyze uniformly, resulting in low monitoring efficiency.
Through self-developed methods, it obtains hardware information of the TPU cluster, including temperature, power consumption, and memory usage, updates the monitoring panel in real time, supports multiple data sources and custom configurations, provides a unified analysis platform, configures alarm rules for timely notification of problems, and predicts operating status through machine learning to improve the flexibility and scalability of the monitoring system.
It enables real-time monitoring of TPU clusters, improves the efficiency of data collection and display, quickly locates problems, reduces system construction costs, improves resource utilization and decision-making accuracy, reduces downtime, and enhances system stability and reliability.
Smart Images

Figure CN119271505B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing, and in particular to a cluster monitoring method and related devices for large language model training. Background Art
[0002] With the rapid development of artificial intelligence (AI), large language models (LLMs) have demonstrated remarkable performance in natural language processing. Training LLMs requires extensive computing resources and data, typically relying on high-performance computing clusters. To meet these high computing demands, computing centers can be built. These centers are composed of application-specific integrated circuits (ASICs) (Tensor Processing Units (TPUs)) designed to accelerate machine learning tasks, forming a massive computing cluster to support LLM training. TPUs are designed to efficiently process tensor (multidimensional array) operations, a core component of deep learning and neural network computing. TPUs are primarily used to accelerate machine learning and deep learning tasks, particularly those requiring extensive matrix or tensor operations, such as image recognition and natural language processing.
[0003] The performance and stability of the computing cluster are crucial during LLM training. Because training requires long, high-load runs, any issue in any step can lead to training interruption or inaccurate results. Therefore, real-time monitoring of the TPU computing center cluster has become a key technical challenge in large-scale model training. Summary of the Invention
[0004] In response to the technical problems existing in the prior art, this application provides a cluster monitoring method and related devices for large language model training, which are used to achieve real-time monitoring of TPU clusters, improve cluster monitoring efficiency, and the scalability and flexibility of the cluster monitoring system, and reduce the cost of setting up the cluster monitoring system.
[0005] In a first aspect, an embodiment of the present application provides a cluster monitoring method for large language model training, the method comprising:
[0006] Obtaining raw operational data of a target object; the target object is one or more TPU nodes in a TPU cluster of a dedicated integrated circuit; the TPU nodes are used to accelerate training and inference tasks of a large language model;
[0007] Convert the raw operating data into hardware information to be processed of the target object; the conversion method of the hardware information to be processed is configured based on multiple indicator dimensions to be monitored in the TPU cluster; the multiple indicator dimensions include at least one of the following: temperature, real-time power consumption, and memory usage of the target object in the TPU cluster;
[0008] Real-time display information is generated based on the hardware information to be processed, and is updated in real time to the monitoring panel for display; the real-time display information includes at least: hardware monitoring information and computing status information corresponding to the TPU cluster.
[0009] In a second aspect, an embodiment of the present application provides a cluster monitoring device for large language model training, the device comprising at least the following units:
[0010] An acquisition unit is configured to acquire raw operating data of a target object; the target object is one or more TPU nodes in a TPU cluster of a dedicated integrated circuit; the TPU nodes are used to accelerate training tasks and inference tasks of a large language model;
[0011] A conversion unit is configured to convert the raw operating data into hardware information to be processed of a target object; the conversion method of the hardware information to be processed is configured based on multiple indicator dimensions to be monitored in the TPU cluster; the multiple indicator dimensions include at least one of the following: temperature, real-time power consumption, and memory usage of the target object in the TPU cluster;
[0012] The display unit is configured to generate real-time display information based on the hardware information to be processed, and update it in real time to the monitoring panel for display; the real-time display information at least includes: hardware monitoring information and computing status information corresponding to the TPU cluster.
[0013] In a third aspect, an embodiment of the present application provides an electronic device, comprising:
[0014] at least one processor, memory, and input-output unit;
[0015] The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the cluster monitoring method for large language model training of the first aspect.
[0016] In a fourth aspect, a computer-readable storage medium is provided, comprising instructions, which, when executed on a computer, cause the computer to execute the cluster monitoring method for large language model training according to the first aspect.
[0017] The beneficial effect of the present application is that it provides a cluster monitoring method and related devices for large language model training. In this technical solution, first, the original operation data of the target object is obtained. The target object is one or more TPU nodes in a dedicated integrated circuit TPU cluster; the TPU nodes are used to accelerate the training and reasoning tasks of the large language model. Then, the original operation data is converted into the hardware information to be processed of the target object. The conversion method of the hardware information to be processed is obtained based on the configuration of multiple indicator dimensions to be monitored in the TPU cluster; the multiple indicator dimensions include at least one of the following: temperature, real-time power consumption, and memory usage of the target object in the TPU cluster. Finally, real-time display information is generated based on the hardware information to be processed, and is updated in real time to the monitoring panel for display. The real-time display information includes at least: hardware monitoring information and computing status information corresponding to the TPU cluster. The technical solution of this application realizes real-time monitoring of the TPU cluster, improves cluster monitoring efficiency, as well as the scalability and flexibility of the cluster monitoring system, and reduces the cost of setting up the cluster monitoring system. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 This is a flow chart of a cluster monitoring method for large language model training according to an embodiment of the present application;
[0019] Figure 2 This is a schematic diagram of the structure of a cluster monitoring device for large language model training according to an embodiment of the present application;
[0020] Figure 3 This is a schematic structural diagram of an electronic device according to an embodiment of the present application;
[0021] Figure 4 It is a structural diagram of a medium device according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0023] With the rapid development of artificial intelligence (AI), large language models (LLMs) have demonstrated remarkable performance in natural language processing. Training LLMs requires extensive computing resources and data, typically relying on high-performance computing clusters. To meet this high computing demand, many companies and research institutions have established computing centers. These centers are composed of application-specific integrated circuits (ASICs) (Tensor Processing Units (TPUs)) designed to accelerate machine learning tasks, forming large computing clusters to support LLM training. The primary purpose of TPUs is to accelerate machine learning and deep learning tasks, particularly those requiring extensive matrix or tensor operations, such as image recognition and natural language processing. TPUs are designed to efficiently process tensor (multidimensional array) operations, a core component of deep learning and neural network computing. TPUs are primarily used to accelerate machine learning and deep learning tasks, particularly those requiring extensive matrix or tensor operations, such as image recognition and natural language processing.
[0024] During LLM training, the performance and stability of the computing cluster are crucial. Because the training process requires long, high-load runs, any problem in any link can lead to training task interruption or inaccurate results.
[0025] Therefore, how to monitor the TPU computing center cluster in real time has become one of the technical problems that need to be solved during the training of large models.
[0026] In order to solve at least one technical problem in the related art, an embodiment of the present application provides a cluster monitoring method and related devices for large language model training. In this technical solution, first, the original operation data of the target object is obtained. The target object is one or more TPU nodes in a dedicated integrated circuit TPU cluster; the TPU nodes are used to accelerate the training tasks and inference tasks of the large language model. Then, the original operation data is converted into the hardware information to be processed of the target object. The conversion method of the hardware information to be processed is obtained based on the configuration of multiple indicator dimensions to be monitored in the TPU cluster; the multiple indicator dimensions include at least one of the following: temperature, real-time power consumption, and memory usage of the target object in the TPU cluster. Finally, real-time display information is generated based on the hardware information to be processed, and updated to the monitoring panel in real time for display. The real-time display information includes at least: hardware monitoring information and computing status information corresponding to the TPU cluster.
[0027] The following analyzes the improvements of the technical solution of this application compared with the related art from multiple perspectives:
[0028] In response to the problem that monitoring systems in related technologies often rely on general hardware information reading interfaces, making it difficult to meet the high computing density requirements of TPU clusters, the technical solution provided by this application uses independent research and development methods to obtain TPU hardware information, including the number of TPUs, memory block usage, temperature, power consumption, etc., to ensure data accuracy and real-time performance. This application can more finely monitor TPU clusters, improve the efficiency and accuracy of data collection, and provide a reliable data foundation for subsequent real-time monitoring and alarms.
[0029] To address the problem of slow data update and processing speeds in monitoring systems in related technologies, making real-time monitoring difficult, the technical solution provided by this application ensures that hardware information of the TPU cluster can be displayed in a timely manner through real-time data update capabilities, helping to quickly discover and address potential problems. This makes the cluster monitoring solution more real-time, able to quickly respond to changes in hardware status, and enhance the system's real-time monitoring capabilities.
[0030] To address the problems of unfriendly interfaces, non-intuitive data display, and difficulty in quickly locating problems in related technologies, the technical solution provided by this application uses a custom-configured monitoring panel with a reasonable arrangement of monitoring charts and clear data display, making it easy to quickly locate problems. This helps improve the user experience, lowers the threshold for system use, and enables non-professionals to quickly understand and locate problems.
[0031] Some monitoring systems in related technologies lack the ability to scale and adapt when clusters expand or new technologies are introduced. The technical solution provided by this application supports multiple data sources and allows users to customize monitoring panels and charts, enabling the system to quickly adapt and expand when clusters expand or new technologies are introduced. The technical solution provided by this application offers greater flexibility and scalability, adapting to different monitoring needs and environmental changes, and reducing the cost and time of system upgrades.
[0032] The alarm function of the monitoring system in related technologies is often not timely enough and cannot effectively notify relevant personnel to handle problems. The technical solution provided by this application can detect abnormal situations in real time through the configuration of alarm rules and notify relevant personnel through email, SMS, etc., ensuring that problems are handled in a timely manner. This helps to improve the system's response speed and problem-solving efficiency, reduce downtime caused by hardware failures or anomalies, and improve system stability and reliability.
[0033] In related technologies, the data generated by different monitoring tools and systems is scattered and difficult to integrate and analyze uniformly. The technical solution provided by this application integrates data from different monitoring tools and systems to provide a unified analysis platform, solving the problems of data silos and integration difficulties, improving resource utilization and data analysis efficiency, thereby providing more comprehensive data analysis capabilities, helping users better understand and manage TPU clusters, and improving the accuracy and efficiency of decision-making.
[0034] In summary, the technical solution provided by this application significantly improves the resource utilization of TPU clusters during large language model training and inference through independently developed data collection methods, real-time data processing and storage, a flexible and intuitive monitoring interface, powerful scalability and flexible configuration, real-time alerting capabilities, and a data silo solution. This allows for timely detection and resolution of potential issues, reduces downtime, and improves system stability and reliability. Compared to related technologies, this application provides a more real-time, more aesthetically pleasing, and more scalable monitoring solution that better meets the monitoring needs of large-scale TPU clusters.
[0035] The technical solution of the present application and the cluster monitoring solution for large language model training provided in the embodiments of the present application can also be executed by an electronic device, which can be a server, a server cluster, or a cloud server. The electronic device can also be a terminal device such as a mobile phone, a computer, a tablet computer, a wearable device, or a dedicated device (such as a dedicated terminal device with a cluster monitoring method system for large language model training). These electronic devices can also be equipped with the chips introduced in the above embodiments. Alternatively, these electronic devices can also be installed with a service program for executing the cluster monitoring solution for large language model training.
[0036] Figure 1 A flow chart of a cluster monitoring method for large language model training provided in an embodiment of the present application is shown as follows: Figure 1 As shown, the method includes the following steps:
[0037] 101, obtaining the original operation data of the target object;
[0038] 102, converting the original operation data into hardware information to be processed of a target object;
[0039] 103 : Generate real-time display information according to the hardware information to be processed, and update the real-time display information to the monitoring panel for display.
[0040] Among them, TPU is a custom hardware chip designed to accelerate machine learning (especially deep learning) tasks. The main purpose of TPU is to provide higher computing efficiency and lower power consumption than traditional CPUs and GPUs, thereby significantly improving performance when training and reasoning large-scale neural network models. TPU is deeply optimized for matrix operations, which are the core operations of deep learning algorithms such as convolutional neural networks and recurrent neural networks. TPU is designed to quickly process large amounts of data and reduce data transmission delays within the chip, thereby increasing overall computing speed. TPU performs well in deep learning training and reasoning tasks, can significantly improve computing efficiency and reduce power consumption, and is widely used in cloud services and edge computing scenarios. By providing efficient hardware acceleration, TPU has promoted the development and application of deep learning technology.
[0041] In an embodiment of the present application, the target object is one or more Tensor Processing Unit (TPU) nodes in a TPU cluster. Further optionally, the TPU nodes are used to accelerate the training and inference tasks of large language models.
[0042] A TPU cluster is a high-performance computing cluster composed of multiple TPUs, specifically designed to accelerate machine learning and deep learning tasks. These clusters excel at processing large datasets, training complex neural network models, and performing efficient parallel computations. By connecting multiple TPU nodes together, a TPU cluster can simultaneously process multiple computing tasks, significantly improving computational throughput and efficiency. Each TPU node is equipped with High Bandwidth Memory (HBM), ensuring rapid data transfer during processing and reducing latency. The TPU has an instruction set optimized for deep learning, enabling more efficient execution of common neural network operations.
[0043] The TPU cluster adopts a modular design, which allows the number of nodes to be easily expanded or reduced as needed to adapt to different computing requirements. Through the cloud computing platform, users can dynamically adjust the usage of TPU resources as needed to achieve elastic scaling. The TPU cluster connects each node through a high-speed interconnection network (such as NVLink or a customized high-speed network) to ensure efficient data transmission between nodes. By optimizing the communication protocol and architecture, the TPU cluster can achieve low-latency data exchange and is suitable for parallel computing tasks that require frequent communication. The TPU cluster supports multiple deep learning frameworks, such as TensorFlow and PyTorch, and users can easily migrate existing models to the TPU for acceleration. Through automated tools and APIs, users can manage and optimize the use of TPU clusters and simplify the deployment and maintenance process.
[0044] In practical applications, TPU clusters are suitable for training large-scale deep learning models, such as image classification, speech recognition, and natural language processing, significantly reducing training time. Leveraging distributed training technology, TPU clusters can train multiple models simultaneously, further improving training efficiency. In production environments, TPU clusters efficiently perform model inference tasks, providing low latency and high throughput. They are suitable for applications requiring real-time inference, such as speech recognition and image processing. TPU clusters excel in scientific computing and data analysis, and are suitable for large-scale data processing and complex computational tasks. Leveraging their powerful computing power, TPU clusters enable efficient data mining and analysis, uncovering underlying patterns within the data.
[0045] Compared to traditional CPUs and GPUs, TPUs offer higher computational efficiency for deep learning tasks, significantly accelerating training and inference. Through specially optimized hardware design, TPU clusters maintain high computing performance while reducing energy consumption. Users can dynamically adjust TPU resource usage based on actual needs, paying on demand and minimizing resource waste. With the support of cloud service providers, users can easily deploy and manage TPU clusters, reducing operational costs and complexity. TPU clusters support mainstream deep learning frameworks, allowing users to seamlessly migrate existing models and benefit from high-performance acceleration. A rich set of tools and libraries are provided to help users optimize model and cluster configurations, improving overall performance.
[0046] TPU clusters, with their high-performance hardware design, flexible scalability, efficient communication architecture, and robust software support, provide powerful computing capabilities for deep learning and scientific computing. Whether in large-scale model training, real-time inference, or high-performance computing, TPU clusters demonstrate significant advantages and potential, driving the development and application of related technologies. With the support of cloud service providers, users can easily access and use TPU clusters, improving their computing power and efficiency.
[0047] As an optional embodiment, assume that each TPU node in the TPU cluster is deployed with a corresponding register, and the register of each TPU node is used to store the original operating data of the corresponding node. Based on the above assumption, obtaining the original operating data of the target object in 101 can be implemented by reading the original operating data of the target object from the register corresponding to the target object through a preconfigured hardware data interface.
[0048] Wherein, the hardware data interface at least includes: i 2c and / or PCIE. In a TPU cluster, each TPU node is equipped with corresponding registers to store the raw operating data of the corresponding node. These registers can be accessed and read through pre-configured hardware data interfaces such as I²C and PCIe. I²C (Inter-Integrated Circuit) is a serial communication interface used to transmit data between devices. PCIe (Peripheral Component Interconnect Express) is a high-speed serial computer expansion bus standard used to connect peripherals and host systems.
[0049] The registers on each TPU node are used to store the node's raw operating data, including computing status, memory usage, temperature, and power consumption. By reading the data in the registers, the operating status of the TPU node can be monitored in real time, helping with performance analysis and fault diagnosis.
[0050] In an alternative example, let's take reading temperature data from a TPU node via I²C. TPU node A is connected to the host system via the I²C bus. Assume that the I²C address of TPU node A is 0x5A. The host system initializes the I²C controller, setting the I²C bus speed and communication mode. The host system sends a read command to address 0x5A, specifying the address of the temperature register to read (assuming it's 0x01). TPU node A responds to the read request and sends the temperature data (assuming it's 45 degrees Celsius) back to the host system. The host system stores the received data in a buffer and further processes or displays the temperature information.
[0051] In another alternative example, let's take reading the memory usage data of a TPU node via PCIe. TPU node B is connected to the host system via the PCIe bus. Assume that TPU node B's PCIe device ID is 0x1234. Based on this, the host system initializes the PCIe controller, identifies and connects to TPU node B with device ID 0x1234. The host system sends a read request to TPU node B, specifying the address of the memory usage register to read (assuming it is 0x100) and the data length (assuming it is 4 bytes). In response to the read request, TPU node B sends the memory usage data (assuming it is 512MB) back to the host system. The host system stores the received data in a cache and further processes or displays the memory usage information.
[0052] In summary, through pre-configured hardware data interfaces (such as I²C and PCIe), raw operational data can be read from the registers of each TPU node in the TPU cluster. This design not only improves the efficiency and accuracy of data collection, but also enhances the system's real-time monitoring and performance analysis capabilities.
[0053] In 102, the original operation data is converted into hardware information to be processed of the target object.
[0054] In an embodiment of the present application, the conversion method of the hardware information to be processed is obtained based on multiple indicator dimensions to be monitored in the TPU cluster. The multiple indicator dimensions include at least one of the following: temperature, real-time power consumption, and memory usage of the target object in the TPU cluster.
[0055] In a TPU cluster, raw operational data (such as temperature, real-time power consumption, and memory usage) must be converted into hardware information for the target object. This hardware information is configured based on multiple metrics to support subsequent analysis and decision-making. The following describes the specific conversion method and an example.
[0056] Temperature, such as that of a TPU node, indicates the current heat level of the chip. Monitoring temperature can help prevent performance degradation or hardware damage caused by overheating. During operation, temperature sensors monitor the temperature readings of TPU nodes in real time and send the measured values to a central monitoring system via a dedicated data transmission channel. The central monitoring system analyzes the temperature data of multiple TPU nodes using pre-designed data processing algorithms. If the temperature of a node is detected to exceed a preset safety threshold, the system will immediately initiate a predetermined response mechanism, such as reducing the task load of the node, activating the heat dissipation system to enhance cooling, or in extreme cases, automatically shutting down the node to protect the hardware from overheating and damage.
[0057] Real-time power consumption, such as that of a TPU node, indicates the chip's current energy consumption. Monitoring power consumption can help optimize energy use, reduce costs, and prevent damage to hardware caused by excessive power consumption. Real-time power consumption monitoring is a key component of TPU cluster management, aiming to monitor the energy consumption of each TPU node.
[0058] High-performance power consumption sensors capture real-time power consumption data from each node and transmit it to a central monitoring system via I²C or PCIe interfaces. Optionally, pre-defined analysis algorithms process this data to calculate key metrics such as average, maximum, and minimum power consumption, and map these metrics to different energy consumption levels (e.g., low, medium, and high). When a node's power consumption exceeds a preset threshold, the system immediately takes optimization measures, such as adjusting the distribution of computing tasks, reducing the load on high-power nodes, or optimizing the execution strategy of computing tasks to reduce overall energy consumption. This not only helps reduce energy costs but also effectively prevents hardware overheating and damage caused by excessive power consumption, thereby ensuring the efficient and stable operation of the TPU cluster. Through real-time power consumption monitoring and optimization, the TPU cluster can achieve efficient energy utilization while maintaining high performance, extending hardware lifespan, and reducing operational costs.
[0059] Memory usage, such as that of a TPU node, indicates the current utilization of memory resources. Monitoring memory usage can help identify memory bottlenecks and optimize resource allocation for computing tasks. Memory usage monitoring is an important part of TPU cluster management, designed to monitor the memory resource utilization of each node in real time. Through the memory monitoring module, the system can obtain the current memory usage data of each TPU node and transmit it to the central monitoring system through the network interface. The monitoring system uses a built-in algorithm to calculate the average, peak, and minimum memory usage values and map this data to predefined memory usage levels (such as low, medium, and high).
[0060] When it detects that the memory usage of a node is approaching or exceeding the preset threshold, the system automatically triggers optimization measures. For example, it adjusts the task allocation strategy, transfers some computing tasks to nodes with more abundant memory resources, or optimizes data processing algorithms to reduce memory usage. This real-time monitoring and optimization mechanism helps identify and resolve memory bottlenecks, ensures the efficient execution of computing tasks, and maximizes the utilization of the TPU cluster's computing resources. By accurately monitoring and dynamically adjusting memory usage, the TPU cluster can maintain stable performance under high load conditions, improve resource utilization, and provide reliable support for complex computing tasks.
[0061] In TPU cluster management, raw operational data (such as temperature, real-time power consumption, and memory usage) is first converted into standardized numerical values for subsequent processing and analysis. For example, temperature data is converted to Celsius or Kelvin, power consumption data is converted to watts (W), and memory usage data is converted to percentages (%) or bytes (Bytes). Next, this raw data is aggregated into an overall metric by calculating the average, maximum, and minimum values of specific metrics across multiple TPU nodes to represent the status of the entire TPU cluster. Finally, this standardized and aggregated data is mapped to predefined states or levels to facilitate system decision-making. For example, temperature data is mapped to "normal," "warning," and "dangerous" levels; power consumption data is mapped to "low," "medium," and "high" levels; and memory usage data is mapped to "low," "medium," and "high" levels. This provides strong support for optimizing energy use, preventing hardware damage, and identifying memory bottlenecks.
[0062] In a TPU cluster, by converting raw runtime data into the target object's hardware information to be processed, multiple metrics (such as temperature, real-time power consumption, and memory usage) can be effectively monitored and managed. This conversion method not only improves the usability and understandability of the data, but also provides a solid foundation for subsequent performance analysis and decision-making.
[0063] For example, the original operating data is: TPU node A -50°C, node B -55°C, and node C -60°C. The original data unit (°C) is retained. The average temperature is 55°C, the maximum temperature is 60°C, and the minimum temperature is 50°C. The average temperature of 55°C is mapped to the "Warning" level (assuming the threshold is 50°C normal and 60°C warning). The maximum temperature of 60°C is mapped to the "Warning" level, and the minimum temperature of 50°C is mapped to the "Normal" level.
[0064] For example, the original operating data shows TPU node A at 150W, node B at 160W, and node C at 170W. The original data unit (W) is retained. The average power consumption is 160W, the maximum power consumption is 170W, and the minimum power consumption is 150W. The average power consumption of 160W is mapped to the "medium" level (assuming the thresholds are 150W low, 160W medium, and 170W high). The maximum power consumption of 170W is mapped to the "high" level, and the minimum power consumption of 150W is mapped to the "low" level.
[0065] For example, the original running data shows TPU node A at 80%, node B at 85%, and node C at 90%. The original data units (%) are retained. Average memory usage is 85%, maximum memory usage is 90%, and minimum memory usage is 80%. Average memory usage of 85% is mapped to the "Medium" level (assuming the thresholds are 80% for Low, 85% for Medium, and 90% for High). Maximum memory usage of 90% is mapped to the "High" level, and minimum memory usage of 80% is mapped to the "Low" level.
[0066] Further optionally, in step 102, converting the raw operating data into the target object's hardware information to be processed can be implemented by: cleaning the read raw operating data; performing data conversion calculations on the raw operating data according to the data calculation methods corresponding to the respective data types to obtain corresponding data endpoints; and constructing the hardware information to be processed using the data endpoints corresponding to the raw operating data. In this embodiment of the present application, the data endpoints are the key indicator data to be extracted from the raw operating data.
[0067] In TPU cluster management, raw operating data needs to be cleaned, converted, and constructed to be converted into the target object's hardware information to be processed. Assume that the raw operating data read from TPU nodes A, B, and C is in TPU cluster management. The raw operating data needs to be cleaned, converted, and constructed to be converted into the target object's hardware information to be processed.
[0068] Assume that the raw operating data read from TPU nodes A, B, and C is as follows: TPU node A has a temperature of 50°C, power consumption of 150W, and memory utilization of 80%; TPU node B has a temperature of 55°C, power consumption of 160W, and memory utilization of 85%; and TPU node C has a temperature of 60°C, power consumption of 170W, and memory utilization of 90%. The raw operating data is cleaned to remove abnormal data points and noise. For example, the data is checked to ensure it is within a reasonable range and obviously unreasonable data (such as data with negative temperature values or zero power consumption) is removed. The raw operating data is converted and calculated according to the data calculation methods corresponding to the respective data types to obtain the corresponding data endpoints: the temperature indicator remains in the original data units (°C), with the data endpoints being 50°C, 55°C, and 60°C; the power consumption indicator remains in the original data units (W), with the data endpoints being 150W, 160W, and 170W; and the memory utilization indicator remains in the original data units (%), with the data endpoints being 80%, 85%, and 90%.
[0069] The hardware information to be processed is constructed using the data endpoints corresponding to the original running data: In the temperature data endpoint, the average temperature is 55°C, the maximum temperature is 60°C, and the minimum temperature is 50°C. The temperature status is mapped to an average temperature of 55°C warning (assuming the threshold is 50°C normal and 60°C warning), a maximum temperature of 60°C warning, and a minimum temperature of 50°C normal. In the power consumption data endpoint, the average power consumption is 160W, the maximum power consumption is 170W, and the minimum power consumption is 150W. The power consumption status is mapped to an average power consumption of 160W medium (assuming the thresholds are 150W low, 160W medium, and 170W high), a maximum power consumption of 170W high, and a minimum power consumption of 150W low. In the memory usage data endpoint, the average memory usage is 85%, the maximum memory usage is 90%, and the minimum memory usage is 80%. The memory usage status is mapped to an average memory usage of 85% medium (assuming the thresholds are 80% low, 85% medium, and 90% high), a maximum memory usage of 90% high, and a minimum memory usage of 80% low. The final hardware information to be processed includes the following: For temperature information, node A is at a normal temperature of 50°C, node B is at a warning temperature of 55°C, and node C is at a warning temperature of 60°C. The average temperature is at a warning temperature of 55°C, the maximum temperature is at a warning temperature of 60°C, and the minimum temperature is at a normal temperature of 50°C. For power consumption information, node A is at a low of 150W, node B is at a medium of 160W, and node C is at a high of 170W. The average power consumption is at a medium of 160W, the maximum power consumption is at a high of 170W, and the minimum power consumption is at a low of 150W. For memory usage information, node A is at a low of 80%, node B is at a medium of 85%, and node C is at a high of 90%. The average memory usage is at a medium of 85%, the maximum memory usage is at a high of 90%, and the minimum memory usage is at a low of 80%.
[0070] Through the above data conversion and construction process, the hardware information to be processed, including temperature, power consumption, and memory usage, is obtained. This information provides basic data support for subsequent system decision-making and management.
[0071] As an optional embodiment, in step 102, after converting the raw operating data into the target object's hardware information to be processed, the hardware information to be processed may also be converted into a corresponding multidimensional hardware information sequence in a multidimensional data model and stored in a storage space corresponding to the multidimensional tag. The multidimensional hardware information sequence is associated with multiple indicator dimension tags.
[0072] In TPU cluster management, raw operational data is first converted into hardware information to be processed, such as temperature, power consumption, and memory usage. This hardware information is then converted into a multidimensional hardware information sequence within a multidimensional data model and stored in the storage space corresponding to the multidimensional labels. Specifically, dimension labels (such as node labels, metric labels, statistical labels, and status labels) are determined, and the data points in each dimension are constructed into a multidimensional data sequence. For example, for the temperature metric, the constructed multidimensional hardware information sequence includes the temperature data for nodes A, B, and C, as well as the average, maximum, and minimum temperature data, and is associated with a status label (such as normal or warning). These multidimensional hardware information sequences are then stored in the corresponding storage space. For example, the temperature information sequence is stored in "temperature_storage," the power consumption information sequence is stored in "power_storage," and the memory usage information sequence is stored in "memory_usage_storage." In this way, the hardware information to be processed is converted into a multidimensional hardware information sequence and stored in the storage space corresponding to the multidimensional labels, supporting subsequent query, analysis, and decision-making.
[0073] As an optional embodiment, in 102, after converting the hardware information to be processed into a corresponding multidimensional hardware information sequence in the multidimensional data model, it is also possible to detect in real time whether the multidimensional hardware information sequence contains target hardware information that meets the alarm conditions; the detection indicators of the multidimensional hardware information sequence are composed of alarm thresholds and / or alarm conditions of different types of hardware information; if the multidimensional hardware information sequence contains target hardware information that meets the alarm conditions, alarm information of the target object is generated based on the target hardware information; the alarm information is displayed in real time on the monitoring panel to prompt the user to promptly deal with the operating risks existing in the TPU cluster.
[0074] For example, after converting the hardware information to be processed into a multidimensional hardware information sequence in a multidimensional data model, the system can detect in real time whether the multidimensional hardware information sequence contains target hardware information that meets the alarm conditions. Assume that the multidimensional hardware information sequence includes temperature, power consumption, and memory usage information. For example, the temperatures of nodes A, B, and C are 50°C, 55°C, and 60°C, respectively; the power consumption is 150 W, 160 W, and 170 W, respectively; and the memory usage is 80%, 85%, and 90%, respectively. The alarm thresholds are set as follows: temperature warning of 55°C and above, severe warning of 60°C and above; power consumption warning of 160 W and above, severe warning of 170 W and above; memory usage warning of 85% and above, severe warning of 90% and above. The system detects that the temperature, power consumption, and memory usage of nodes B and C meet warning or critical warning conditions, generates alarm information, and displays it in real time on the monitoring panel. For example: Node B: Temperature 55°C, Warning; Node C: Temperature 60°C, Critical Warning; Node B: Power Consumption 160W, Warning; Node C: Power Consumption 170W, Critical Warning; Node B: Memory Usage 85%, Warning; Node C: Memory Usage 90%, Critical Warning. These alarms allow users to promptly identify and address operational risks in the TPU cluster.
[0075] Further, optionally, in the above steps, before real-time detection of whether the multi-dimensional hardware information sequence contains target hardware information that meets the alarm conditions, alarm thresholds and / or alarm conditions for different types of hardware information can be configured. Specifically, monitoring targets for collecting raw operating data are set in a pre-set configuration file; alarm rules are defined in a configuration management file; the configuration management file is applied to the configuration file, and the alarm rules are loaded using the configuration file.
[0076] In an embodiment of the present application, the alarm rules include at least: alarm thresholds and / or alarm conditions for different types of hardware information. Further optionally, the alarm thresholds and / or alarm conditions for different types of hardware information include at least one of the following: memory usage exceeding a preset memory capacity upper limit, temperature below a preset temperature lower limit, temperature above a preset temperature upper limit, or node load exceeding a preset load value.
[0077] Before checking in real time whether a multi-dimensional hardware information sequence contains target hardware information that meets alarm conditions, configuring alarm thresholds and / or alarm conditions for different types of hardware information is a crucial step. First, in a pre-configured configuration file, specify the monitoring targets for collecting raw operational data, such as the temperature, power consumption, and memory usage of nodes A, B, and C. Next, define alarm rules in the configuration management file, including alarm thresholds and / or alarm conditions for different types of hardware information, such as temperature warnings of 55°C and above and critical warnings of 60°C and above, power consumption warnings of 160W and above and critical warnings of 170W and above, and memory usage warnings of 85% and above and critical warnings of 90% and above. Furthermore, alarm rules can include specific conditions such as memory usage exceeding a preset memory capacity limit (e.g., 80%), temperature falling below a preset lower temperature limit (e.g., 0°C), temperature exceeding a preset upper temperature limit (e.g., 70°C), and node load exceeding a preset load value (e.g., 85%). Applying the configuration management file and loading the alarm rules in the configuration file ensures that the system can recognize and handle these alarm conditions.
[0078] For example, a sample configuration file may contain detailed configurations of monitoring targets and alarm rules, ensuring that when target hardware information that meets the conditions is detected, alarm information is generated and displayed, prompting users to promptly address operational risks.
[0079] To ensure Prometheus can correctly collect relevant metrics and trigger alerts, first ensure that the monitoring targets for memory usage, temperature, and power consumption are correctly configured in the prometheus.yml configuration file. Next, define alert rules in a YAML file, such as alerts.yml, and include it in prometheus.yml. Alert rules should include trigger conditions for memory usage exceeding 64GB, temperature below 5°C or above 90°C, and power consumption exceeding 500W. In prometheus.yml, add or modify the rule_files section to load alerts.yml. Finally, verify that the rules are activated by simulating relevant conditions (such as high memory usage, abnormal temperature, or high power consumption) on the Rules page of the Prometheus web UI. Verify that alerts are generated correctly. These steps ensure that Prometheus can promptly detect and address potential operational risks.
[0080] Further optionally, in 102, before converting the original operating data into the target object's hardware information to be processed, the application scenario of the large language model can also be obtained; based on the real-time environmental data and / or scenario attributes in the application scenario, the operating status of the large language model is predicted to obtain a real-time operating status prediction value of the large language model; adaptive configuration is performed according to the real-time operating status prediction value to obtain multiple indicator dimensions for TPU cluster matching.
[0081] For example, when using Prometheus and Grafana for hardware information monitoring, combining the application scenarios of large language models can further improve the accuracy and adaptability of monitoring and alerting. When the large language model is running, its current application scenario is first obtained. The application scenario may include different computing tasks, load types, real-time environmental factors (such as network latency, temperature, humidity, etc.), and scenario attributes (such as task priority, data sensitivity, etc.). Based on the real-time environmental data and scenario attributes obtained in the application scenario, machine learning and predictive models are used to predict the operation of the large language model. These predictions may include but are not limited to memory usage trends, CPU load predictions, temperature change trends, etc.
[0082] Based on the predicted values of the operating status, dynamically adjust Prometheus's monitoring indicators and Grafana's dashboard configuration to match the current operating status. For example, if it is predicted that memory usage will increase significantly, the memory usage alarm threshold can be adjusted in advance. If the temperature is predicted to rise significantly, the sensitivity of the temperature alarm can be increased or the cooling strategy can be adjusted. Prometheus continues to be responsible for collecting and storing hardware information such as memory usage, temperature, and power consumption. At the same time, Prometheus can also collect data related to the application scenario and input it into the predictive model. Create dynamic dashboards in Grafana to display customized monitoring data based on the current application scenario and predicted values. For example, when the memory usage is predicted to exceed 64GB, the dashboard can display a highlighted prompt. Use Grafana's alarm function to set alarm rules based on predicted values so that relevant personnel can be notified in advance when potential problems are predicted.
[0083] Prometheus and Alertmanager can continue to be responsible for rule triggering and notification, while Grafana's alerting function can serve as a supplement to provide more intuitive alert information. In combination with predictive models, alerts can be configured to be adaptive, adjusting the sensitivity and threshold of alerts based on the predicted values. Suppose a scenario in which a large language model is used for a high-priority natural language processing task. First, obtain real-time network latency and priority information for the task. Based on this information, predict the model's operation status, especially memory usage and CPU load. If it is predicted that memory usage will increase rapidly, Prometheus can adjust the memory usage alert rules in advance, and Grafana can display the relevant memory usage trend chart and set a highlighted prompt. If it is predicted that the CPU load will increase significantly, Grafana can display a time series chart of CPU usage and issue an alert in advance.
[0084] In this way, by combining real-time data and prediction models of application scenarios, more accurate and adaptive monitoring and alarm configuration can be achieved, improving the stability and reliability of large language models during runtime.
[0085] In 103, real-time display information is generated based on the hardware information to be processed, and is updated in real time to the monitoring panel for display. In the embodiment of the present application, the real-time display information at least includes: hardware monitoring information and computing status information corresponding to the TPU cluster.
[0086] In this embodiment, the real-time display information includes hardware monitoring information and computing status information of the TPU cluster. Hardware monitoring information covers key performance indicators such as memory usage, CPU load, network traffic, temperature, and power consumption. This information is collected and stored by Prometheus and displayed in a visual form in Grafana.
[0087] For example, memory usage is updated in real time in the form of a line graph through the "node_memory_MemTotal_bytes - node_memory_MemFree_bytes" metric, in GB, showing the current memory usage and historical trends; CPU load is updated in real time in the form of a line graph through the "rate(node_cpu_seconds_total{mode=
[0088] The "system"}[5m])" metric is displayed in the form of a bar graph in seconds, showing the CPU load in system mode in real time. The network traffic is displayed in the form of a line graph through the "rate(node_network_receive_bytes_total[5m])" metric in MB / s, showing the network receive traffic in real time. The temperature is displayed in the form of a single value through the "node_hwmon_temp_celsius" metric in °C, showing the current temperature in real time and providing color warnings (for example, red indicates too high and green indicates normal). The power consumption is displayed in the form of a line graph through the "node_power_usage_watts" metric in watts, showing the power consumption and historical trends in real time. Computation status information includes task execution status, data processing speed, and error rate.
[0089] For example, the task execution status is updated in real time through the "task_status{state="running"}" indicator in a single-value display format, showing the number of tasks currently running. The data processing speed is displayed in the form of a line graph through the "rate(data_processing_bytes_total[5m])" indicator in MB / s, showing the data processing speed in real time. The error rate is displayed in the form of a line graph through the "rate(error_count_total[5m]) / rate(total_count_total[5m])" indicator, showing the error rate in real time.
[0090] In Grafana, you can create a dedicated dashboard that displays task execution status and node temperature at the top, memory usage trends and data processing speed on the left, CPU load and network traffic on the right, and power consumption and error rates at the bottom. By following these steps and examples, you can create a dynamic, real-time monitoring dashboard in Grafana to help operations teams monitor and manage the operational status of the TPU cluster in real time.
[0091] As an optional embodiment, in step 103, generating real-time display information based on the hardware information to be processed and updating it in real time to the monitoring panel for display can be implemented as follows:
[0092] According to different types of hardware information in the hardware information to be processed, the corresponding data display method is dynamically configured; based on the configured data display method, the information arrangement method in the monitoring panel is constructed; based on the hardware information to be processed, the hardware monitoring information and computing status information corresponding to the TPU cluster are determined; and the determined hardware monitoring information and computing status information are respectively loaded into the constructed monitoring panel.
[0093] In an optional embodiment, the corresponding data display method is dynamically configured based on different types of hardware information. For example, information that changes over time, such as memory usage and network traffic, is better displayed using a line chart, while instantaneous values such as CPU load and task execution status are better displayed using a bar chart or single value display.
[0094] Example configuration:
[0095] Memory Usage: Line Chart
[0096] CPU Load: Histogram
[0097] Network Traffic: Line Chart
[0098] Temperature: Single value display
[0099] Power consumption: line graph
[0100] Task execution status: single value display
[0101] Data processing speed: line chart
[0102] Error Rate: Line Chart
[0103] Based on the configured data display method, construct the information layout in the monitoring panel. Ensure that key information is easy to view and understand. For example, important metrics (such as task execution status and temperature) can be placed at the top, time series data (such as memory usage and data processing speed) can be placed on the left and right, and power consumption and error rate can be placed at the bottom.
[0104] Top: Task execution status (single value display), TPU node temperature (single value display).
[0105] Left: Memory usage trend (line chart), data processing speed (line chart).
[0106] Right: CPU load (bar chart), network traffic (line chart).
[0107] Bottom: Power consumption (line graph), error rate (line graph).
[0108] Based on the hardware information to be processed, the hardware monitoring information and computing status information corresponding to the TPU cluster are determined. This information is collected through Prometheus and stored in a time series database.
[0109] Taking hardware monitoring information as an example, memory usage: `node_memory_MemTotal_bytes - node_memory_MemFree_bytes`; CPU load: `rate(node_cpu_seconds_total{mode=
[0110] "system"}[5m])`; Network traffic: `rate(node_network_receive_bytes_total[5m])`; Temperature: `node_hwmon_temp_celsius`; Power consumption: `node_power_usage_watts`.
[0111] The calculation status information is as follows: task execution status: `task_status{state="running"}`; data processing speed: `rate(data_processing_bytes_total[5m])`; error rate: `rate(error_count_total[5m]) / rate(total_count_total[5m])`.
[0112] Load the determined hardware monitoring information and computing status information into the built monitoring dashboards. Ensure that this information is updated in real time and displayed on the dashboards. Configure Prometheus as a data source for Grafana to ensure that Grafana can query and retrieve data from Prometheus. Create new line charts, bar charts, and single value display panels in Grafana. Configure Prometheus queries for each panel, using the metrics previously determined. For example, for the memory usage trend panel, configure the query to be `node_memory_MemTotal_bytes - node_memory_MemFree_bytes`, in GB, displaying time series data. Arrange the created panels on the dashboard according to the previously designed layout. Adjust the style of each panel, such as color, font size, and chart title, to ensure that the information is clear and easy to read. Ensure that the data in each panel automatically refreshes and updates in real time. Regularly review and adjust the panel configuration to adapt to changes in hardware information and updated metrics.
[0113] By following the steps and examples above, you can dynamically configure data display in Grafana, build the information layout in the monitoring panel, determine the hardware monitoring information and computing status of the TPU cluster, and load this information into the monitoring panel for real-time display. This method helps the operations team monitor and manage the operating status of the TPU cluster in real time, ensuring efficient and stable system operation.
[0114] In an embodiment of the present application, first, the original operating data of the target object is obtained. The target object is one or more TPU nodes in a dedicated integrated circuit TPU cluster; the TPU nodes are used to accelerate the training and reasoning tasks of large language models. Then, the original operating data is converted into hardware information to be processed for the target object. The conversion method of the hardware information to be processed is obtained based on the configuration of multiple indicator dimensions to be monitored in the TPU cluster; the multiple indicator dimensions include at least one of the following: temperature, real-time power consumption, and memory usage of the target object in the TPU cluster. Finally, real-time display information is generated based on the hardware information to be processed, and is updated in real time to the monitoring panel for display. The real-time display information includes at least: hardware monitoring information and computing status information corresponding to the TPU cluster. The technical solution of the present application realizes real-time monitoring of the TPU cluster, improves cluster monitoring efficiency, as well as the scalability and flexibility of the cluster monitoring system, and reduces the cost of setting up the cluster monitoring system.
[0115] The technical solution of this application has significant innovation and practicality in monitoring and managing TPU clusters. Through the following key technical points, real-time monitoring and efficient management of TPU clusters are achieved:
[0116] First, this application uses a proprietary method to obtain TPU hardware information. This method not only ensures data accuracy and completeness, but also allows for customized acquisition of specific hardware information based on specific needs, providing a solid foundation for subsequent data processing and presentation. Second, this application pushes processed hardware information to a Prometheus-based data processing application for storage in real time. This real-time push mechanism ensures timely data updates and traceability. Leveraging Prometheus's powerful data processing capabilities, it can efficiently process massive amounts of hardware information, providing data support for subsequent analysis and presentation. Third, this application uses Grafana to configure a beautiful and elegant monitoring panel that displays the hardware information and computing status of the TPU cluster in real time. Through this carefully designed monitoring panel, users can intuitively view key hardware metrics and computing status, promptly identify and resolve issues, and thus improve system stability and reliability. This real-time, visual monitoring method is particularly important in the context of large-scale TPU clusters. In addition, this application also configures Prometheus alerting rules to enable real-time alerting and processing. This real-time alarm mechanism can promptly notify operation and maintenance personnel of possible abnormal situations in the system, so that they can take quick measures to avoid the expansion of problems and system downtime, further improving the stability and reliability of the system.
[0117] In the technical solution of the present application, firstly, the resource utilization rate of the TPU cluster in the training and reasoning process of large language models is significantly improved. Through real-time monitoring and precise data analysis, resources can be better allocated and scheduled, and resource utilization efficiency can be improved. Secondly, potential problems can be discovered and resolved in a timely manner, downtime can be reduced, and the continuous and stable operation of the system can be ensured. Finally, the present application provides a more real-time, more beautiful, and more scalable monitoring solution, which can better meet the monitoring needs of large-scale TPU clusters and has obvious advantages compared with the existing technology. In summary, the technical solution of the present application significantly improves the monitoring efficiency and system stability of the TPU cluster through independently developed hardware information acquisition methods, real-time data push and processing, beautiful monitoring panel design, and real-time alarm processing, providing strong support for the management and operation of large-scale TPU clusters.
[0118] In another embodiment of the present application, a cluster monitoring device for large language model training is also provided. Figure 2 Said device comprises the following units:
[0119] An acquisition unit is configured to acquire raw operating data of a target object; the target object is one or more TPU nodes in a TPU cluster of a dedicated integrated circuit; the TPU nodes are used to accelerate training tasks and inference tasks of a large language model;
[0120] A conversion unit is configured to convert the raw operating data into hardware information to be processed of a target object; the conversion method of the hardware information to be processed is configured based on multiple indicator dimensions to be monitored in the TPU cluster; the multiple indicator dimensions include at least one of the following: temperature, real-time power consumption, and memory usage of the target object in the TPU cluster;
[0121] The display unit is configured to generate real-time display information based on the hardware information to be processed, and update it in real time to the monitoring panel for display; the real-time display information at least includes: hardware monitoring information and computing status information corresponding to the TPU cluster.
[0122] Further optionally, each TPU node in the TPU cluster is deployed with a corresponding register, and the register of each TPU node is used to store the original running data of the corresponding node; the acquisition unit acquires the original running data of the target object and is configured to:
[0123] The original operation data of the target object is read from the register corresponding to the target object through a pre-configured hardware data interface; the hardware data interface at least includes: i2c and / or PCIE.
[0124] Further optionally, the conversion unit converts the original operation data into hardware information to be processed of the target object, and is configured to:
[0125] Cleaning the read raw operating data;
[0126] According to the data calculation method corresponding to each data type, the original operation data is converted and calculated to obtain the corresponding data endpoint; wherein the data endpoint is the key indicator data to be extracted from the original operation data;
[0127] The hardware information to be processed is constructed using the data endpoint corresponding to the original operating data.
[0128] Further optionally, a storage unit is further included, configured to: after the conversion unit converts the original operation data into the hardware information to be processed of the target object, convert the hardware information to be processed into a corresponding multidimensional hardware information sequence in the multidimensional data model, and store the sequence in the storage space corresponding to the multidimensional tag;
[0129] The multi-dimensional hardware information sequence is associated with multiple indicator dimension tags.
[0130] Further optionally, the system further includes an alarm unit configured to: after the conversion unit converts the hardware information to be processed into a corresponding multidimensional hardware information sequence in the multidimensional data model, detect in real time whether the multidimensional hardware information sequence contains target hardware information that meets the alarm condition; the detection index of the multidimensional hardware information sequence is composed of alarm thresholds and / or alarm conditions for different types of hardware information;
[0131] If the multi-dimensional hardware information sequence includes target hardware information that meets the alarm condition, generating alarm information of the target object based on the target hardware information;
[0132] The alarm information is displayed in real time on the monitoring panel to prompt the user to promptly handle the operational risks existing in the TPU cluster.
[0133] Further optionally, a configuration unit is further included, which is configured to:
[0134] Before the detection unit detects in real time whether the multi-dimensional hardware information sequence contains target hardware information that meets the alarm condition, a monitoring target for collecting raw operation data is set in a pre-set configuration file;
[0135] Alarm rules are defined in the configuration management file; the alarm rules include at least: alarm thresholds and / or alarm conditions for different types of hardware information, the alarm thresholds and / or alarm conditions for different types of hardware information include at least one of the following: memory usage exceeds a preset memory capacity upper limit, temperature is lower than a preset temperature lower limit, temperature is higher than a preset temperature upper limit, and node load exceeds a preset load value;
[0136] The configuration management file is applied to the configuration file, and the alarm rules are loaded using the configuration file.
[0137] Further optionally, the method further includes a prediction unit configured to: before the conversion unit converts the original operation data into the to-be-processed hardware information of the target object, obtain the application scenario of the large language model;
[0138] Based on real-time environmental data and / or scene attributes in the application scenario, the operation status of the large language model is predicted to obtain a real-time operation status prediction value of the large language model;
[0139] Adaptive configuration is performed according to the real-time running status prediction value to obtain multiple indicator dimensions matching the TPU cluster.
[0140] Further optionally, the display unit generates real-time display information according to the hardware information to be processed, and updates the real-time display information to the monitoring panel for display, and is configured to:
[0141] Dynamically configure corresponding data display modes according to different types of hardware information in the hardware information to be processed;
[0142] Based on the configured data display method, construct the information arrangement method in the monitoring panel;
[0143] Determine hardware monitoring information and computing status information corresponding to the TPU cluster based on the hardware information to be processed;
[0144] The determined hardware monitoring information and computing status information are loaded into the monitoring panel that has been constructed.
[0145] The system can implement various steps in the above method embodiments, which will not be expanded here.
[0146] In the embodiment of the present application, a cluster monitoring device for large language model training is used to achieve real-time monitoring of the TPU cluster, improve the cluster monitoring efficiency, the scalability and flexibility of the cluster monitoring system, and reduce the cost of setting up the cluster monitoring system.
[0147] See also Figure 3 , Figure 3 This is a schematic diagram of an embodiment of an electronic device provided in an embodiment of the present application. Figure 3As shown, an embodiment of the present application provides an electronic device 500, including a memory 510, a processor 520 and a computer program 511 stored in the memory 510 and executable on the processor 520. When the processor 520 executes the computer program 511, the following steps are implemented: obtaining the original operating data of the target object; the target object is one or more TPU nodes in a dedicated integrated circuit TPU cluster; the TPU nodes are used to accelerate the training tasks and inference tasks of large language models; converting the original operating data into hardware information to be processed of the target object; the conversion method of the hardware information to be processed is obtained based on the configuration of multiple indicator dimensions to be monitored in the TPU cluster; the multiple indicator dimensions include at least one of the following: temperature, real-time power consumption, and memory usage of the target object in the TPU cluster; generating real-time display information based on the hardware information to be processed, and updating it to the monitoring panel for display in real time; the real-time display information includes at least: hardware monitoring information and computing status information corresponding to the TPU cluster.
[0148] See also Figure 4 , Figure 4 This is a schematic diagram of an embodiment of a computer-readable storage medium provided in an embodiment of the present application. Figure 4 As shown, this embodiment provides a computer-readable storage medium 600, on which a computer program 611 is stored. When the computer program 611 is executed by a processor, it implements the following steps: obtaining the original operating data of the target object; the target object is one or more TPU nodes in a dedicated integrated circuit TPU cluster; the TPU nodes are used to accelerate the training tasks and inference tasks of large language models; converting the original operating data into hardware information to be processed of the target object; the conversion method of the hardware information to be processed is obtained based on the configuration of multiple indicator dimensions to be monitored in the TPU cluster; the multiple indicator dimensions include at least one of the following: temperature, real-time power consumption, and memory usage of the target object in the TPU cluster; generating real-time display information based on the hardware information to be processed, and updating it to the monitoring panel in real time for display; the real-time display information includes at least: hardware monitoring information and computing status information corresponding to the TPU cluster.
[0149] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0150] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0151] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0152] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if such changes and modifications of the present application fall within the scope of the claims of the present application and their equivalents, the present application is intended to include such changes and modifications.
Claims
1. A cluster monitoring method for large language model training, characterized in that: The method comprises: Obtain the original running data of the target object; the target object is one or more TPU nodes in a dedicated integrated circuit TPU cluster; the TPU node is used to accelerate the training and reasoning tasks of large language models; TPU is a customized hardware chip designed to accelerate machine learning tasks, used to accelerate the processing of large-scale data to reduce the transmission delay of data within the chip; each TPU node in the TPU cluster is deployed with a corresponding register, and the register of each TPU node is used to store the original running data of the corresponding node; the original running data of the target object is read from the register corresponding to the target object through a pre-configured hardware data interface; the hardware data interface at least includes: i 2 c and / or PCIE; Obtain the application scenario of the large language model; based on the real-time environmental data and scenario attributes in the application scenario, predict the operation status of the large language model to obtain the real-time operation status prediction value of the large language model; adaptively configure according to the real-time operation status prediction value to obtain multiple indicator dimensions matched by the TPU cluster; the real-time environmental data includes at least temperature, humidity, and network delay, and the scenario attributes include at least task priority and data sensitivity; convert the original operation data into the hardware information to be processed of the target object; the conversion method of the hardware information to be processed is obtained based on the configuration of multiple indicator dimensions to be monitored in the TPU cluster; the multiple indicator dimensions include at least one of the following: temperature, real-time power consumption, and memory usage of the target object in the TPU cluster; Converting the hardware information to be processed into a corresponding multidimensional hardware information sequence in a multidimensional data model and storing the sequence in a storage space corresponding to the multidimensional tag; wherein the multidimensional hardware information sequence is associated with a plurality of indicator dimension tags; Real-time detection of whether the multi-dimensional hardware information sequence contains target hardware information that meets the alarm conditions; the detection indicators of the multi-dimensional hardware information sequence are composed of alarm thresholds and / or alarm conditions for different types of hardware information; if the multi-dimensional hardware information sequence contains target hardware information that meets the alarm conditions, generating alarm information for the target object based on the target hardware information; and displaying the alarm information in real time on the monitoring panel to prompt the user to promptly address operational risks in the TPU cluster; Generate real-time display information based on the hardware information to be processed, and update it in real time to the monitoring panel for display; the real-time display information includes at least: hardware monitoring information and computing status information corresponding to the TPU cluster; Before the real-time detection of whether the multi-dimensional hardware information sequence contains target hardware information that meets the alarm conditions, the step of configuring alarm thresholds and / or alarm conditions for different types of hardware information also includes: setting a monitoring target for collecting original operation data in a pre-set configuration file; defining alarm rules in a configuration management file; the alarm rules include at least: alarm thresholds and / or alarm conditions for different types of hardware information, and the alarm thresholds and / or alarm conditions for different types of hardware information include at least one of the following: memory usage exceeds a preset memory capacity upper limit, temperature is lower than a preset temperature lower limit, temperature is higher than a preset temperature upper limit, and node load is greater than a preset load value; applying the configuration management file in the configuration file, and using the configuration file to load the alarm rules.
2. The cluster monitoring method for large language model training according to claim 1, characterized in that: The converting the original operation data into hardware information to be processed of the target object includes: Cleaning the read raw operating data; According to the data calculation method corresponding to each data type, the original operation data is converted and calculated to obtain the corresponding data endpoint; wherein the data endpoint is the key indicator data to be extracted from the original operation data; The hardware information to be processed is constructed using the data endpoint corresponding to the original operating data.
3. The cluster monitoring method for large language model training according to claim 1, characterized in that: Generating real-time display information according to the hardware information to be processed and updating it in real time to the monitoring panel for display includes: Dynamically configure corresponding data display modes according to different types of hardware information in the hardware information to be processed; Based on the configured data display method, construct the information arrangement method in the monitoring panel; Determine hardware monitoring information and computing status information corresponding to the TPU cluster based on the hardware information to be processed; The determined hardware monitoring information and computing status information are loaded into the monitoring panel that has been constructed.
4. A cluster monitoring device for large language model training, characterized in that: The device comprises the following units, wherein: An acquisition unit is configured to acquire the original running data of a target object; the target object is one or more TPU nodes in a dedicated integrated circuit TPU cluster; the TPU nodes are used to accelerate the training and reasoning tasks of large language models; TPU is a customized hardware chip designed to accelerate machine learning tasks, used to accelerate the processing of large-scale data to reduce the transmission delay of data within the chip; each TPU node in the TPU cluster is deployed with a corresponding register, and the register of each TPU node is used to store the original running data of the corresponding node; the original running data of the target object is read from the register corresponding to the target object through a pre-configured hardware data interface; the hardware data interface at least includes: i 2 c and / or PCIE; A prediction unit is configured to obtain an application scenario in which the large language model is located; predict the operation status of the large language model based on real-time environmental data and scenario attributes in the application scenario to obtain a real-time operating status prediction value of the large language model; and adaptively configure the system according to the real-time operating status prediction value to obtain multiple indicator dimensions for TPU cluster matching; the real-time environmental data includes at least temperature, humidity, and network latency, and the scenario attributes include at least task priority and data sensitivity; A conversion unit is configured to convert the raw operating data into hardware information to be processed of a target object; the conversion method of the hardware information to be processed is configured based on multiple indicator dimensions to be monitored in the TPU cluster; the multiple indicator dimensions include at least one of the following: temperature, real-time power consumption, and memory usage of the target object in the TPU cluster; A storage unit is configured to convert the hardware information to be processed into a corresponding multidimensional hardware information sequence in a multidimensional data model and store the converted information in a storage space corresponding to the multidimensional tag; wherein the multidimensional hardware information sequence is associated with a plurality of indicator dimension tags; An alarm unit is configured to detect in real time whether the multidimensional hardware information sequence contains target hardware information that meets the alarm conditions; the detection indicators of the multidimensional hardware information sequence are composed of alarm thresholds and / or alarm conditions for different types of hardware information; if the multidimensional hardware information sequence contains target hardware information that meets the alarm conditions, generate alarm information for the target object based on the target hardware information; and display the alarm information in real time on the monitoring panel to prompt the user to promptly address operational risks in the TPU cluster; A display unit is configured to generate real-time display information based on the hardware information to be processed, and update the real-time display information to the monitoring panel for display; the real-time display information includes at least: hardware monitoring information and computing status information corresponding to the TPU cluster; The configuration unit, when configuring alarm thresholds and / or alarm conditions for different types of hardware information, is configured to: before detecting in real time whether the multidimensional hardware information sequence contains target hardware information that meets the alarm conditions, set a monitoring target for collecting original operation data in a pre-set configuration file; define alarm rules in a configuration management file; the alarm rules include at least: alarm thresholds and / or alarm conditions for different types of hardware information, and the alarm thresholds and / or alarm conditions for different types of hardware information include at least one of the following: memory usage exceeds a preset memory capacity upper limit, temperature is lower than a preset temperature lower limit, temperature is higher than a preset temperature upper limit, and node load is greater than a preset load value; apply the configuration management file in the configuration file, and use the configuration file to load the alarm rules.
5. An electronic device, characterized in that: include: Memory for storing computer software programs; A processor is used to read and execute the computer software program, thereby implementing the cluster monitoring method for large language model training as described in any one of claims 1-3.
6. A computer-readable storage medium, characterized in that The storage medium stores a computer software program, which, when executed by a processor, implements the cluster monitoring method for large language model training according to any one of claims 1 to 3.
7. A chip, characterized in that: The chip is loaded with a computer software program and / or a hardware unit, and the computer software program and / or the hardware unit are used to implement the cluster monitoring method for large language model training as described in any one of claims 1-3.
Citation Information
Patent Citations
GPU computing power resource scheduling method and device
CN118885273A
Hardware-based predictive fault detection and analysis
US20220253337A1