Supercomputing center performance data return method and system and computing equipment
By dividing the computing nodes of the supercomputing center into multiple subclusters and deploying data acquisition and aggregation components, the technical difficulties of performance data acquisition, backhaul and storage of super-large-scale information technology innovation computing centers are solved, and fast, efficient and complete data processing and storage are achieved, reducing the risk of data leakage and network bandwidth pressure.
Patent Information
- Application Number
- CN202510147619.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
AI Technical Summary
The ultra-large-scale information and innovation computing center has technical difficulties in performance data acquisition, backhaul and storage, including diverse data sources, huge scale, high real-time requirements, network bandwidth limitations, and data security and privacy issues.
By dividing the compute nodes of the supercomputing center into multiple subclusters, each subcluster containing multiple compute nodes and a management node, the data acquisition component and data aggregation component are deployed. The data acquisition component regularly collects performance data and converts it to a predetermined format. The data aggregation component performs aggregation and compression processing, sends it to the data receiving component through the management network, and finally decompresses and converts the data into the target format through the data processing component for storage.
It realizes fast, efficient and complete collection, back-passing and storage of supercomputing center performance data, reduces data transmission and storage pressure, saves storage space, reduces the risk of data leakage, and reduces the pressure on managing network bandwidth.
Smart Images

Figure CN120066894A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method for transmitting performance data of a supercomputing center, a system for transmitting performance data of a supercomputing center, and a computing device. Background Art
[0002] A super-large-scale information technology innovation computing center refers to a large infrastructure that integrates a large number of domestic information technology innovation computing resources, network resources, and storage resources. It usually has thousands or even hundreds of thousands of computing nodes, and through technologies such as parallel computing and distributed computing, it can achieve a computing speed of hundreds of millions of times per second or even higher. There are certain technical difficulties in collecting the micro-architecture data and performance data of a super-large-scale information technology innovation computing center, which are mainly reflected in the following aspects: 1) The data sources are diverse and the data scale is huge. For example, in a supercomputing center with 200,000 nodes, each node generates 15 KB of performance data per second, so the data volume per second is as high as 2.86 GB; 2) The real-time requirement is high. In order to timely discover performance problems and optimize system performance, it is necessary to collect performance data in real time or on time, and the requirements for the data collection speed and transmission delay are very high; 3) Network bandwidth limitation. In order not to affect the use of the high-speed network for data collection, it is required to transmit the collected data through the management network. The bandwidth of the management network is usually 1 Gbps, and the transmission of a large amount of performance data through the management network will bring unprecedented bandwidth pressure to the management network; 4) The performance data may contain sensitive information, such as user-customized performance index data, confidential hardware information of users, etc., and there are data security and privacy problems.
[0003] However, the existing data collection frameworks often cannot meet the technical requirements for data collection of supercomputing centers, and there are the following problems: 1) They cannot provide a distributed collection architecture, which is not flexible enough for super-large-scale computing cluster scenarios; 2) The amount of collected data is large and not compressed, and the bandwidth requirement for the management network during data transmission is very high, resulting in a large data transmission pressure; 3) The server side cannot receive an extremely large amount of data, and often the program memory overflows due to the large amount of data; 4) It will occupy a large amount of storage resources, resulting in a large data storage pressure; 5) Using a common data format, there is a risk of data leakage.
[0004] How to quickly, efficiently, and completely collect, transmit back, and store the performance data of a supercomputing center, so as to provide solid data support for the user side to analyze the performance data and optimize the hardware of the supercomputing center, is an urgent problem to be solved.
[0005] Therefore, a method for transmitting performance data of a supercomputing center is needed to solve the problems existing in the above technical solutions. Summary of the Invention
[0006] To this end, the present invention provides a method for transmitting performance data of a supercomputing center and a system for transmitting performance data of a supercomputing center to solve or at least alleviate the problems described above.
[0007] According to one aspect of the present invention, there is provided a method for transmitting performance data of a supercomputing center. Each computing node of the supercomputing center is adapted to be divided into a plurality of sub-clusters. Each sub-cluster respectively includes a plurality of the computing nodes and a management node communicatively connected to the plurality of computing nodes. Data acquisition components are respectively deployed on each of the computing nodes in each sub-cluster, and a data aggregation component is deployed on the management node in each sub-cluster. The method includes: for each sub-cluster, through each data acquisition component on each of the computing nodes in the sub-cluster, periodically acquire performance data of each of the computing nodes in the sub-cluster, and after converting the performance data into performance data in a predetermined format, send it to the data aggregation component on the management node in the sub-cluster; through the data aggregation component on the management node in the sub-cluster, aggregate the performance data in the predetermined format of each of the computing nodes in the sub-cluster, and perform compression processing to obtain compressed data of the sub-cluster; through a data receiving component, receive the compressed data of each sub-cluster sent by each data aggregation component via a management network; through a data processing component, decompress and perform format conversion on the compressed data of each sub-cluster to obtain data in a target format, and store the data in the target format.
[0008] Optionally, according to the method for transmitting performance data of a supercomputing center of the present invention, the management nodes in each sub-cluster respectively correspond to different current sending times. The method further includes: through the data aggregation component on the management node in the sub-cluster, based on the current sending time corresponding to the management node, send the compressed data of the sub-cluster to the data receiving component via the management network.
[0009] Optionally, according to the method for transmitting performance data of a supercomputing center of the present invention, aggregating the performance data in the predetermined format of each of the computing nodes in the sub-cluster and performing compression processing to obtain the compressed data of the sub-cluster includes: through the data aggregation component on the management node in the sub-cluster, aggregating the performance data in the predetermined format of each of the computing nodes in the sub-cluster, and performing compression processing using the lz4 compression algorithm to obtain the compressed data of the sub-cluster.
[0010] Optionally, according to the method for transmitting performance data of the supercomputing center of the present invention, the data receiving component is a Kafka cluster; through the data receiving component, compressed data of each sub-cluster sent by each data aggregation component via the management network is received, including: through the data receiving component, compressed data of each sub-cluster sent by each data aggregation component via the management network is received, and the compressed data of each sub-cluster is written into the WAL write-ahead log for subscription by the data processing component.
[0011] Optionally, according to the method for transmitting performance data of the supercomputing center of the present invention, through the data processing component, the compressed data of each sub-cluster is decompressed and format-converted to obtain target format data, including: through the data processing component, the compressed data of each sub-cluster is decompressed to obtain decompressed data in a predetermined format, and the decompressed data is expanded into a single line of data to obtain target format data.
[0012] Optionally, according to the method for transmitting performance data of the supercomputing center of the present invention, the target format data includes first format metadata and target format performance data, and the first format metadata includes a cluster code, a sub-cluster code, and a computing node name; storing the target format data includes: storing the first format metadata in a MySQL database and storing the target format performance data in a ClickHouse database.
[0013] Optionally, according to the method for transmitting performance data of the supercomputing center of the present invention, the data receiving component is communicatively connected to a plurality of data processing components, and the data processing components include a queue data consumption component, a queue data conversion component, and a queue data storage component that are sequentially coupled; through the data processing component, the compressed data of each sub-cluster is decompressed and format-converted to obtain target format data, and the target format data is stored, including: through the queue data consumption component, subscribing to the compressed data of one or more topics in the WAL write-ahead log and forwarding the compressed data of each topic to the queue data conversion component corresponding to the topic; through the queue data conversion component, decompressing the compressed data to obtain decompressed data in a predetermined format, and expanding the decompressed data into a single line of data to obtain target format data, where the target format data includes first format metadata and target format performance data; through the queue data storage component, storing the first format metadata in a MySQL database and storing the target format performance data in a ClickHouse database.
[0014] Optionally, according to the method for transmitting back the performance data of the supercomputing center of the present invention, the predetermined format includes a name, a tags part, and a columes part; the tags part includes two first elements, and the two first elements respectively represent the environment to which it belongs and the type to which the data belongs; the columes part includes nine second elements, which respectively represent the supercomputing center, the cluster, the sub-cluster, the cabinet, the chassis, the host, the CPU family, the CPU domain, and the specified CPU.
[0015] Optionally, according to the method for transmitting back the performance data of the supercomputing center of the present invention, the performance data includes CPU utilization rate, memory utilization rate, disk utilization rate, disk I / O rate, management network I / O rate, high-speed network rate, and custom performance metrics.
[0016] Optionally, according to the method for transmitting back the performance data of the supercomputing center of the present invention, each of the sub-clusters further includes a switch, and each of the computing nodes in the sub-cluster is adapted to communicate via the switch.
[0017] According to one aspect of the present invention, there is provided a system for transmitting back the performance data of a supercomputing center. Each of the computing nodes of the supercomputing center is adapted to be divided into a plurality of sub-clusters. Each of the sub-clusters respectively includes a plurality of the computing nodes and a management node communicatively connected to the plurality of computing nodes. The system includes: a plurality of data acquisition components deployed on each of the computing nodes in each of the sub-clusters, adapted to periodically acquire the performance data of each of the computing nodes in the sub-cluster, and after converting the performance data into performance data in a predetermined format, send it to a data aggregation component on the management node in the sub-cluster; a data aggregation component deployed on the management node in each of the sub-clusters, adapted to aggregate the performance data in a predetermined format of each of the computing nodes in the sub-cluster, and perform compression processing to obtain the compressed data of the sub-cluster; a data receiving component communicatively connected to the management nodes in each of the sub-clusters via a management network, adapted to receive the compressed data of each of the sub-clusters sent by each of the data aggregation components via the management network; and a data processing component communicatively connected to the data receiving component, adapted to decompress and perform format conversion on the compressed data of each of the sub-clusters to obtain data in a target format, and store the data in the target format.
[0018] The performance data feedback system of the supercomputing center according to the present invention includes a plurality of the data processing components, and further includes a data storage component, wherein the data storage component is communicatively connected to each of the data processing components; the data processing component includes a queue data consumption component, a queue data conversion component, and a queue data transfer and storage component that are sequentially coupled; the data storage component includes a MySQL database and a Clickhouse database; the queue data consumption component is adapted to subscribe to the compressed data of one or more topics in the WAL write-ahead log, and forward the compressed data of each topic to the queue data conversion component corresponding to the topic; the queue data conversion component is adapted to decompress the compressed data to obtain decompressed data in a predetermined format, and expand the decompressed data into a row of data to obtain target format data, wherein the target format data includes first format metadata and target format performance data; the queue data transfer and storage component is adapted to store the first format metadata in the MySQL database and store the target format performance data in the Clickhouse database.
[0019] According to one aspect of the present invention, there is provided a computing device, including: at least one processor; a memory storing program instructions, wherein the program instructions are configured to be executed by the at least one processor, and the program instructions include instructions for executing the supercomputing center performance data feedback method as described above.
[0020] According to one aspect of the present invention, there is provided a computer program product, including computer programs / instructions, wherein when the computer programs / instructions are executed by a processor, the method as described above is implemented.
[0021] According to one aspect of the present invention, there is provided a readable storage medium storing program instructions, which when read and executed by a computing device, cause the computing device to execute the supercomputing center performance data feedback method as described above.
[0022] According to the technical solution of the present invention, a method for transmitting performance data of a supercomputing center is provided. By dividing each computing node of the supercomputing center into multiple sub-clusters, data acquisition components are deployed on the computing nodes in each sub-cluster to periodically collect the performance data of each computing node and convert it into a predetermined format. A data aggregation component is deployed on the management node in the sub-cluster to aggregate and compress the performance data in the predetermined format of each computing node in the sub-cluster at the same moment, and then send it to the data receiving component via the management network. The data processing component decompresses the compressed data of each sub-cluster and performs format conversion to store the data in the target format. Based on this, not only is the distributed acquisition and transmission of the performance data of the supercomputing center realized, which can reduce data transmission failures and disperse data transmission pressure; moreover, by agreeing on the data format, the performance data in the predetermined format is transmitted and stored, which can save storage space, reduce data transmission and storage pressure, and can reduce the risk of data leakage; the data compression mechanism can greatly reduce the bandwidth pressure on the management network. It can be seen that the present invention can realize the fast, efficient and complete acquisition, transmission and storage of the performance data of the supercomputing center.
[0023] In addition, according to the technical solution of the present invention, the data aggregation components on the management nodes of each sub-cluster respectively send data to the data receiving component via the management network based on different current sending times. In this way, the bandwidth pressure on the management network is further reduced, and it is possible to avoid the management nodes in multiple sub-clusters sending data via the management network simultaneously and exceeding the bandwidth of the management network.
[0024] In addition, implementing the data receiving component based on the Kafka cluster can achieve an efficient and high-performance data receiving service, avoiding data receiving blockage and loss.
[0025] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present invention more obvious and understandable, the following specifically illustrates the specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] To achieve the above and related purposes, certain illustrative aspects are described herein in conjunction with the following description and drawings, which indicate various ways in which the principles disclosed herein can be practiced, and all aspects and their equivalent aspects are intended to fall within the scope of the claimed subject matter. By reading the following detailed description in conjunction with the drawings, the above and other purposes, features and advantages of the present disclosure will become more apparent. Throughout the present disclosure, the same reference numerals generally refer to the same components or elements.
[0027] Figure 1Shows a schematic diagram of a supercomputing center performance data feedback system 100 provided according to an embodiment of the present invention;
[0028] Figure 2 Shows a schematic diagram of a computing device 200 provided according to an embodiment of the present invention;
[0029] Figure 3 Shows a flowchart of a supercomputing center performance data feedback method 300 provided according to an embodiment of the present invention;
[0030] Figure 4 Shows a schematic diagram of a data receiving component 130 and a data processing component 140 in some embodiments of the present invention. Detailed implementation manners
[0031] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0032] Aiming at the problems of large amount of performance data to be collected in the supercomputing center, large data transmission and storage pressure, and the risk of data leakage, the embodiments of the present invention provide a supercomputing center performance data feedback method, which can perform distributed collection and transmission of supercomputing center performance data, reduce data transmission and storage pressure, save storage space, and can reduce the risk of data leakage, and finally can realize fast, efficient, and complete collection, feedback, and storage of supercomputing center performance data.
[0033] The supercomputing center performance data feedback method provided according to the embodiments of the present invention can be implemented in a supercomputing center performance data feedback system. The following introduces the supercomputing center performance data feedback system of the present invention.
[0034] Figure 1 Shows a schematic diagram of a supercomputing center performance data feedback system 100 provided according to an embodiment of the present invention. In some embodiments, the supercomputing center can be implemented as an ultra-large-scale information technology innovation computing center, but the present invention is not limited thereto.
[0035] It should be noted that the supercomputing center (cluster) has multiple computing nodes, generally with more than 200,000 computing nodes. If all computing nodes transmit performance data to the data receiving server at the same time, then 200,000 data transmission requests of 15KB will be generated at each data transmission time point, and the total data transmission size is 200,000 * 15KB = 2.86GB. In this way, not only does the total data volume exceed the bandwidth of the management network, which is 1Gbps, but also the number of concurrent requests will cause great network pressure on the switch and the data receiving server.
[0036] Therefore, in the embodiments of the present invention, as Figure 1 shown, all computing nodes of the supercomputing center are logically divided so that each computing node of the supercomputing center is divided into multiple sub-clusters. In other words, each computing node of the supercomputing center can be divided into multiple sub-clusters, where each sub-cluster respectively includes multiple computing nodes and a management node communicatively connected to the multiple computing nodes. Here, for each sub-cluster, the management node in the sub-cluster can be communicatively connected to each computing node in the sub-cluster respectively. In some embodiments, each sub-cluster further includes at least one switch, and the computing nodes in the sub-cluster can communicate with each other via the switch in the sub-cluster.
[0037] For example, for a supercomputing center with more than 200,000 computing nodes, every 4,096 computing nodes can be respectively divided into a logically sub-cluster, so that ultimately about 49 sub-clusters (sub-cluster 1, sub-cluster 2,..., sub-cluster 49) can be formed.
[0038] In the embodiments of the present invention, a data acquisition component 110 is respectively deployed on each computing node in each sub-cluster, and a data aggregation component 120 is deployed on the management node in each sub-cluster. Among them, the data acquisition component 110 on each computing node in the sub-cluster is used to acquire the performance data of the computing node and can send the performance data of the computing node it acquires to the data aggregation component 120 on the management node in the sub-cluster. Furthermore, the data aggregation component 120 can aggregate and compress the performance data of each computing node in the sub-cluster.
[0039] As Figure 1As shown, in an embodiment of the present invention, the supercomputing center performance data feedback system 100 includes various data acquisition components 110 deployed on each computing node in each sub-cluster, and a data aggregation component 120 deployed on the management node in each sub-cluster. In addition, the supercomputing center performance data feedback system 100 further includes a data receiving component 130 and a data processing component 140. Among them, the management nodes in each sub-cluster can be communicatively connected to the data receiving component 130 via a management network, and the data receiving component 130 is communicatively connected to the data processing component 140.
[0040] As Figure 1 shown, in some embodiments, C can be used to represent a computing node. For example, multiple computing nodes (e.g., 4096 computing nodes) in each sub-cluster can be respectively represented as C1, C2, … C4096. M can be used to represent a management node, and S can be used to represent a switch.
[0041] In the supercomputing center performance data feedback system 100 of the present invention, for each sub-cluster, each data acquisition component 110 deployed on each computing node in the sub-cluster can periodically collect the performance data of each computing node in the sub-cluster, and after converting the performance data of each computing node in the sub-cluster it collects into performance data in a predetermined format, send it to the data aggregation component 120 deployed on the management node in the sub-cluster. The data aggregation component 120 deployed on the management node in the sub-cluster can aggregate the performance data (predetermined format performance data) of each computing node in the sub-cluster collected by each data acquisition component 110 in the sub-cluster at the same moment, and perform compression processing to obtain the compressed data of the sub-cluster. Subsequently, the data aggregation component 120 on the management node can send the compressed data of the sub-cluster to the data receiving component 130 via the management network. The data receiving component 130 can receive the compressed data of each sub-cluster sent by each data aggregation component 120 via the management network, and after confirming the receipt of the compressed data of each sub-cluster, the data receiving component 130 can send a data consumption event to the data processing component 140. The data processing component 140 can decompress and convert the format of the compressed data of each sub-cluster to obtain data in a target format, and then store the data in the target format.
[0042] In some embodiments, each data acquisition component 110 on each computing node can collect the performance data of the corresponding computing node every minute.
[0043] In some embodiments, the management nodes in each sub-cluster respectively correspond to different current sending times. For example, the current sending time corresponding to the management node M1 in sub-cluster 1 is the current 1 second, and the current sending time corresponding to the management node M2 in sub-cluster 2 is the current 2 seconds. After the data aggregation component 120 on the management node in the sub-cluster aggregates the predetermined format performance data of each computing node in the sub-cluster collected at the same moment by each data collection component 110 on each computing node in the sub-cluster and performs compression processing to obtain the compressed data of the sub-cluster, the data aggregation component 120 on the management node in the sub-cluster can, based on the current sending time corresponding to this management node, send the compressed data of the sub-cluster to the data receiving component 130 via the management network. In this way, it is possible to avoid multiple management nodes in multiple sub-clusters simultaneously sending data to the data receiving component 130 via the management network and exceeding the bandwidth of the management network, which is 1 Gbps.
[0044] In some embodiments, as Figure 1 shown, the supercomputing center performance data feedback system 100 includes a plurality of distributed data processing components 140. The data receiving component 130 can be communicatively connected to the plurality of data processing components 140 to perform distributed data processing through the plurality of data processing components 140.
[0045] In some embodiments, as Figure 1 shown, the supercomputing center performance data feedback system 100 further includes a data storage component 150. The data storage component 150 can be communicatively connected to each data processing component 140. The data processing component 140 can store the target format data in the data storage component 150.
[0046] In some embodiments, the data processing component 140 includes a MySQL database and a ClickHouse database. Among them, the MySQL database can be implemented as a MySQL cluster, and the MySQL cluster includes a MySQL proxy and a plurality of MySQL nodes. The ClickHouse database can be implemented as a ClickHouse cluster, and the ClickHouse cluster includes a ClickHouse proxy and a plurality of ClickHouse nodes.
[0047] In some embodiments, as Figure 1 shown, the supercomputing center performance data feedback system 100 further includes a data display layer 160 communicatively connected to the data storage component 150. The data display layer 160 can obtain the target format data from the data storage component 150 and can display the target format data so as to be displayed to the user via the client. In this way, it is convenient for the user side to analyze the performance data of each computing node of the supercomputing center and provide solid data support for optimizing the hardware of the supercomputing center.
[0048] In an embodiment of the present invention, the supercomputing center performance data feedback system 100 can be configured to execute the supercomputing center performance data feedback method 300, and the supercomputing center performance data feedback method 300 of the present invention will be described in detail below. Among them, the data acquisition component 110, the data aggregation component 120, the data receiving component 130, and the data processing component 140 in the supercomputing center performance data feedback system 100 can respectively execute steps 310 to 340 in the supercomputing center performance data feedback method 300. For the specific execution logics of the data acquisition component 110, the data aggregation component 120, the data receiving component 130, and the data processing component 140 in the supercomputing center performance data feedback system 100, reference can be made to the relevant descriptions of steps 310 to 340 in the supercomputing center performance data feedback method 300 below.
[0049] According to the supercomputing center performance data feedback system 100 of the present invention, by dividing each computing node of the supercomputing center into multiple sub-clusters, by deploying a data acquisition component on the computing nodes in each sub-cluster to periodically acquire the performance data of each computing node and convert it into a predetermined format, by deploying a data aggregation component on the management node in the sub-cluster to aggregate and compress the performance data in a predetermined format of each computing node at the same moment in the sub-cluster, and then sending it to the data receiving component via the management network, and then decompressing the compressed data of each sub-cluster by the data processing component and performing format conversion to target format data for storage. Based on this, not only is the distributed acquisition and transmission of the supercomputing center performance data realized, which can reduce data transmission failures and disperse data transmission pressure; moreover, by agreeing on the data format and feeding back and storing the performance data in a predetermined format, it is possible to save storage space, reduce data transmission and storage pressure, and reduce the risk of data leakage; through the data compression mechanism, the bandwidth pressure on the management network can be greatly reduced. It can be seen that the present invention can achieve fast, efficient, and complete acquisition, feedback, and storage of the performance data of the supercomputing center.
[0050] In some embodiments, each of the above-mentioned computing nodes, management nodes, data receiving component 130, and data processing component 140 can be respectively implemented as a computing device described below, so that the supercomputing center performance data feedback method 300 of the present invention can be executed in the computing device.
[0051] Figure 2 Shows a schematic diagram of a computing device 200 provided according to an embodiment of the present invention. As Figure 2As shown, in a basic configuration, computing device 200 includes at least one processing unit 202 and system memory 204. According to one aspect, depending on the configuration and type of the computing device, processing unit 202 may be implemented as a processor. System memory 204 includes, but is not limited to, volatile storage (e.g., random access memory), non-volatile storage (e.g., read-only memory), flash memory, or any combination of such memories. According to one aspect, operating system 205 is included in system memory 204.
[0052] According to one aspect, operating system 205 is suitable for controlling the operation of computing device 200, for example. Additionally, the examples are practiced in conjunction with a graphics library, other operating systems, or any other application programs, and are not limited to any particular application or system. In Figure 2 this basic configuration is illustrated by those components within the dashed lines. According to one aspect, computing device 200 has additional features or functionality. For example, according to one aspect, computing device 200 includes additional data storage devices (removable and / or non-removable), such as magnetic disks, optical disks, or magnetic tapes. Such additional storage is Figure 2 illustrated by removable storage device 209 and non-removable storage device 210 in
[0053] As stated above, according to one aspect, program modules 203 are stored in system memory 204. According to one aspect, program modules 203 may include one or more application programs, and the present invention does not limit the type of application programs. For example, application programs may include: email and contact applications, word processing applications, spreadsheet applications, database applications, slide show applications, painting or computer-aided applications, web browser applications, etc.
[0054] In an embodiment according to the present invention, program modules 203 include multiple program instructions for executing the supercomputing center performance data feedback method 300 of the present invention.
[0055] According to one aspect, the examples may be practiced in a circuit including discrete electronic components, a package or integrated electronic chip containing logic gates, a circuit utilizing a microprocessor, or on a single chip containing electronic components or a microprocessor. For example, it may be via one in which Figure 2Each or many of the components shown in [the figure] may be implemented as a system-on-chip (SOC) integrated on a single integrated circuit. According to one aspect, such an SOC device may include one or more processing units, graphics units, communication units, system virtualization units, and various application functions, all integrated (or "burned") onto a chip substrate as a single integrated circuit. When operating via an SOC, the functions described herein may be operated via dedicated logic integrated with other components of computing device 200 on a single integrated circuit (chip). Embodiments of the present invention may also be practiced using other technologies capable of performing logical operations (such as AND, OR, and NOT), including but not limited to mechanical, optical, fluidic, and quantum technologies. Additionally, embodiments of the present invention may be practiced within a general-purpose computer or in any other circuit or system.
[0056] According to one aspect, computing device 200 may also have one or more input devices 212, such as a keyboard, mouse, pen, voice input device, touch input device, etc. Output devices 214 may also be included, such as a display, speaker, printer, etc. The foregoing devices are examples and other devices may also be used. Computing device 200 may include one or more communication connections 216 that allow communication with other computing devices 218. Examples of suitable communication connections 216 include but are not limited to: RF transmitter, receiver, and / or transceiver circuits; Universal Serial Bus (USB), parallel, and / or serial ports.
[0057] As used herein, the term computer-readable medium includes computer storage media. Computer storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, or program modules). System memory 204, removable storage device 209, and non-removable storage device 210 are all examples of computer storage media (i.e., memory storage). Computer storage media may include random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, disk storage or other magnetic storage devices, or any other article that can be used to store information and can be accessed by computing device 200. According to one aspect, any such computer storage media may be part of computing device 200. Computer storage media does not include carrier waves or other propagated data signals.
[0058] According to one aspect, a communication medium is implemented by computer-readable instructions, data structures, program modules, or other data in a modulated data signal (e.g., a carrier wave or other transmission mechanism), and includes any information delivery medium. According to one aspect, the term "modulated data signal" describes a signal having one or more sets of characteristics or a signal that has been altered in a manner that encodes information in the signal. By way of example and not limitation, communication media include wired media such as a wired network or a direct wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media.
[0059] In an embodiment according to the present invention, the computing device 200 is configured to execute the supercomputing center performance data feedback method 300. The computing device 200 includes one or more processors and one or more readable storage media storing program instructions, which when configured to be executed by the one or more processors, enable the computing device to execute the supercomputing center performance data feedback method 300 in the embodiments of the present invention.
[0060] In an embodiment of the present invention, the supercomputing center performance data feedback method 300 may be executed by the aforementioned supercomputing center performance data feedback system 100.
[0061] According to the foregoing, in an embodiment of the present invention, each computing node of the supercomputing center may be divided into multiple sub-clusters, where each sub-cluster respectively includes multiple computing nodes and a management node communicatively connected to the multiple computing nodes. Data acquisition components 110 are respectively deployed on each of the computing nodes in each sub-cluster, and a data aggregation component 120 is deployed on the management node in each sub-cluster.
[0062] The supercomputing center performance data feedback method 300 in the embodiments of the present invention will be described in detail below.
[0063] Figure 3 FIG. shows a flowchart of a supercomputing center performance data feedback method 300 provided according to an embodiment of the present invention. As Figure 3 shown, the supercomputing center performance data feedback method 300 includes the following steps 310 to 340.
[0064] Step 310: For each sub-cluster, the performance data of each computing node in the sub-cluster can be periodically collected through each data acquisition component 110 on each computing node in the sub-cluster, and the performance data of each computing node can be converted into performance data in a predetermined format, and the performance data in the predetermined format of each computing node can be sent to the data aggregation component 120 deployed on the management node in the sub-cluster.
[0065] In some embodiments, each data acquisition component 110 on each computing node may collect the performance data of the corresponding computing node every minute, and convert the performance data into performance data in a predetermined format.
[0066] In some embodiments, the performance data includes but is not limited to CPU usage rate, memory usage rate, disk usage rate, disk I / O rate, management network I / O rate, high-speed network rate, custom performance metrics (such as microarchitecture performance metrics).
[0067] Step 320: Through the data aggregation component 120 on the management node in the sub-cluster, aggregate the performance data in the predetermined format of each computing node in the sub-cluster collected at the same time by each data acquisition component 110 on each computing node in the sub-cluster, and perform compression processing to obtain the compressed data of the sub-cluster. Subsequently, the data aggregation component 120 on the management node may send the compressed data of the sub-cluster to the data receiving component 130 via the management network.
[0068] Step 330: Receive the compressed data of each sub-cluster sent by each data aggregation component 120 via the management network through the data receiving component 130. And after the data receiving component 130 confirms that it has received the compressed data of each sub-cluster, it may send a data consumption event to the data processing component 140.
[0069] Step 340: Decompress and convert the format of the compressed data of each sub-cluster through the data processing component 140 to obtain the data in the target format, and store the data in the target format. In some embodiments, the data processing component 140 may store the data in the target format in the data storage component 150.
[0070] In the embodiments of the present invention, the data processing component 140 can obtain the decompressed data in the predetermined format by decompressing the compressed data of each sub-cluster. Furthermore, the decompressed data in the predetermined format can be converted into the data in the target format and stored.
[0071] It should be noted that by converting the performance data of each computing node into a predetermined format, and backhauling and storing the performance data in the predetermined format, the present invention can save storage space, reduce the network and database storage pressure, and reduce the risk of data leakage.
[0072] In some embodiments, in step 320, after aggregating the performance data in the predetermined format of each computing node in the sub-cluster collected at regular intervals, the data aggregation component 120 on the management node in the sub-cluster may perform compression processing using the lz4 compression algorithm to obtain the compressed data of the sub-cluster.
[0073] It should be noted that for a sub-cluster with 4096 computing nodes, the data aggregation component 120 deployed on the management node needs to summarize the performance data from all nodes (4096 nodes) in the sub-cluster at the same moment. At this time, the amount of data that the data aggregation component 120 needs to process is approximately 4096 * 15 KB = 60 MB. By using the lz4 compression algorithm to compress the performance data from all nodes in the sub-cluster (lossless compression), the compressed data of the sub-cluster obtained after compression is approximately 1 / 10 of the original data, about 6 MB.
[0074] If the data aggregation components 120 on the management nodes in 49 sub-clusters simultaneously send the compressed data of each sub-cluster to the data receiving component 130 via the management network, a data volume of approximately 6 MB * 49 = 294 MB will be generated, and this data volume will also far exceed the bandwidth of the management network, which is 1 Gbps.
[0075] In view of this, in some embodiments, different current sending times can be set for the management nodes in each sub-cluster, that is, the management nodes in each sub-cluster correspond to different current sending times. For example, the current sending time corresponding to the management node M1 in sub-cluster 1 is the current 1 second, the current sending time corresponding to the management node M2 in sub-cluster 2 is the current 2 seconds, and so on. Moreover, it is required that the data aggregation component 120 on the management node in each sub-cluster send data to the data receiving component 130 according to the current sending time corresponding to this management node. In this way, it can be ensured that the amount of data sent and received simultaneously through the management network per second is approximately 6 MB, thereby avoiding the situation where the management nodes in multiple sub-clusters simultaneously send data to the data receiving component 130 via the management network and exceeding the bandwidth of the management network, which is 1 Gbps.
[0076] Specifically, in step 320, after the data aggregation component 120 on the management node in the sub-cluster aggregates the performance data in a predetermined format of each computing node in the sub-cluster regularly collected by each data collection component 110 at the same moment in the sub-cluster and performs compression processing to obtain the compressed data of the sub-cluster, the data aggregation component 120 on the management node in the sub-cluster can, based on the current sending time corresponding to this management node, send the compressed data of the sub-cluster to the data receiving component 130 via the management network. In this way, it can be avoided that the management nodes in multiple sub-clusters simultaneously send data to the data receiving component 130 via the management network and exceed the bandwidth of the management network, which is 1 Gbps.
[0077] Correspondingly, in step 330, the data receiving component 130 can receive the compressed data of each sub-cluster sent by each data aggregation component 120 based on the current sending time corresponding to the management node in the sub-cluster where it is located.
[0078] In some embodiments, the predetermined format may include a name (the value corresponding to name may be the name of a computing node), a tags section, and a columes section, and may also include a timestamp and a value.
[0079] Among them, the tags section contains two first elements, and the two first elements can be distinguished by different subscripts, respectively representing the environment to which it belongs and the type to which the data belongs. For example, among the two first elements in the tags section, the first element with subscript 0 represents the environment to which it belongs; the first element with subscript 1 represents the type to which the data belongs, and when the value of the first element with subscript 1 is 0, it means the data is performance data (applicable to the embodiments of the present invention), and when the value of the first element with subscript 1 is 1, it means the data is microarchitecture data.
[0080] The columes section contains a total of nine second elements, and the nine second elements can be distinguished by different subscripts, respectively representing the supercomputing center, the cluster (the corresponding value may be the cluster code), the sub-cluster (the corresponding value may be the sub-cluster code), the cabinet, the chassis, the host, the CPU family, the CPU domain, and the specified CPU. For example, among the nine second elements in the columes section, the second element with subscript 0 represents the supercomputing center, the second element with subscript 1 represents the cluster, the second element with subscript 2 represents the sub-cluster, the second element with subscript 3 represents the cabinet, the second element with subscript 4 represents the chassis, the second element with subscript 5 represents the host, the second element with subscript 6 represents the CPU family, the second element with subscript 7 represents the CPU domain, and the second element with subscript 8 represents the specified CPU.
[0081] In a specific embodiment, the predetermined format may be represented as follows:
[0082] {
[0083] "name":"string”,
[0084] "tags":[int,int],
[0085] "columes":[int,int,int,int,int,int,int,int,int],
[0086] "timestamp":long,
[0087] "value":double
[0088] }
[0089] In some embodiments, in step 340, the data processing component 140 decompresses the compressed data of each sub-cluster, and decompressed data in a predetermined format can be obtained. Furthermore, the data processing component 140 can convert the decompressed data in the predetermined format into target format data according to the conversion rule. Specifically, the data processing component 140 can expand the decompressed data in the predetermined format into a row of data, thereby obtaining the target format data. Here, it should be understood that the data to be expanded includes the data of each part such as name, tags, columes, timestamp, and value in the decompressed data in the predetermined format, and the respective element data in tags and columes need to be correspondingly expanded.
[0090] After that, in step 350, the data processing component 140 can store the target format data in the data storage component 150. In some embodiments, the data processing component 140 includes a MySQL database and a ClickHouse database. Among them, the MySQL database can be implemented as a MySQL cluster, and the MySQL cluster includes a MySQL proxy and multiple MySQL nodes. The ClickHouse database can be implemented as a ClickHouse cluster, and the ClickHouse cluster includes a ClickHouse proxy and multiple ClickHouse nodes. The target format data obtained after the above conversion includes first format metadata and target format performance data. The first format metadata can include, for example, cluster encoding, sub-cluster encoding, and computing node name.
[0091] It should be noted that the first format is a data format that can be recognized by the MySQL database, and the target format is a data format that can be recognized by the ClickHouse database. Based on this, in step 350, the data processing component 140 can store the first format metadata in the MySQL database and store the target format performance data in the ClickHouse database.
[0092] In a specific embodiment, the data processing component 140 can store the target format performance data in the ClickHouse database in the int64 format, which can greatly increase the data compression ratio and reduce the data storage pressure.
[0093] Figure 4 The schematic diagram of the data receiving component 130 and the data processing component 140 in some embodiments according to the present invention is shown.
[0094] In some embodiments, such as Figure 4As shown, the data receiving component 130 can be implemented as a Kafka cluster including multiple Kafka nodes. Based on this, an efficient and high-performance data receiving service can be implemented to avoid data receiving blockage and loss.
[0095] In some embodiments, in step 330, after the data receiving component 130 (Kafka cluster) receives the compressed data of each sub-cluster sent by each data aggregation component 120 via the management network, the compressed data of each sub-cluster can be written into the WAL (Write Ahead Log) pre-write log for the data processing component 140 (queue data consumption component) to subscribe to the compressed data of each sub-cluster in the WAL pre-write log. It should be noted that the WAL pre-write log is a common means in database systems to ensure the atomicity and durability of data operations. Since WAL logs are all sequential write operations, the data write performance is very good and it is very suitable for high-concurrency data write scenarios.
[0096] In addition, after writing the compressed data of each sub-cluster into the WAL pre-write log, the data receiving component 130 can also send receipt information to each data aggregation component 120 to indicate that the compressed data of each sub-cluster sent by each data aggregation component 120 has been received.
[0097] In some embodiments, as Figure 1 shown, the supercomputing center performance data feedback system 100 includes multiple data processing components 140, and the data receiving component 130 can be communicatively connected to the multiple data processing components 140 to perform distributed data processing through the multiple data processing components 140.
[0098] Figure 4 Only one data processing component 140 communicatively connected to the data receiving component 130 is exemplarily shown in Figure 4 As shown, the data processing component 140 includes a queue data consumption component, a queue data conversion component, and a queue data storage component that are sequentially coupled.
[0099] In some embodiments, in step 340, the specific process of decompressing and converting the format of the compressed data of each sub-cluster by the data processing component 140 to obtain the target format data and storing the target format data is as follows: The queue data consumption component subscribes to the compressed data of one or more topics in the WAL write-ahead log. After obtaining the compressed data of one or more topics from the WAL write-ahead log, the compressed data can be routed and forwarded according to the topic to forward the compressed data of each topic to the queue data conversion component corresponding to the topic. Subsequently, the queue data conversion component decompresses the compressed data to obtain the decompressed data in a predetermined format, and then the decompressed data in the predetermined format can be expanded into a row of data, that is, the target format data is obtained. Subsequently, the target format data is sent to the queue data storage component. Among them, the target format data includes the first format metadata and the target format performance data. Finally, the queue data storage component can store the target format data in the data storage component 150. Specifically, the first format metadata in the target format data is stored in the MySQL database through the queue data storage component, and the target format performance data in the target format data is stored in the Clickhouse database. In a specific embodiment, the queue data storage component of the data processing component 140 can write the target format performance data into the Clickhouse database in the int64 format, which can greatly increase the data compression ratio and reduce the data storage pressure.
[0100] In some embodiments, after the queue data conversion component decompresses and converts the format of the compressed data to obtain the target format data, it can also save the target format data in the memory. Specifically, the target format data can be persistently saved according to the following rules: If the target format data is greater than or equal to 1024 pieces, the target format data is batch-saved once; if the target format data is less than 1024 pieces but the save time is greater than 1 minute, the target format data is batch-saved once; if the memory space occupied by the target format data is greater than 4MB, the target format data is actively batch-saved once.
[0101] In some embodiments, the data display layer 160 can obtain the target format data from the data storage component 150 and can display the target format data so as to be displayed to the user via the client. In this way, it is convenient for the user side to analyze the performance data of each computing node of the supercomputing center and provide solid data support for optimizing the hardware of the supercomputing center.
[0102] According to the supercomputing center performance data feedback method 300 of the present invention, by dividing each computing node of the supercomputing center into multiple sub-clusters, deploying data acquisition components on the computing nodes in each sub-cluster to regularly collect the performance data of each computing node and convert it into a predetermined format, deploying data aggregation components on the management nodes in the sub-clusters to aggregate and compress the performance data in the predetermined format of each computing node at the same moment in the sub-cluster, and then sending it to the data receiving component via the management network, and then decompressing the compressed data of each sub-cluster through the data processing component and performing format conversion to target format data for storage. Based on this, not only is the distributed acquisition and transmission of the supercomputing center performance data realized, which can reduce data transmission failures and disperse data transmission pressure; moreover, by agreeing on the data format and feedbacking and storing the performance data in the predetermined format, it can save storage space, reduce data transmission and storage pressure, and reduce the risk of data leakage; through the data compression mechanism, the bandwidth pressure on the management network can be greatly reduced. It can be seen that the present invention can realize the fast, efficient, and complete acquisition, feedback, and storage of the performance data of the supercomputing center.
[0103] In addition, according to the technical solution of the present invention, the data aggregation components on each sub-cluster management node respectively send data to the data receiving component via the management network based on different current sending times. In this way, the bandwidth pressure on the management network is further reduced, and it can avoid the management nodes in multiple sub-clusters from simultaneously sending data via the management network and exceeding the bandwidth of the management network.
[0104] In addition, implementing the data receiving component based on the Kafka cluster can achieve an efficient and high-performance data receiving service and avoid data receiving blocking and loss.
[0105] In addition, embodiments of the present invention also disclose: A8. The method according to any one of A1 - A7, wherein the predetermined format includes a name, a tags part, and a columes part; the tags part includes two first elements, and the two first elements respectively represent the environment to which it belongs and the type to which the data belongs; the columes part includes nine second elements, which respectively represent a supercomputing center, a cluster, a sub - cluster, a cabinet, a chassis, a host, a CPU family, a CPU domain, and a specified CPU. A9. The method according to any one of A1 - A8, wherein the performance data includes CPU usage rate, memory usage rate, disk usage rate, disk I / O rate, management network I / O rate, high - speed network rate, and custom performance metrics. A10. The method according to any one of A1 - A9, wherein each of the sub - clusters further includes a switch, and the computing nodes in the sub - cluster are adapted to communicate with each other via the switch. B14. A computer program product, including a computer program / instructions, wherein when the computer program / instructions are executed by a processor, the method according to any one of A1 - A10 is implemented. C15. A readable storage medium storing program instructions, when the program instructions are read and processed by a computing device, the computing device is caused to process the method according to any one of A1 - A10.
[0106] The various technologies described herein can be implemented in combination with hardware or software, or a combination thereof. Thus, the method and device of the present invention, or certain aspects or parts of the method and device of the present invention, may take the form of program code (i.e., instructions) embedded in a tangible medium, such as a removable hard disk, a USB flash drive, a floppy disk, a CD - ROM, or any other machine - readable storage medium. When the program is loaded into a machine such as a computer and executed by the machine, the machine becomes a device for practicing the present invention.
[0107] In the case where the program code is executed on a programmable computer, a mobile terminal generally includes a processor, a processor - readable storage medium (including volatile and non - volatile memories and / or storage elements), at least one input device, and at least one output device. Among them, the memory is configured to store the program code; the processor is configured to execute the method for transmitting performance data of the supercomputing center of the present invention according to the instructions in the program code stored in the memory.
[0108] By way of example and not limitation, readable media include readable storage media and communication media. Readable storage media store information such as computer - readable instructions, data structures, program modules, or other data. Communication media generally embody computer - readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and include any information - carrying medium. A combination of any of the above is also included within the scope of readable media.
[0109] In the specification provided herein, the algorithms and displays are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems may also be used in conjunction with the examples of the present invention. The structure required to construct such systems will be apparent from the above description. In addition, the present invention is not directed to any particular programming language. It should be understood that the teachings of the present invention can be implemented in various programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the present invention.
[0110] In the specification provided herein, numerous specific details are set forth. However, it can be understood that embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0111] Similarly, it should be understood that, for the purpose of streamlining the present disclosure and aiding in the understanding of one or more of the various inventive aspects, in the foregoing description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof.
[0112] Those skilled in the art will appreciate that the modules or units or components of the devices in the examples disclosed herein may be arranged in the devices as described in that embodiment, or alternatively may be located in one or more devices different from those of the example. The modules in the foregoing examples may be combined into one module or further divided into multiple sub-modules.
[0113] Unless otherwise specified, the use of ordinal numbers such as "first", "second", "third", etc. to describe ordinary objects merely indicates different instances of similar objects and is not intended to imply that the objects so described must have a given order in time, space, ranking, or in any other manner.
Claims
1. A method for transmitting performance data of a supercomputing center, wherein each computing node of the supercomputing center is suitable for being divided into a plurality of subclusters, each of the subclusters respectively comprises a plurality of the computing nodes and a management node in communication connection with the plurality of the computing nodes, each computing node in each of the subclusters is respectively deployed with a data collection component, and each management node in each of the subclusters is deployed with a data aggregation component; The method comprises: For each of the sub-clusters, the performance data of each computing node in the sub-cluster is collected periodically through each data collection component on each computing node in the sub-cluster, and the performance data is converted into a predetermined format and then sent to the data aggregation component on the management node in the sub-cluster; Aggregating the performance data of the predetermined format of each computing node in the sub-cluster through a data aggregation component on the management node in the sub-cluster, and performing compression processing to obtain compressed data of the sub-cluster; Receiving, by a data receiving component, compressed data of each of the sub-clusters sent by each of the data aggregation components via a management network; The compressed data of each sub-cluster is decompressed and format-converted through a data processing component to obtain target format data, and the target format data is stored.
2. The method of claim 1, wherein: The management nodes in each sub-cluster correspond to different current sending times; The method further comprises: The compressed data of the sub-cluster is sent to the data receiving component via the management network through the data aggregation component on the management node in the sub-cluster based on the current sending time corresponding to the management node.
3. The method according to claim 1 or 2, wherein: Aggregating the performance data of the predetermined format of each computing node in the sub-cluster through a data aggregation component on the management node in the sub-cluster, and performing compression processing to obtain compressed data of the sub-cluster, including: The performance data of the predetermined format of each computing node in the sub-cluster is aggregated by a data aggregation component on the management node in the sub-cluster, and is compressed using the lz4 compression algorithm to obtain compressed data of the sub-cluster.
4. The method according to any one of claims 1 to 3, wherein: The data receiving component is a Kafka cluster; receiving the compressed data of each sub-cluster sent by each data aggregation component via the management network through the data receiving component, including: The compressed data of each sub-cluster sent by each data aggregation component via the management network is received through the data receiving component, and the compressed data of each sub-cluster is written into the WAL write-ahead log for subscription by the data processing component.
5. The method according to any one of claims 1 to 4, wherein: Decompressing and formatting the compressed data of each sub-cluster through a data processing component to obtain target format data includes: The compressed data of each of the sub-clusters is decompressed by a data processing component to obtain decompressed data in a predetermined format, and the decompressed data is expanded into a row of data to obtain data in a target format.
6. The method according to any one of claims 1 to 5, wherein: The target format data includes first format metadata and target format performance data, wherein the first format metadata includes a cluster code, a subcluster code, and a computing node name; Storing the target format data includes: The first format metadata is stored in a MySQL database, and the target format performance data is stored in a Clickhouse database.
7. The method of claim 4, wherein: The data receiving component is in communication with a plurality of data processing components, wherein the data processing components include a queue data consumption component, a queue data conversion component, and a queue data transfer component coupled in sequence; Decompressing and formatting the compressed data of each of the sub-clusters through a data processing component to obtain target format data, and storing the target format data, including: Subscribe to the compressed data of one or more topics in the WAL write-ahead log through the queue data consumption component, and forward the compressed data of each topic to the queue data conversion component corresponding to the topic; Decompressing the compressed data through a queue data conversion component to obtain decompressed data in a predetermined format, and expanding the decompressed data into a row of data to obtain target format data, wherein the target format data includes first format metadata and target format performance data; The first format metadata is stored in the MySQL database through the queue data transfer component, and the target format performance data is stored in the Clickhouse database.
8. A supercomputing center performance data transmission system, wherein each computing node of the supercomputing center is suitable for being divided into multiple sub-clusters, each of the sub-clusters comprises multiple computing nodes and a management node in communication connection with the multiple computing nodes, the system comprising: Each data collection component deployed on each computing node in each of the sub-clusters is adapted to periodically collect performance data of each computing node in the sub-clusters, convert the performance data into performance data in a predetermined format, and then send the performance data to a data aggregation component on a management node in the sub-clusters; A data aggregation component deployed on the management node in each of the sub-clusters is adapted to aggregate the performance data of the predetermined format of each computing node in the sub-cluster and perform compression processing to obtain compressed data of the sub-cluster; A data receiving component, which is communicatively connected to the management node in each sub-cluster via the management network, and is adapted to receive the compressed data of each sub-cluster sent by each data aggregation component via the management network; The data processing component is communicatively connected with the data receiving component, and is adapted to decompress and convert the format of the compressed data of each of the sub-clusters to obtain target format data, and store the target format data.
9. The system of claim 8, wherein: The system includes a plurality of the data processing components, and also includes a data storage component, wherein the data storage component is in communication with each of the data processing components; the data processing component includes a queue data consumption component, a queue data conversion component, and a queue data transfer component coupled in sequence; the data storage component includes a MySQL database and a Clickhouse database; The queue data consumption component is adapted to subscribe to compressed data of one or more topics in the WAL write-ahead log, and forward the compressed data of each topic to the queue data conversion component corresponding to the topic; The queue data conversion component is adapted to decompress the compressed data to obtain decompressed data in a predetermined format, and expand the decompressed data into a row of data to obtain target format data, wherein the target format data includes first format metadata and target format performance data; The queue data transfer component is suitable for storing the first format metadata in a MySQL database and storing the target format performance data in a Clickhouse database.
10. A computing device comprising: at least one processor; and A memory storing program instructions, wherein the program instructions are configured to be processed by the at least one processor, and the program instructions include instructions for processing the method according to any one of claims 1 to 7.