Method for monitoring GPFS in real time based on middleware
By deploying the monitoring agent Agent on the GPFS cluster nodes, real-time monitoring and automatic alarms are realized, and the problem of management and monitoring relying on manual operations in the existing technology is solved, and system management efficiency and stability are improved.
Patent Information
- Application Number
- CN202510102427.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, the management and monitoring of GPFS file system rely on manual operations, resulting in high degree of human intervention, lack of real-time monitoring and automatic alarm mechanisms, and the inability to promptly detect and deal with file system failures or capacity exceeding the threshold.
Using a middleware-based real-time monitoring method, the lightweight monitoring agent Agent is deployed on the GPFS cluster node, and the CLI command collection is regularly called to collect multi-dimensional data, and the data is encapsulated into standardized messages. It consumes and stores it through message queues to realize real-time data display and automatic alarms.
Real-time monitoring and automatic alarm of GPFS file system are realized, which reduces human intervention, improves system management efficiency and stability, promptly detects and handles abnormal situations, and avoids production accidents.
Smart Images

Figure CN119938450A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of cloud computing and distributed file system management, and in particular relates to a method for real-time monitoring GPFS based on middleware. Background Art
[0002] In the field of cloud computing, GPFS file systems are widely used in high-performance computing and big data processing scenarios. However, in the prior art, the management and monitoring of GPFS file systems mainly rely on manual operations, such as creating, configuring, and mounting file systems through GUI interfaces or CLI commands. This approach has the following technical defects: High degree of human intervention: When you need to obtain data such as GPFS cluster status or file system capacity, you need to manually operate through the GUI or CLI, which is a complex and time-consuming process.
[0003] Lack of real-time monitoring and automatic alarm mechanism: When the GPFS cluster or file system fails or the capacity reaches the threshold, the existing method cannot automatically issue an alarm, resulting in the inability to take timely measures, which may cause serious production accidents.
[0004] Therefore, a more intelligent and automated technical means is needed to achieve real-time monitoring and management of the GPFS file system. Summary of the invention
[0005] In view of this, the present invention aims to propose a method for real-time monitoring of GPFS based on middleware to solve the problems of complex operation and inability to monitor and warn in real time in the prior art, and to improve system management efficiency and stability.
[0006] To achieve the above object, the technical solution of the present invention is achieved as follows: A method for real-time monitoring of GPFS based on middleware, comprising: Deploy a lightweight monitoring agent on all nodes of the GPFS cluster and set the agent to start with the node server; The Agent calls a preset CLI command set at fixed time intervals to collect multi-dimensional data of the GPFS cluster and file system; Encapsulating the data collected by the Agent into a standardized message with a timestamp and a unique identifier; The backend management system consumes the standardized messages in real time through the consumer module of the message queue, and performs format parsing, data verification and preprocessing on the messages; The backend management system stores the processed monitoring data in a database that supports time series query and aggregate analysis; Based on the stored data, the backend management system displays the operating status of the GPFS cluster and file system in real time through the Web interface and generates trend analysis charts.
[0007] Furthermore, the Agent calls a preset CLI command set at fixed time intervals to collect multi-dimensional data of the GPFS cluster and file system, including cluster health status, total capacity, used capacity, available capacity, number of client accesses, and read and write performance indicators.
[0008] Furthermore, when an abnormal state is detected, the background management system triggers a multi-channel alarm through a built-in alarm module; The multi-channel alarm includes email, SMS, and instant message push, notifying designated personnel to handle; The abnormal state includes capacity usage exceeding a threshold or node disconnection.
[0009] Furthermore, the CLI command set of the Agent includes preset template files of multiple commands, and users can add new commands by modifying the template files; the template files are stored in the local configuration directory of the Agent, and the Agent reads and parses the template files to generate an execution command queue to support latency analysis, file system fragmentation rate detection and I / O throughput statistics.
[0010] Furthermore, the Agent includes a local storage module. When the message queue middleware is unavailable, the Agent stores the collected data in a local cache directory in JSON format and sets a resend mechanism to ensure that the cached data is resent in timestamp order after the message queue is restored.
[0011] Furthermore, the message queue middleware adopts the Kafka distributed architecture and ensures high availability of data based on partition and copy mechanisms; message producers and consumers manage subscription and distribution processes through a distributed coordinator.
[0012] Furthermore, the background management system includes a rule engine module, which supports the configuration of threshold rule files through an interface. The rule files are stored in JSON format and include monitoring indicator names, threshold ranges, detection cycles, and alarm levels. The rule engine module matches the rule files when processing monitoring data, and generates alarm events when abnormalities are found.
[0013] Furthermore, the time series database adopts a hierarchical storage design, and the monitoring data is divided into hot and cold data according to the time window when writing; hot data is stored on high-performance disks for fast query and aggregate analysis, and cold data is regularly compressed and migrated to low-cost storage media to optimize data storage resources. Furthermore, the web interface is developed based on a front-end and back-end separation architecture. The front-end dynamically renders data through the React framework, and the back-end provides a query interface based on the RESTful API. The interface includes a cluster overview view and a node detail view. Users can switch to the historical performance indicator details page of the node by clicking the node icon.
[0014] Furthermore, the alarm module combines the trend analysis of historical data and uses time series algorithms (such as ARIMA model) to predict the trend of monitoring indicator data; when the predicted value exceeds the predefined threshold, an early warning message is generated and sent through the configured mail server or message push interface. Furthermore, the background management system includes a report generation module, and the report template is stored in XML format. The template defines the indicator name, time range and grouping method to be displayed; the user selects the template and time range through the Web interface to generate daily, weekly or monthly reports, and exports them in PDF or CSV format.
[0015] Furthermore, the system realizes data separation of multiple clusters by adding a cluster identification field in the message queue; the consumption module of the background management system writes the data into an independent database table according to the identification, and realizes cross-cluster monitoring through the cluster switching function.
[0016] Furthermore, the system provides data access capabilities through a RESTful API interface, which supports GET and POST request methods, allowing users to query real-time monitoring data, alarm information and historical performance indicators, and provides data writing functions to support third-party systems to upload external indicators.
[0017] Furthermore, the Agent deployment program includes a node automatic discovery module, which automatically identifies newly added cluster nodes by parsing cluster configuration files (such as / opt / gpfs / etc / node.conf), dynamically generates node monitoring configuration files and starts corresponding Agent instances, thereby achieving deployment without manual intervention.
[0018] Furthermore, the present solution discloses an electronic device, including a processor and a memory that is communicatively connected to the processor and is used to store instructions executable by the processor, wherein the processor is used to execute a method for real-time monitoring of GPFS based on middleware.
[0019] Furthermore, the present solution discloses a server, comprising at least one processor, and a memory communicatively connected to the processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor so that the at least one processor executes a method for real-time monitoring of GPFS based on middleware.
[0020] Furthermore, the present solution discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, a method for real-time monitoring of GPFS based on middleware is implemented.
[0021] Compared with the prior art, the method for real-time monitoring of GPFS based on middleware described in the present invention has the following beneficial effects: (1) The method for real-time monitoring of GPFS based on middleware described in the present invention realizes automatic real-time dynamic display of GPFS file system monitoring data; (2) The method for real-time monitoring of GPFS based on middleware described in the present invention realizes the monitoring of the entire life cycle of the GPFS file system, facilitates the subsequent analysis of the usage of GPFS, and further optimizes the subsequent usage plan; (3) The method of real-time monitoring of GPFS based on middleware described in the present invention can immediately issue an alarm when abnormal information such as cluster and file system failures, or capacity usage reaching a threshold, and human intervention can be carried out to avoid causing larger production accidents. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings constituting a part of the present invention are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings: Figure 1 A schematic diagram of the GPFS file system monitoring system architecture according to an embodiment of the present invention; Figure 2 The figure is a schematic diagram of the GPFS file system monitoring process described in an embodiment of the present invention. DETAILED DESCRIPTION
[0023] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features in the embodiments may be combined with each other.
[0024] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0025] Currently, GPFS file systems are generally created, configured, mounted, and other operations are performed manually through the GUI interface or CLI commands to achieve daily maintenance and use of GPFS. Similarly, the health of the GPFS cluster and file system, capacity usage, and various dimensions of information such as clients using each file system are all viewed in the GUI interface or obtained by logging into the GPFS cluster node background and executing CLI commands. Human intervention is required, which is not only complex and cumbersome; at the same time, when a fault occurs or some thresholds are reached, an alarm cannot be issued in time, which may cause serious production accidents.
[0026] A middleware-based real-time monitoring solution for GPFS includes the following technical points: 1. Deploy agents on all nodes of the GPFS cluster; 2. The Agent executes CLI commands on the cluster nodes every 2 seconds to complete the collection of a series of required monitoring data; 3. The Agent encapsulates the collected data into a specific message and sends it to the message queue middleware server; 4. The background management system consumes messages from the message queue and stores them in InfluxDB after processing; 5. The backend management system displays the various dimensions of data of the GPFS cluster and file system in real time, and immediately issues an alarm if any abnormal information occurs; The present invention is a solution for real-time monitoring of GPFS based on middleware, comprising the following steps: S01. Deploy the agent on all nodes of the GPFS cluster and start it with the server. S02. At 2s intervals, the agent executes CLI commands on the server to obtain the health status, total capacity, used capacity, available capacity, and client data of the GPFS cluster and a single file system, as well as the data of the clients using the file system. S03, the Agent encapsulates the acquired data into a specific message and sends it to the message queue middleware server; S04. The background management system consumes messages from the message queue; S05. The background management system parses the message and stores it in InfluxDB; S06. The backend management system displays the GPFS cluster and the data of various dimensions of the file system in real time. If any abnormal information is found, an alarm will be issued immediately and human intervention will be carried out. The present invention can enable the GPFS file system monitoring data to be automatically displayed in real time and dynamically without human intervention; at the same time, the entire life cycle of the entire GPFS file system is monitored, which is convenient for subsequent data analysis and further optimization of the usage plan; when a fault occurs or reaches a certain threshold, it can be discovered and handled in time to avoid causing a larger production accident; Those of ordinary skill in the art will appreciate that the units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0027] In the several embodiments provided in the present application, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the division of the units described above is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The above-mentioned units may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present invention.
[0028] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein by equivalents. These modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and specification of the present invention.
[0029] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for real-time monitoring of GPFS based on middleware, characterized in that: include: Deploy a lightweight monitoring agent on all nodes of the GPFS cluster and set the agent to start with the node server; The Agent calls a preset CLI command set at fixed time intervals to collect multi-dimensional data of the GPFS cluster and file system; Encapsulating the data collected by the Agent into a standardized message with a timestamp and a unique identifier; The backend management system consumes the standardized messages in real time through the consumer module of the message queue, and performs format parsing, data verification and preprocessing on the messages; The backend management system stores the processed monitoring data in a database that supports time series query and aggregate analysis; Based on the stored data, the backend management system displays the operating status of the GPFS cluster and file system in real time through the Web interface and generates trend analysis charts.
2. A method for real-time monitoring of GPFS based on middleware according to claim 1, characterized in that: The Agent calls a preset CLI command set at fixed time intervals to collect multi-dimensional data of the GPFS cluster and file system, including cluster health status, total capacity, used capacity, available capacity, number of client accesses, and read and write performance indicators.
3. The method for real-time monitoring of GPFS based on middleware according to claim 1, characterized in that: When an abnormal state is detected, the background management system triggers multi-channel alarms through the built-in alarm module; The multi-channel alarm includes email, SMS, and instant message push, notifying designated personnel to handle; The abnormal state includes capacity usage exceeding a threshold or node disconnection.
4. The method for real-time monitoring of GPFS based on middleware according to claim 1, characterized in that: The CLI command set of the Agent includes preset template files of multiple commands. Users can add new commands by modifying the template files. The template files are stored in the local configuration directory of the Agent. After the Agent reads and parses the template files, it generates an execution command queue to support latency analysis, file system fragmentation rate detection and I / O throughput statistics.
5. The method for real-time monitoring of GPFS based on middleware according to claim 1, characterized in that: The Agent includes a local storage module. When the message queue middleware is unavailable, the Agent stores the collected data in a local cache directory in JSON format and sets a resend mechanism to ensure that the cached data is resent in timestamp order after the message queue is restored.
6. The method for real-time monitoring of GPFS based on middleware according to claim 1, characterized in that: The message queue middleware adopts the Kafka distributed architecture and ensures high availability of data based on partition and copy mechanisms; message producers and consumers manage subscription and distribution processes through a distributed coordinator.
7. The method for real-time monitoring of GPFS based on middleware according to claim 1, characterized in that: The alarm module combines trend analysis of historical data and uses time series algorithm to predict the trend of monitoring indicator data; When the predicted value exceeds the predefined threshold, an early warning message is generated and sent through the configured mail server or message push interface.
8. An electronic device, comprising a processor and a memory connected to the processor for storing instructions executable by the processor, characterized in that: The processor is used to execute the method for real-time monitoring of GPFS based on middleware as described in any one of claims 1-7 above.
9. A server, characterized in that: It includes at least one processor and a memory communicatively connected to the processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor so that the at least one processor executes a method for real-time monitoring of GPFS based on middleware as described in any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for real-time monitoring of GPFS based on middleware described in any one of claims 1 to 7 is implemented.