Disk performance monitoring method and device, equipment and medium
The eBPF program-based method addresses the limitations of traditional disk I/O monitoring by providing real-time, fine-grained monitoring with visualization, enhancing system performance and fault detection.
Patent Information
- Application Number
- CN202510441340.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-15
AI Technical Summary
Traditional disk I/O monitoring tools are difficult to provide fine-grained metric collection, real-time monitoring of data and flexible monitoring strategies, and may have adverse effects on system performance under high load environments.
The EBPF program is used to run the monitoring program in the kernel state, collect disk read and write operation data through the kernel state driver, and analyze and display performance indicators by the user state driver, and monitor it using a visual graphical interface.
It realizes fine-grained disk I/O monitoring and real-time analysis, improves monitoring efficiency and flexibility, helps optimize system performance and locates faults and monitors potential threats without increasing system burden.
Smart Images

Figure CN120315643A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of detection technology, and in particular, to a method, device, equipment and medium for monitoring disk performance. Background Art
[0002] In modern computer systems, disk read and write (I / O) performance has an important impact on the overall system performance, especially in scenarios involving a large number of data read and write operations, such as database management systems, large-scale data processing platforms, and virtualization environments, etc.; traditional disk I / O monitoring tools include iostat, vmstat, sar. Although these monitoring tools can provide system-level disk I / O performance statistical information, most of these disk monitoring tools are implemented through user-space applications and have the following deficiencies:
[0003] Traditional monitoring tools usually can only provide high-level system metrics and are difficult to collect fine-grained metrics for specific processes or specific disk I / O requests; most traditional monitoring tools use a polling mechanism to monitor disks and are difficult to provide real-time I / O monitoring data according to the read and write conditions of the disks; traditional frequent system calls may introduce additional overhead, especially in high-load environments, which may have an adverse impact on system performance; traditional monitoring tools are difficult to flexibly customize monitoring strategies and data collection methods according to specific monitoring requirements. Summary of the Invention
[0004] This application provides a method, device, equipment and medium for monitoring disk performance. The method includes obtaining a kernel-mode driver and a user-mode driver; mounting the kernel-mode driver to a node corresponding to disk read and write operations according to the monitoring requirements for the disk; collecting disk read and write operation data through the kernel-mode driver, where the disk read and write operation data includes the start time and end time of read and write requests, and the number of bytes transferred; the user-mode driver determines disk read and write operation performance metric data according to the disk read and write operation data, and the disk read and write operation performance metric data includes read and write request latency, read and write operation data volume, and read and write operation request size; displaying the disk read and write operation performance metric data through a visual graphical interface, and monitoring the disk according to the display result. This application uses an eBPF program to run a monitoring program in the disk kernel mode, which can capture and analyze disk read and write events in real time at extremely low cost, resulting in high monitoring efficiency and flexibility.
[0005] This application provides a method for monitoring disk performance. The method is applied to a disk performance monitoring system and includes:
[0006] Obtaining a kernel-mode driver and a user-mode driver;
[0007] Mount the kernel-mode driver to the node corresponding to the disk read and write operation according to the monitoring requirements of the disk;
[0008] Collect disk read and write operation data through the kernel-mode driver, where the disk read and write operation data includes the start time and end time of the read and write request, and the number of bytes transferred;
[0009] The user-mode driver determines the disk read and write operation performance metric data according to the disk read and write operation data. The disk read and write operation performance metric data includes the read and write request latency time, the amount of read and write operation data, and the read and write operation request size;
[0010] Display the disk read and write operation performance metric data through a visual graphical interface, and monitor the disk according to the display result.
[0011] This application also provides a disk performance monitoring device, including:
[0012] An acquisition module for acquiring the kernel-mode driver and the user-mode driver;
[0013] A mounting module for mounting the kernel-mode driver to the node corresponding to the disk read and write operation according to the monitoring requirements of the disk;
[0014] A collection module for collecting disk read and write operation data through the kernel-mode driver, where the disk read and write operation data includes the start time and end time of the read and write request, and the number of bytes transferred;
[0015] A determination module for determining the disk read and write operation performance metric data according to the disk read and write operation data. The disk read and write operation performance metric data includes the read and write request latency time, the amount of read and write operation data, and the read and write operation request size;
[0016] A monitoring module for displaying the disk read and write operation performance metric data through a visual graphical interface, and monitoring the disk according to the display result.
[0017] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of the disk performance monitoring method when executing the computer program. The method includes:
[0018] Acquire the kernel-mode driver and the user-mode driver;
[0019] Mount the kernel-mode driver to the node corresponding to the disk read and write operation according to the monitoring requirements of the disk;
[0020] Collect disk read and write operation data through the kernel-mode driver, where the disk read and write operation data includes the start time and end time of the read and write request, and the number of bytes transferred;
[0021] The user-mode driver determines the disk read / write operation performance metric data based on the disk read / write operation data, and the disk read / write operation performance metric data includes read / write request latency time, read / write operation data volume, and read / write operation request size;
[0022] The disk read / write operation performance metric data is displayed through a visual graphical interface, and the disk is monitored according to the display result.
[0023] This application also provides a computer-readable storage medium in which a computer program is stored. When the computer program is executed by a processor, the steps of the disk performance monitoring method are implemented. The method includes:
[0024] Obtain the kernel-mode driver and the user-mode driver;
[0025] Mount the kernel-mode driver to the node corresponding to the disk read / write operation according to the monitoring requirements for the disk;
[0026] Collect disk read / write operation data through the kernel-mode driver, where the disk read / write operation data includes the start time and end time of the read / write request, and the number of bytes transferred;
[0027] The user-mode driver determines the disk read / write operation performance metric data based on the disk read / write operation data, and the disk read / write operation performance metric data includes read / write request latency time, read / write operation data volume, and read / write operation request size;
[0028] The disk read / write operation performance metric data is displayed through a visual graphical interface, and the disk is monitored according to the display result.
[0029] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the disk performance monitoring method are implemented. The method includes:
[0030] Obtain the kernel-mode driver and the user-mode driver;
[0031] Mount the kernel-mode driver to the node corresponding to the disk read / write operation according to the monitoring requirements for the disk;
[0032] Collect disk read / write operation data through the kernel-mode driver, where the disk read / write operation data includes the start time and end time of the read / write request, and the number of bytes transferred;
[0033] The user-mode driver determines the disk read / write operation performance metric data based on the disk read / write operation data, and the disk read / write operation performance metric data includes read / write request latency time, read / write operation data volume, and read / write operation request size;
[0034] The performance index data of disk read and write operations is displayed through a visual graphical interface, and the disk is monitored according to the display results.
[0035] Through the present application, since the method includes obtaining a kernel-mode driver and a user-mode driver; mounting the kernel-mode driver to a node corresponding to disk read and write operations according to the monitoring requirements of the disk; collecting disk read and write operation data through the kernel-mode driver, where the disk read and write operation data includes the start time and end time of read and write requests, and the number of bytes of data transmission; the user-mode driver determines disk read and write operation performance index data according to the disk read and write operation data, and the disk read and write operation performance index data includes read and write request latency, read and write operation data volume, and read and write operation request size; the performance index data of disk read and write operations is displayed through a visual graphical interface, and the disk is monitored according to the display results. Therefore, the present application uses an eBPF program to run a monitoring program in the disk kernel mode, which can capture and analyze disk read and write events in real time at extremely low cost, resulting in high monitoring efficiency and flexibility.
[0036] The technical solution of the present application can provide fine-grained monitoring, real-time analysis, and visual display of disk read and write (I / O) operations through an eBPF program without increasing the system burden, thereby helping users optimize system performance, locate disk failures, and monitor potential security threats.
[0037] The technical solution of the present application can be mounted on different kernel functions or event points according to monitoring requirements, and the monitoring content and scope can be flexibly defined; the technical solution of the present application is implemented in a pure software manner without special hardware requirements, saving development costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is the first flowchart of the disk performance monitoring method provided by the embodiment of the present application;
[0040] Figure 2 It is the architecture diagram of the disk performance monitoring system provided by the embodiment of the present application;
[0041] Figure 3 It is the second flowchart of the disk performance monitoring method provided by the embodiment of the present application;
[0042] Figure 4 It is the structure diagram of the disk performance monitoring device provided by the embodiment of the present application;
[0043] Figure 5 An exemplary system provided for the embodiments of the present application that can be used to implement the various embodiments described in the present application. Detailed implementation manners
[0044] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0045] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0046] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0047] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the disk performance monitoring method depends, the specific application environment architecture or specific hardware architecture is described herein.
[0048] The eBPF program is a sandbox mechanism running in the Linux kernel, which allows user-defined programs to be inserted without modifying the kernel code, realizing real-time monitoring and tracing of kernel behaviors; using eBPF technology, monitoring programs can be mounted on kernel functions or event points related to disk I / O, capturing detailed information of disk I / O in the kernel state in real time and saving the data; after reading the disk I / O data saved in the kernel state in the user state, performing calculations, analysis, and aggregation, and finally docking with the visualization graphical interface to intuitively display the read / write I / O related metrics of the disk.
[0049] It can be understood that the eBPF program: a program written by users that can run in the Linux kernel, which is essentially a virtual machine running in the kernel state, enabling users to load and run custom programs in the kernel, and allowing users and the kernel to interact with data as needed and manage the behaviors of the kernel at any time.
[0050] Disk I / O (Input / Output): Refers to the read and write operations of a computer system on a disk. Operations such as database operations and file reading and writing will generate disk read and write operations.
[0051] Flame graph: A tool graph for visualizing and analyzing performance. Through the flame graph, performance bottlenecks in the system can be intuitively discovered.
[0052] Embodiments of the present application provide a method for monitoring disk performance, as Figure 1 、 2 shown. The method is applied to a disk performance monitoring system and includes:
[0053] Obtain the kernel-mode driver (kernel-mode eBPF program) and the user-mode driver (user-mode eBPF program);
[0054] Mount the kernel-mode driver to the node corresponding to the disk read and write operation according to the monitoring requirements for the disk;
[0055] Collect disk read and write operation data through the kernel-mode driver. Among them, the disk read and write operation data includes the start time and end time of the read and write request, and the number of bytes of data transmission;
[0056] The user-mode driver determines the disk read and write operation performance metric data according to the disk read and write operation data. The disk read and write operation performance metric data includes the read and write request latency time, the amount of read and write operation data, and the read and write operation request size;
[0057] Display the disk read and write operation performance metric data through a visual graphical interface, and monitor the disk according to the display result.
[0058] It can be understood that the present application proposes a method for monitoring disk performance based on the eBPF program. By designing the eBPF program, it is possible to provide fine-grained monitoring, real-time analysis, and visual display of disk I / O operations without significantly increasing the system burden, thereby helping users optimize system performance, locate faults, and monitor potential security threats.
[0059] First, write and design the eBPF program to capture disk I / O events in the kernel; the eBPF program can be mounted on the kernel function entry (kprobe), function return point (kretprobe), or specific kernel event (tracepoint) related to disk I / O to obtain key events such as I / O request submission, processing, and completion; specific monitoring points can be configured according to actual needs to flexibly adapt to different application scenarios.
[0060] Deploy the eBPF program on each node. The eBPF program is responsible for collecting data and storing it in the eBPFMap in the kernel space. The eBPFMap is one of the most important data structures in the eBPF technology, implemented in the kernel state, storing data in the form of key-value, and can perform efficient operations such as querying, inserting, and deleting. Before storing data in the eBPFMap, preliminary filtering and aggregation can be performed on the data to reduce the computational burden of subsequent analysis.
[0061] Secondly, the eBPF program in the user state is responsible for reading the data in the eBPFMap and performing further in-depth analysis, such as calculating I / O latency, throughput, read / write ratio, etc. The analysis calculation results can generate detailed log records for subsequent auditing and analysis.
[0062] Finally, the processed data can be docked with the visualization and graphical interface to display the I / O performance metrics of the disk in a more intuitive way in real time, facilitating the timely discovery of problems caused by disk I / O performance and taking corresponding measures.
[0063] An embodiment of the present application provides a disk performance monitoring method, which is applied to a disk performance monitoring system, such as Figure 3 as shown, the method includes:
[0064] Step S01, obtain the kernel state driver and the user state driver;
[0065] Step S02, mount the kernel state driver to the node corresponding to the disk read and write operation according to the monitoring requirements of the disk.
[0066] Specifically, write a kernel state eBPF program and mount it at the kernel function entry (kprobe) or function return point (kretprobe) or specific kernel event (tracepoint) related to disk I / O.
[0067] Step S021, in response to monitoring the latency time and / or throughput and / or read / write request size data of the disk read and write operation, mount the kernel state driver to the kernel function entry corresponding to the disk read and write operation;
[0068] In response to monitoring the completion time and / or read / write request latency time of each disk read and write operation request, mount the kernel state driver to the kernel function return point corresponding to the disk read and write operation;
[0069] In response to monitoring the read and write events and / or flush events of the disk, mount the kernel state driver to the target kernel event corresponding to the disk read and write operation.
[0070] Specifically, when it is necessary to monitor in real time metrics such as the latency, throughput, and request size of disk I / O, the kernel-mode eBPF program is attached to the kernel functions related to disk I / O.
[0071] When it is necessary to measure the completion time of each disk I / O request and calculate the latency, the kernel-mode eBPF program is attached to the return point of the kernel function of disk I / O.
[0072] When it is necessary to monitor specific disk I / O events (such as read, write, flush, discard, etc.), the kernel-mode eBPF program is attached to specific events in block_tracepoints.
[0073] Step S03, collect disk read and write operation data through the kernel-mode driver. Among them, the disk read and write operation data includes the start time and end time of the read and write requests, and the number of bytes of data transfer.
[0074] Specifically, when the kernel-mode eBPF program is running, it will automatically collect the following key data when a specified kernel event is triggered: the start time and end time of the I / O request, the I / O operation type (read / write), the target device (disk) information, the number of bytes of data transfer, the process ID, and the thread ID, etc.
[0075] Step S04, filter and aggregate the disk read and write operation data;
[0076] Store the filtered and aggregated disk read and write operation data in the kernel-mode driver Map;
[0077] Filtering and aggregating the disk read and write operation data includes:
[0078] Judge whether the latency time of the disk read and write operation is greater than the first threshold (1 - 2s);
[0079] If so, retain the data with the latency time of the disk read and write operation greater than the first threshold; if not, delete the disk read and write operation data;
[0080] Aggregate the filtered disk read and write operation data according to the process ID and / or the target disk information.
[0081] Specifically, after collecting the data, it will perform preliminary filtering and aggregation on the data, only retain the data with the I / O latency exceeding the first threshold, aggregate the filtered disk read and write operation data according to the process ID or the target disk to reduce the data volume and improve the efficiency of subsequent data analysis, and write the filtered and aggregated data into the eBPFMap.
[0082] Step S05: The user-mode driver determines the disk read / write operation performance metric data based on the disk read / write operation data. The disk read / write operation performance metric data includes read / write request latency, read / write operation data volume, and read / write operation request size.
[0083] Specifically, the user-mode eBPF program is responsible for reading the data in the eBPFMap, calculating various disk I / O performance metrics, and recording the analysis and calculation results in the form of logs. The metrics include I / O latency, I / O read / write data volume, I / O request count, etc.
[0084] Step S051: Calculate the latency corresponding to the disk read / write request based on the start time and end time of the disk read / write request;
[0085] Calculate the disk read / write operation data volume based on the number of bytes transferred;
[0086] Calculate the disk read / write operation data volume based on the number of bytes transferred, including:
[0087] Obtain the total number of read bytes and the total number of written bytes through the kernel-mode driver;
[0088] Take the sum of the total number of read bytes and the total number of written bytes as the total disk read / write operation data volume;
[0089] Calculate the disk read / write operation request size based on the number of bytes transferred;
[0090] Calculate the disk read / write operation request size based on the number of bytes transferred, including:
[0091] Obtain the total number of read bytes, the total number of written bytes, the number of read requests, and the number of write requests through the kernel-mode driver;
[0092] Calculate the disk read operation request size through the formula: average read request size = total number of read bytes / number of read requests;
[0093] Calculate the disk write operation request size through the formula: average write request size = total number of written bytes / number of write requests.
[0094] Step S06: Convert the read / write request latency, read / write operation data volume, and read / write request size data into a log form and store it;
[0095] Aggregate the disk read / write operation performance metric data according to the disk type and process ID.
[0096] Specifically, finally aggregate the above data according to the disk type and PID (process), as shown in Table 1 below:
[0097] Table 1
[0098]
[0099]
[0100] As shown in the above example, it shows the index data such as disks named sr0 and sr1, the I / O read and write data volume, the number of I / O requests, and PID at certain intervals of time, I / O latency.
[0101] Step S07, display the performance index data of disk read and write operations through a visual graphical interface, and monitor the disk according to the display results.
[0102] Specifically, by docking with the visual graphical interface, the I / O metrics of the system disk can be intuitively and real-time displayed.
[0103] Step S071, display the trend of changes in disk read and write operation latency time and / or throughput index data over time through a time series graph;
[0104] Display the proportion of different types of disk read and write operation data through a pie chart;
[0105] Display the request density of disk read and write operation data within a unit time period through a flame graph.
[0106] Specifically, display the trend of changes in metrics such as I / O latency and throughput over time through a time series graph; display the proportion of different types of disk I / O operations, such as the proportion of read / write operations, through a pie chart; display the I / O operation density within a specific time period through a flame graph.
[0107] At the same time, the visual graphical interface increases the custom display time range, data type, and filtering conditions. After the above graphical interface display, the I / O performance status of the current system disk can be understood more quickly, potential problems can be discovered in a timely manner, and corresponding measures can be taken.
[0108] Step S08, optimize the disk according to the disk HDD performance monitoring problem;
[0109] Optimize the disk according to the disk performance monitoring problem, including:
[0110] When the disk utilization rate approaches the second threshold (100%), upgrade the disk and / or configure a redundant array for the disk and / or disperse the disk read and write operation load;
[0111] When the latency time of disk read and write operations is greater than the third threshold (20 - 50 ms), optimize the disk read and write operation scheduler and / or upgrade the disk hardware and / or reduce the read and write operations of small files;
[0112] When the number of read / write requests of the disk HDD per unit time is greater than the fourth threshold (400 - 500 times), the disk read / write request load is dispersed and / or disk caching is used and / or the disk read / write operation mode of the application is optimized;
[0113] When the backlog depth of the disk read / write requests in the queue is greater than the fifth threshold (10), the scheduler for the disk read / write operations is optimized and / or the number of disks is increased and / or the disk read / write operation load is dispersed;
[0114] When the time for the disk to process each read / write request is greater than the sixth threshold (20 ms), the disk is upgraded and / or the disk file system is optimized and / or the disk kernel parameters are adjusted.
[0115] Here, in the present application, by designing an eBPF program to obtain disk I / O related operation metrics in the kernel state, and processing and analyzing the metrics in the user - state eBPF program and finally presenting them through a graphical interface, it is possible to provide fine - grained monitoring and real - time analysis of disk I / O operations with low overhead.
[0116] In addition, according to the monitoring requirements of the disk, mounting the kernel - state driver to the node corresponding to the disk read / write operation further includes:
[0117] When it is necessary to monitor the error requests of the disk, the kernel - state eBPF program is mounted to the block_rq_complete event (a tracepoint in the kernel) to diagnose the disk health;
[0118] When it is necessary to analyze the queue depth of the disk, the kernel - state eBPF program is mounted to the blk_mq_start_request function (a function on the multi - queue read / write path) to balance the disk load and optimize the scheduling;
[0119] When it is necessary to trace the blocking read / write model of the disk, the kernel - state eBPF program is mounted to the block_bio_queue (request queue) to analyze the blocking data structure and monitor the disk file system layer;
[0120] When it is necessary to associate the process - level read / write operations of the disk, the kernel - state eBPF program is mounted to the kprobe / __blk_account_io_start address to analyze the disk application behavior by obtaining the process ID of the current process and the disk read / write parameters.
[0121] Through the eBPF program, low - overhead and high - precision disk monitoring can be achieved, which is applicable to the full - link observation from development and debugging to the production environment.
[0122] Without departing from the technical solution of the present application, several improvements and optimizations can be made to the method for monitoring disk performance provided by the embodiments of the present application, and these improvements and optimizations should also be regarded as the protection scope of the present application.
[0123] The beneficial effects brought by the technical solution provided by the embodiments of the present application are as follows:
[0124] The present application uses an eBPF program to run a monitoring program in the disk kernel state, which can capture and analyze disk read and write events in real time at extremely low cost, making the monitoring efficiency and flexibility high.
[0125] The technical solution of the present application can provide fine-grained monitoring, real-time analysis, and visualization display of disk read and write (I / O) operations through the eBPF program without increasing the system burden, thereby helping users optimize system performance, locate disk failures, and monitor potential security threats.
[0126] The technical solution of the present application can be mounted on different kernel functions or event points according to monitoring requirements, and the monitoring content and scope can be flexibly defined; the technical solution of the present application is implemented in a pure software manner without special hardware requirements, saving development costs.
[0127] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner.
[0128] The embodiments of the present application also provide a disk performance monitoring device, as Figure 4 shown, the device includes: an acquisition module, a mounting module, a collection module, a filtering module, a determination module, an aggregation module, a monitoring module, and a post-processing module.
[0129] In this embodiment, the acquisition module is used to acquire the kernel-state driver and the user-state driver;
[0130] The mounting module is used to mount the kernel-state driver to the node corresponding to the disk read and write operation according to the monitoring requirements of the disk;
[0131] The collection module is used to collect disk read and write operation data through the kernel-state driver, where the disk read and write operation data includes the start time and end time of the read and write requests, and the number of bytes of data transmission;
[0132] The determination module is used to determine the disk read and write operation performance index data according to the disk read and write operation data, and the disk read and write operation performance index data includes the read and write request latency time, the read and write operation data volume, and the read and write operation request size;
[0133] The monitoring module is used to display the performance index data of disk read and write operations through a visual graphical interface, and monitor the disk according to the display results.
[0134] In this embodiment, the mounting module is used to mount the kernel-mode driver to the kernel function entry corresponding to the disk read and write operation in response to monitoring the latency time and / or throughput and / or read and write request size data of the disk read and write operation;
[0135] In response to monitoring the completion time and / or read and write request latency time of each read and write operation request of the disk, mount the kernel-mode driver to the kernel function return point corresponding to the disk read and write operation;
[0136] In response to monitoring the read and write events and / or refresh events of the disk, mount the kernel-mode driver to the target kernel event corresponding to the disk read and write operation.
[0137] In one embodiment, the determination module is used to calculate the latency time corresponding to the disk read and write request according to the start time and end time of the disk read and write request;
[0138] Calculate the disk read and write operation data volume according to the number of bytes of data transmission;
[0139] Calculating the disk read and write operation data volume according to the number of bytes of data transmission includes:
[0140] Obtain the total number of read bytes and the total number of written bytes through the kernel-mode driver;
[0141] Take the sum of the total number of read bytes and the total number of written bytes as the total disk read and write operation data volume;
[0142] Calculate the disk read and write operation request size according to the number of bytes of data transmission;
[0143] Calculating the disk read and write operation request size according to the number of bytes of data transmission includes:
[0144] Obtain the total number of read bytes, the total number of written bytes, the number of read requests, and the number of write requests through the kernel-mode driver;
[0145] Calculate the disk read operation request size through the formula: average read request size = total number of read bytes / number of read requests;
[0146] Calculate the disk write operation request size through the formula: average write request size = total number of written bytes / number of write requests.
[0147] In one embodiment, the filtering module is used to filter and aggregate the disk read and write operation data;
[0148] Store the filtered and aggregated disk read and write operation data in the kernel driver Map;
[0149] Filter and aggregate the disk read and write operation data, including:
[0150] Judge whether the latency time of the disk read and write operation is greater than the first threshold;
[0151] If so, retain the data with the latency time of the disk read and write operation greater than the first threshold; if not, delete the disk read and write operation data;
[0152] Aggregate the filtered disk read and write operation data according to the process ID and / or target disk information.
[0153] In one embodiment, the aggregation module is used to convert the read and write request latency time, read and write operation data volume, and read and write request size data into a log form and store them;
[0154] Aggregate the disk read and write operation performance metric data according to the disk type and process ID.
[0155] In one embodiment, the monitoring module is used to display the trend of the disk read and write operation latency time and / or throughput metric data changing with time through a time series graph;
[0156] Display the occupancy ratio of different types of disk read and write operation data through a pie chart;
[0157] Display the request density of the disk read and write operation data within a unit time period through a flame graph.
[0158] In one embodiment, the post-processing module is used to optimize the disk according to the disk performance monitoring problems;
[0159] Optimize the disk according to the disk performance monitoring problems, including:
[0160] When the disk utilization rate is close to the second threshold, upgrade the disk and / or configure a redundant array for the disk and / or disperse the disk read and write operation load;
[0161] When the latency time of the disk read and write operation is greater than the third threshold, optimize the disk read and write operation scheduler and / or upgrade the disk hardware and / or reduce the read and write operations of small files;
[0162] When the number of read and write requests per unit time of the disk is greater than the fourth threshold, disperse the disk read and write request load and / or use disk caching and / or optimize the application disk read and write operation mode;
[0163] When the backlog depth of disk read / write requests in the queue is greater than the fifth threshold, optimize the scheduler for disk read / write operations and / or increase the number of disks and / or disperse the load of disk read / write operations;
[0164] When the time for the disk to process each read / write request is greater than the sixth threshold, upgrade the disk and / or optimize the disk file system and / or adjust the disk kernel parameters.
[0165] Specifically, this application captures disk I / O events in the kernel through an eBPF program, writes the collected data into an eBPF map after preprocessing filtering and aggregation in the kernel state, reads the data in the eBPF map in the user-state eBPF program and analyzes it, calculates key metric data reflecting disk I / O performance such as I / O latency, I / O read / write data volume, I / O request count, etc., and conducts data docking with the graphical interface to dynamically and real-time display the I / O performance metrics of the system disk, mainly including the following steps:
[0166] The kernel-state eBPF program is mounted in kernel functions related to disk I / O, such as kprobe, tracepoint, etc. When there is a disk I / O operation, relevant data related to disk I / O is captured when a specified kernel event is triggered, and the data is filtered and aggregated and then written into the eBPF map;
[0167] The user-state eBPF program reads the data in the eBPF map in real-time and calculates, analyzes, and aggregates the data to finally obtain key metrics related to disk I / O;
[0168] Dock the key metrics related to disk I / O with the visualization graphical interface, display the obtained disk I / O key metric data in the form of a time series graph, pie chart, and flame graph, and the data display time range, data type, and filtering conditions can be customized to timely discover potential problems and take countermeasures.
[0169] The beneficial effects brought by the technical solution provided by the embodiments of this application are:
[0170] This application uses an eBPF program to run a monitoring program in the disk kernel state, which can capture and analyze disk read / write events in real-time at extremely low cost, with high monitoring efficiency and flexibility.
[0171] The technical solution of this application can, through the eBPF program, provide fine-grained monitoring, real-time analysis, and visualization display of disk read / write (I / O) operations without increasing the system burden, thereby helping users optimize system performance, locate disk failures, and monitor potential security threats.
[0172] The technical solution of this application can be mounted on different kernel functions or event points according to monitoring requirements, and the monitoring content and scope can be flexibly defined; the technical solution of this application is implemented in a pure software manner without special hardware requirements, saving development costs.
[0173] For the description of the features in the corresponding embodiment of the disk performance monitoring device, reference can be made to the relevant description in the corresponding embodiment of the disk performance monitoring method, which will not be elaborated here one by one.
[0174] An embodiment of this application also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in the embodiment of the disk performance monitoring method. The method includes:
[0175] Obtain the kernel-mode driver and the user-mode driver;
[0176] Mount the kernel-mode driver to the node corresponding to the disk read and write operation according to the monitoring requirements for the disk;
[0177] Collect disk read and write operation data through the kernel-mode driver. Among them, the disk read and write operation data includes the start time and end time of the read and write requests, and the number of bytes of data transmission;
[0178] The user-mode driver determines the disk read and write operation performance index data according to the disk read and write operation data. The disk read and write operation performance index data includes the read and write request latency time, the amount of read and write operation data, and the size of the read and write operation requests;
[0179] Display the disk read and write operation performance index data through a visual graphical interface, and monitor the disk according to the display result.
[0180] As Figure 5 shown, an embodiment of this application also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. Among them, the computer program is configured to execute the steps in the embodiment of the disk performance monitoring method when running. The method includes:
[0181] Obtain the kernel-mode driver and the user-mode driver;
[0182] Mount the kernel-mode driver to the node corresponding to the disk read and write operation according to the monitoring requirements for the disk;
[0183] Collect disk read and write operation data through the kernel-mode driver. Among them, the disk read and write operation data includes the start time and end time of the read and write requests, and the number of bytes of data transmission;
[0184] The user-mode driver determines the disk read / write operation performance metric data based on the disk read / write operation data. The disk read / write operation performance metric data includes read / write request latency, read / write operation data volume, and read / write operation request size.
[0185] The disk read / write operation performance metric data is displayed through a visual graphical interface, and the disk is monitored according to the display result.
[0186] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: USB flash drive, read-only memory (ROM for short), random access memory (RAM for short), external hard disk, magnetic disk, or optical disc, etc., various media that can store computer programs.
[0187] An embodiment of the present application also provides a computer program product. The computer program product includes a computer program. When the computer program is executed by a processor, the steps in the embodiment of the disk performance monitoring method are implemented. The method includes:
[0188] Obtain the kernel-mode driver and the user-mode driver;
[0189] Mount the kernel-mode driver to the node corresponding to the disk read / write operation according to the monitoring requirements for the disk;
[0190] Collect disk read / write operation data through the kernel-mode driver. Among them, the disk read / write operation data includes the start time and end time of the read / write request, and the number of bytes of data transmission;
[0191] The user-mode driver determines the disk read / write operation performance metric data based on the disk read / write operation data. The disk read / write operation performance metric data includes read / write request latency, read / write operation data volume, and read / write operation request size;
[0192] The disk read / write operation performance metric data is displayed through a visual graphical interface, and the disk is monitored according to the display result.
[0193] Another embodiment of the present application also provides a computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the steps in the embodiment of the disk performance monitoring method are implemented. The method includes:
[0194] Obtain the kernel-mode driver and the user-mode driver;
[0195] Mount the kernel-mode driver to the node corresponding to the disk read / write operation according to the monitoring requirements for the disk;
[0196] Collect disk read and write operation data through a kernel-mode driver, where the disk read and write operation data includes the start time and end time of the read and write requests, and the number of bytes transferred during data transmission;
[0197] The user-mode driver determines the disk read and write operation performance metric data based on the disk read and write operation data. The disk read and write operation performance metric data includes the read and write request latency time, the amount of read and write operation data, and the size of the read and write operation requests;
[0198] Display the disk read and write operation performance metric data through a visual graphical interface, and monitor the disk according to the display results.
[0199] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0200] The above has introduced in detail a disk performance monitoring method, device, equipment, and medium provided by this application. Specific examples have been used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can still be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for monitoring disk performance, characterized in that, The method is applied to a disk performance monitoring system, and the method includes: Obtain a kernel-mode driver and a user-mode driver; Mount the kernel-mode driver to a node corresponding to the disk read / write operation according to the monitoring requirements for the disk; collect disk read / write operation data through the kernel-mode driver, where the disk read / write operation data includes the start time and end time of a read / write request, and the number of bytes of data transmission; The user-mode driver determines disk read / write operation performance metric data according to the disk read / write operation data, where the disk read / write operation performance metric data includes read / write request latency, read / write operation data volume, and read / write operation request size; display the disk read / write operation performance metric data through a visualization graphical interface, and monitor the disk according to the display result.
2. The disk performance monitoring method according to claim 1, characterized in that, The mounting of the kernel-mode driver to a node corresponding to the disk read / write operation according to the monitoring requirements for the disk includes: In response to monitoring the latency time and / or throughput and / or read / write request size data of the disk read / write operation, mount the kernel-mode driver to the kernel function entry corresponding to the disk read / write operation; In response to monitoring the completion time and / or read / write request latency of each read / write operation request of the disk, mount the kernel-mode driver to the kernel function return point corresponding to the disk read / write operation; In response to monitoring the read / write event and / or refresh event of the disk, mount the kernel-mode driver to the target kernel event corresponding to the disk read / write operation.
3. The disk performance monitoring method according to claim 1, wherein, The determining of the disk read / write operation performance metric data according to the disk read / write operation data includes: Calculate the latency time corresponding to the disk read / write request according to the start time and end time of the disk read / write request; Calculate the disk read / write operation data volume according to the number of bytes of data transmission; The calculating of the disk read / write operation data volume according to the number of bytes of data transmission includes: Obtain the total number of read bytes and the total number of written bytes through the kernel-mode driver; Use the sum of the total number of read bytes and the total number of written bytes as the total disk read / write operation data volume; Calculate the disk read / write operation request size according to the number of bytes of data transmission; The calculating of the disk read / write operation request size according to the number of bytes of data transmission includes: Obtain the total number of read bytes, the total number of written bytes, the number of read requests, and the number of write requests through the kernel-mode driver; Calculate the disk read operation request size through the formula: average read request size = total number of read bytes / number of read requests; Calculate the disk write operation request size through the formula: average write request size = total number of written bytes / number of write requests.
4. The disk performance monitoring method according to claim 1, characterized in that, After collecting the disk read / write operation data through the kernel-mode driver, it includes: Filter and aggregate the disk read / write operation data; The filtering and aggregating of the disk read / write operation data includes: Judge whether the latency time of the disk read / write operation is greater than a first threshold; If so, retain the data of the disk read / write operation with a latency greater than the first threshold; if not, delete the disk read / write operation data; Aggregate the filtered disk read / write operation data according to the process ID and / or target disk information; Store the filtered and aggregated disk read / write operation data in the kernel-mode driver Map.
5. The disk performance monitoring method according to claim 1, wherein After determining the disk read / write operation performance metric data based on the disk read / write operation data, it includes: Convert the read / write request latency, read / write operation data volume, and read / write request size data into a log format and store them; aggregate the disk read / write operation performance metric data according to the disk type and process ID.
6. The disk performance monitoring method according to claim 1, wherein Display the disk read / write operation performance metric data through a visual graphical interface and monitor the disk according to the display result, including: display the trend of the disk read / write operation latency and / or throughput metric data changing over time through a time series graph; display the proportion of different types of disk read / write operation data through a pie chart; Display the request density of the disk read / write operation data within a unit time period through a flame graph.
7. The disk performance monitoring method according to claim 1, wherein After displaying the disk read / write operation performance metric data through a visual graphical interface and monitoring the disk according to the display result, it includes: Optimize the disk according to the disk performance monitoring problem; The optimizing the disk according to the disk performance monitoring problem includes: When the disk utilization rate approaches the second threshold, upgrade the disk and / or configure a redundant array for the disk and / or disperse the disk read / write operation load; When the latency of the disk read / write operation is greater than the third threshold, optimize the disk read / write operation scheduler and / or upgrade the disk hardware and / or reduce the read / write operations of small files; When the number of read / write requests per unit time of the disk is greater than the fourth threshold, disperse the disk read / write request load and / or use disk caching and / or optimize the disk read / write operation mode; When the backlog depth of the disk read / write requests in the queue is greater than the fifth threshold, optimize the disk read / write operation scheduler and / or increase the number of disks and / or disperse the disk read / write operation load; When the time for the disk to process read / write requests is greater than the sixth threshold, upgrade the disk and / or optimize the disk file system and / or adjust the disk kernel parameters.
8. A disk performance monitoring device, characterized in that, The device includes: An acquisition module, configured to acquire a kernel-mode driver and a user-mode driver; A mounting module, configured to mount the kernel-mode driver to the node corresponding to the disk read / write operation according to the monitoring requirement of the disk; An acquisition module, configured to acquire disk read / write operation data through the kernel-mode driver, where the disk read / write operation data includes the start time and end time of the read / write request, and the number of bytes of data transmission; A determination module, configured to determine disk read / write operation performance index data according to the disk read / write operation data, where the disk read / write operation performance index data includes read / write request latency, read / write operation data volume, and read / write operation request size; A monitoring module, configured to display the disk read / write operation performance index data through a visual graphical interface, and monitor the disk according to the display result.
9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the disk performance monitoring method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program implements the steps of the disk performance monitoring method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Cited By
Process-level power consumption analysis method and device, electronic equipment and storage medium
CN121412076A
NVMe / TCP full-link performance data acquisition and processing method and system
CN122240383A