EBPF-based database performance optimization method, apparatus and device, and medium
By deploying eBPF probes in the database kernel, obtaining and analyzing performance data, establishing traffic and session ID mapping, and combining machine learning models, the performance loss and positioning difficulties of existing tools are solved, and efficient database performance optimization and intelligent diagnosis are achieved.
Patent Information
- Application Number
- CN202510706031.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-29
AI Technical Summary
Existing database performance monitoring tools rely on APM, resulting in performance loss and stability risks, and cannot effectively correlate application layer SQL with underlying I/O latency, making it difficult to quickly and accurately locate the root cause of slow query problems.
Use the dynamic detection mechanism of eBPF to deploy probes on the critical path of the database kernel, obtain process data, connection data, SQL execution data, disk data, lock data and network data, establish a mapping relationship between network traffic and database session ID, combine machine learning models and root cause positioning rules engine, identify abnormal connections, analyze lock and disk access, and generate optimization suggestions.
It realizes fast and efficient database performance problem analysis, reduces manual intervention, improves system automation level, can accurately identify abnormal connections, warnings for potential risks in advance, and optimizes database performance.
Smart Images

Figure CN120560976A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of database technology, and in particular to an eBPF-based database performance optimization method, device, equipment, and medium. Background Art
[0002] Currently, existing database performance monitoring tools primarily rely on APM (Application Performance Management) code instrumentation or log analysis, requiring modifications to database configuration or application code, which can easily lead to performance degradation and stability risks. Furthermore, they fail to effectively correlate application-layer SQL with underlying I / O latency, making it difficult to quickly and accurately pinpoint the root cause of slow queries. For example, when a Redis (Remote Dictionary Server) GET operation slows down, traditional methods are unable to distinguish whether the cause is network protocol stack latency, storage engine lock contention, or other factors. This increases the time and difficulty of troubleshooting, impacting the overall performance and stability of the system.
[0003] As can be seen from the above, how to quickly and efficiently analyze and optimize database performance issues while minimizing the impact on the running database is an urgent problem that needs to be solved. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a database performance optimization method, apparatus, device, and medium based on eBPF, which can quickly and efficiently analyze and optimize database performance issues while minimizing the impact on the running database. The specific solution is as follows:
[0005] In a first aspect, the present application provides a database performance optimization method based on eBPF, comprising:
[0006] Probes are deployed on the critical path of the database kernel using the dynamic detection mechanism of eBPF to obtain various performance data of the database during operation based on the probes, the performance data is passed to user mode, and a mapping relationship between network traffic and database session ID is established using a first preset data structure of the eBPF; the performance data includes process data, connection data, SQL execution data, disk data, lock data, and network data;
[0007] Identifying abnormal connection conditions of the database based on the mapping relationship and the performance data, and analyzing lock and disk access conditions of the database during operation to obtain corresponding performance issues;
[0008] Based on the performance data, parameter indicators including the number of lock waits per second, I / O queue depth and TCP retransmission rate are determined, and the parameter indicators are input into the machine learning model, and combined with the root cause location rule engine to generate the root cause and optimization suggestions corresponding to the performance problem.
[0009] Optionally, the dynamic detection mechanism of eBPF is used to deploy probes on the critical path of the database kernel to obtain various performance data of the database during operation based on the probes, including:
[0010] Use eBPF's uprobe mechanism to obtain process data including the database process PID;
[0011] Inserting a probe into the network connection detection function of the database kernel based on the uprobe mechanism to obtain connection data including thread PID, connection timestamp and authentication status;
[0012] Using the uprobe mechanism, probes are added to the command scheduling function of the database kernel and each link function in the SQL execution process to extract SQL execution data including SQL text, execution session ID and execution duration information of each link function;
[0013] The storage engine's disk flushing is monitored based on the uprobe mechanism to obtain the disk call frequency, and the eBPF kprobe mechanism is used to insert probes into the preset read and write functions of the database kernel virtual file system to obtain I / O operation information. Disk data is obtained based on the disk call frequency and the I / O operation information.
[0014] Using the uprobe mechanism, a probe is inserted into the lock management function of the storage engine to obtain lock data including lock wait time, transaction ID, and thread context;
[0015] Based on the kprobe mechanism, a probe is inserted into the data transmission function of the transmission control protocol layer in the operating system to obtain network traffic, and the network traffic is filtered through the database process PID to obtain network data.
[0016] Optionally, the identifying an abnormal connection condition of the database based on the mapping relationship and the performance data includes:
[0017] The first client IP address that has been authenticated is stored in a preset whitelist, and the second client IP address that has launched attacks or has frequent illegal operations is stored in a preset blacklist;
[0018] If the IP address to be detected is not in the preset whitelist and the preset blacklist, the number of connections of the IP address to be detected within a preset time period is counted, and it is determined whether the number of connections is greater than the target connection threshold;
[0019] If the number of connections is greater than the target connection threshold, it indicates that the IP address to be detected has an abnormal connection situation, and the IP address to be detected is stored in the preset blacklist, and then the performance data is used to determine whether the abnormal connection situation will cause thread accumulation;
[0020] If the abnormal connection situation will cause thread accumulation, the preset fuse mechanism is used to terminate the abnormal connection.
[0021] Optionally, analyzing the access to locks and disks during operation of the database to identify corresponding performance issues includes:
[0022] aggregating the lock data using a second preset data structure of the eBPF to generate a lock wait relationship graph, and determining whether the lock wait in the lock wait relationship graph exceeds a target lock wait threshold;
[0023] If the lock wait in the lock wait relationship graph exceeds the target lock wait threshold, a database query statement is used to determine a target position corresponding to the target lock wait that exceeds the target lock wait threshold to obtain a corresponding performance problem.
[0024] Optionally, analyzing the access to locks and disks during operation of the database to identify corresponding performance issues includes:
[0025] Determining file information corresponding to a target file involved in a disk read / write operation based on the disk data, and recording a single I / O data volume and a file offset during the disk access process;
[0026] Performing data aggregation dimension statistics on the single I / O data volume and the file offset to obtain the read and write volume per second;
[0027] A performance problem is determined based on the single I / O data volume, the file offset, and the read and write volume per second.
[0028] Optionally, identifying abnormal connection conditions of the database based on the mapping relationship and the performance data, and analyzing lock and disk access conditions of the database during operation to obtain corresponding performance issues, includes:
[0029] Writing the performance data into Prometheus, and integrating the performance data based on Prometheus to construct a data view reflecting the database operation status;
[0030] Generate a monitoring panel based on a preset visualization panel and the data view, and locate target hotspot functions and corresponding function call relationships based on the monitoring panel and using the Huoyanshan tool;
[0031] Corresponding performance issues are determined based on the target hotspot function and the function call relationship.
[0032] Optionally, determining parameter indicators including the number of lock waits per second, I / O queue depth, and TCP retransmission rate based on the performance data, inputting the parameter indicators into a machine learning model, and generating a root cause and optimization suggestion corresponding to the performance issue in combination with a root cause location rule engine, including:
[0033] Based on the performance data and using a preset statistical method, statistically analyze parameter indicators including the number of lock waits per second, I / O queue depth, and TCP retransmission rate, and determine the average values corresponding to the parameter indicators within a preset sliding window to obtain a time window sliding mean;
[0034] A priority strategy is obtained using a machine learning model and a root cause location rule engine, the time window sliding mean is input into the machine learning model, and the root cause and optimization suggestions corresponding to the performance problem are generated based on the priority strategy and the decision tree.
[0035] In a second aspect, the present application provides a database performance optimization device based on eBPF, comprising:
[0036] A performance data acquisition module is configured to deploy probes on the critical path of the database kernel using the dynamic detection mechanism of eBPF, obtain various performance data of the database during operation based on the probes, pass the performance data to user mode, and establish a mapping relationship between network traffic and database session ID using a first preset data structure of eBPF; the performance data includes process data, connection data, SQL execution data, disk data, lock data, and network data;
[0037] An access situation analysis module, configured to identify abnormal connection situations of the database based on the mapping relationship and the performance data, and analyze the access situations of the locks and disks of the database during operation to identify corresponding performance issues;
[0038] A parameter indicator determination module is used to determine parameter indicators including the number of lock waits per second, I / O queue depth and TCP retransmission rate based on the performance data, input the parameter indicators into the machine learning model, and combine with the root cause location rule engine to generate the root cause and optimization suggestions corresponding to the performance problem.
[0039] In a third aspect, the present application provides an electronic device, comprising:
[0040] Memory, used to store computer programs;
[0041] A processor is used to execute the computer program to implement the aforementioned eBPF-based database performance optimization method.
[0042] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program implements the aforementioned eBPF-based database performance optimization method when executed by a processor.
[0043] This application uses the dynamic detection mechanism of eBPF to deploy probes on the critical path of the database kernel, so as to obtain various performance data of the database during operation based on the probes, pass the performance data to the user state, and use the first preset data structure of the eBPF to establish a mapping relationship between network traffic and database session ID; the performance data includes process data, connection data, SQL execution data, disk data, lock data and network data; based on the mapping relationship and the performance data, the abnormal connection status of the database is identified, and the access status of the locks and disks of the database during operation is analyzed to obtain corresponding performance problems; based on the performance data, parameter indicators including the number of lock waits per second, I / O queue depth and TCP retransmission rate are determined, and the parameter indicators are input into the machine learning model, and combined with the root cause location rule engine to generate the root cause and optimization suggestions corresponding to the performance problem.
[0044] As can be seen from the above, this application uses the dynamic detection mechanism of eBPF to obtain various types of performance data in the database kernel, and then can quickly locate the specific session based on the mapping relationship between the network traffic and the database session ID. Based on the mapping relationship and comprehensive performance data, it can accurately identify abnormal connections in the database. Early warning of potential security risks or performance risks to avoid abnormal connections from causing greater impact on the database. In this way, through parameter indicators such as the number of lock waits per second, I / O queue depth, and TCP retransmission rate, the performance status of the database in terms of locks, disks, and networks is intuitively reflected, so that database performance problems can be quantitatively evaluated, providing a clear basis for judging the severity of performance problems, and using machine learning models and root cause location rule engines to analyze performance data, intelligent diagnosis of performance problems and automatic generation of optimization suggestions are achieved, which reduces manual intervention and improves the level of system automation. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0046] Figure 1 This is a flowchart of a database performance optimization method based on eBPF disclosed in this application;
[0047] Figure 2 A client address exception handling flow chart provided for this application;
[0048] Figure 3 This is a specific eBPF-based database performance optimization flowchart disclosed in this application;
[0049] Figure 4 A database performance problem analysis flow chart provided for this application;
[0050] Figure 5 A security audit flowchart provided for this application;
[0051] Figure 6 This is a schematic diagram of the structure of a database performance optimization device based on eBPF disclosed in this application;
[0052] Figure 7 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0054] At present, existing database performance monitoring tools mainly rely on APM, code instrumentation or log analysis, and require modification of database configuration or application code, which can easily lead to performance loss and stability risks; and cannot effectively associate application layer SQL with underlying I / O delays, making it difficult to quickly and accurately locate the root cause of the problem when slow queries occur. To this end, the present application provides a database performance optimization method based on eBPF, which intuitively reflects the performance status of the database in terms of locks, disks, and networks through parameter indicators such as the number of lock waits per second, I / O queue depth, and TCP retransmission rate, thereby quantitatively evaluating database performance problems and providing a clear basis for judging the severity of performance problems. It also uses machine learning models and root cause location rule engines to analyze performance data, realizing intelligent diagnosis of performance problems and automatic generation of optimization suggestions, reducing manual intervention, and improving the level of system automation.
[0055] See also Figure 1 As shown, the embodiment of the present invention discloses a database performance optimization method based on eBPF, including:
[0056] Step S11: Use the dynamic detection mechanism of eBPF to deploy probes on the critical path of the database kernel to obtain various performance data of the database during operation based on the probes, pass the performance data to the user state, and use the first preset data structure of the eBPF to establish a mapping relationship between network traffic and database session ID; the performance data includes process data, connection data, SQL execution data, disk data, lock data and network data.
[0057] In this embodiment, the uprobe mechanism of eBPF (Extended Berkeley Packet Filter) is used to obtain process data including the database process PID for subsequent system function call filtering operations; based on the uprobe mechanism, a probe is inserted into the network connection detection function check_connection() of the database kernel to obtain connection data including thread PID, connection timestamp and authentication status; the uprobe mechanism is used to add probes to the command scheduling function dispatch_command() of the database kernel and each link function in the SQL execution process to extract SQL execution data including SQL text, execution session ID and execution time information of each link function; the various link functions in the SQL execution process include client interaction, SQL parsing, SQL optimization, SQL execution, and storage engine corresponding functions; based on the uprobe mechanism, the fil_flush() flushing of the storage engine InnoDB is monitored The system uses the eBPF kprobe mechanism to obtain the disk call frequency, and uses the eBPF kprobe mechanism to insert probes into the preset read and write functions of the database kernel virtual file system VFS to obtain I / O operation information, and obtains disk data based on the disk call frequency and the I / O operation information; the preset read and write functions include vfs_read() and vfs_write(); the uprobe mechanism is used to insert probes into the lock management function of the storage engine InnoDB to obtain lock data including lock wait time, transaction ID and thread context; the lock management function includes lock_wait() and trx_commit(); based on the kprobe mechanism, probes are inserted into the data transmission function of the transmission control protocol layer in the operating system to obtain network traffic, and the network traffic is filtered by the database process PID to obtain network data; the data transmission functions include tcp_sendmsgtcp() and recvmsg(). Among them, in the transmission control protocol layer, a zero-copy data transfer method is used, using the eBPF ring buffer and the bpf_skb_store_bytes function to complete network packet parsing directly in kernel mode, avoiding user-mode memory copying.
[0058] Specifically, the dynamic detection mechanism of eBPF is used to deploy probes on the critical path of the database kernel to obtain various performance data of the database during operation based on the probes, including: using the uprobe mechanism of eBPF to obtain process data including the database process PID; inserting probes into the network connection detection function of the database kernel based on the uprobe mechanism to obtain connection data including thread PID, connection timestamp and authentication status; using the uprobe mechanism to add probes to the command scheduling function of the database kernel and each link function in the SQL execution process to extract SQL execution data including SQL text, execution session ID and execution time information of each link function; based on The uprobe mechanism is used to monitor the disk flushing of the storage engine to obtain the disk call frequency, and the kprobe mechanism of the eBPF is used to insert probes into the preset read and write functions of the database kernel virtual file system to obtain I / O operation information, and disk data is obtained based on the disk call frequency and the I / O operation information; the uprobe mechanism is used to insert probes into the lock management function of the storage engine to obtain lock data including lock wait time, transaction ID and thread context; based on the kprobe mechanism, a probe is inserted into the data transmission function of the transmission control protocol layer in the operating system to obtain network traffic, and the network traffic is filtered by the database process PID to obtain network data.
[0059] It can be understood that after obtaining the performance data, the data mapping mechanism of eBPF Map is used to trigger the corresponding processing function in the user state in the form of an event, and the performance data is passed to the user state. The first preset data structure of the eBPF Map, namely the LRU_HASH structure, is used to establish a mapping relationship between the five-tuple (source IP, destination IP, port, protocol, PID) and the database session ID to achieve the association between network traffic and SQL operations.
[0060] Step S12: identifying abnormal connection conditions of the database based on the mapping relationship and the performance data, and analyzing access conditions of locks and disks during operation of the database to obtain corresponding performance problems.
[0061] In this embodiment, the IP address to be detected is judged based on the preset whitelist and the preset blacklist, so as to obtain the corresponding abnormal connection status based on the judgment result. Figure 2A client address exception handling flow chart provided in this embodiment first establishes a preset whitelist and a preset blacklist, then stores a first client IP address that has been authenticated for legitimacy in the preset whitelist, and stores a second client IP address that has initiated attack behavior or frequently violated operations in the preset blacklist; if the IP address to be detected is not in the preset whitelist or the preset blacklist, then the number of check_connection() calls for the IP address to be detected within a preset time period is counted to obtain the number of connections, and it is determined whether the number of connections is greater than a target connection threshold; if the number of connections is greater than the target connection threshold, it indicates that the IP address to be detected has an abnormal connection situation, and the corresponding user information is determined based on the IP address to be detected, and the IP address to be detected is stored in the preset blacklist, and then the performance data is used to determine whether the abnormal connection situation will cause thread accumulation; if the abnormal connection situation will cause thread accumulation, the abnormal connection is terminated using a preset fuse mechanism and a corresponding alarm message is generated; if the number of connections is not greater than the target connection threshold, the connection data is updated based on the connection function corresponding to the IP address to be detected, and a connection status table corresponding to the IP address to be detected is displayed. It is worth mentioning that the target connection threshold can be adjusted according to actual conditions and is not specifically limited here.
[0062] Specifically, the abnormal connection situation of the database is identified based on the mapping relationship and the performance data, including: storing the first client IP address that has passed the legitimacy authentication into a preset whitelist, and storing the second client IP address that has launched attack behaviors and has frequent illegal operations into a preset blacklist; if the IP address to be detected is not in the preset whitelist and the preset blacklist, then counting the number of connections of the IP address to be detected within a preset time period, and judging whether the number of connections is greater than the target connection threshold; if the number of connections is greater than the target connection threshold, it indicates that the IP address to be detected has an abnormal connection situation, and the IP address to be detected is stored in the preset blacklist, and then the performance data is used to judge whether the abnormal connection situation will cause thread accumulation; if the abnormal connection situation will cause thread accumulation, the preset fuse mechanism is used to terminate the abnormal connection.
[0063] It is understood that the lock data is aggregated using the second preset data structure bpf_map of eBPF to generate a lock wait relationship graph, identify long wait paths, and then determine whether the lock wait in the lock wait relationship graph exceeds the target lock wait threshold. If the lock wait in the lock wait relationship graph exceeds the target lock wait threshold, the database query statement SHOW ENGINE INNODB STATUS can be used to determine the target location corresponding to the target lock wait that exceeds the target lock wait threshold to identify the corresponding performance issue. The target location includes a specific SQL statement and the thread holding the lock. Specifically, the lock and disk access status of the database during operation to identify the corresponding performance issue includes: aggregating the lock data using the second preset data structure of eBPF to generate a lock wait relationship graph, and determining whether the lock wait in the lock wait relationship graph exceeds the target lock wait threshold. If the lock wait in the lock wait relationship graph exceeds the target lock wait threshold, the database query statement can be used to determine the target location corresponding to the target lock wait that exceeds the target lock wait threshold to identify the corresponding performance issue.
[0064] In this embodiment, based on the disk data and using the struct file pointer, the file information corresponding to the target file involved in the disk read and write operation is determined, and the count parameter is used to record the single I / O data volume and file offset during the disk access process; the file information includes a file identifier and a file name; then, based on the database process PID, the database name, and the file name dimensions and combined with the second preset data structure, data aggregation dimension statistics are performed on the single I / O data volume and the file offset to obtain the read and write volume per second, and performance issues are determined based on the single I / O data volume, the file offset, and the read and write volume per second. Specifically, the lock and disk access status of the database during operation is analyzed to obtain the corresponding performance issues, including: determining the file information corresponding to the target file involved in the disk read and write operation based on the disk data, and recording the single I / O data volume and file offset during the disk access process; performing data aggregation dimension statistics on the single I / O data volume and the file offset to obtain the read and write volume per second; and determining the performance issues based on the single I / O data volume, the file offset, and the read and write volume per second.
[0065] It is understandable that the performance data can be written into Prometheus (a monitoring and alarm tool), and the performance data can be integrated based on the Prometheus to build a data view for reflecting the operation status of the database, realize the visualization of performance data, and then generate a monitoring panel based on the preset visualization panel and the data view, i.e., a Grafana panel, locate the target hotspot function and the corresponding function call path based on the monitoring panel and the Flame Mountain tool, and determine the corresponding performance problem using the target hotspot function and the function call path. Specifically, the abnormal connection of the database is identified based on the mapping relationship and the performance data, and the access to the lock and disk of the database during operation is analyzed to obtain the corresponding performance problem, including: writing the performance data into Prometheus, and integrating the performance data based on the Prometheus to build a data view for reflecting the operation status of the database; generating a monitoring panel based on the preset visualization panel and the data view, locating the target hotspot function and the corresponding function call relationship based on the monitoring panel and the Flame Mountain tool; determining the corresponding performance problem based on the target hotspot function and the function call relationship.
[0066] Step S13: Based on the performance data, determine parameter indicators including the number of lock waits per second, I / O queue depth, and TCP retransmission rate, input the parameter indicators into the machine learning model, and combine the root cause location rule engine to generate the root cause and optimization suggestions corresponding to the performance problem.
[0067] In this embodiment, based on the performance data and using a preset statistical method, parameter indicators including the number of lock waits per second, I / O queue depth and TCP (Transmission Control Protocol) retransmission rate are counted, and the average values corresponding to the parameter indicators in a preset sliding window are determined to obtain a time window sliding mean; the preset sliding window size can be adjusted according to actual conditions, and is set to 30 seconds by default; then a machine learning model and a root cause location rule engine are used to obtain a priority strategy, and the priority strategy can be that disk I / O delay takes precedence over lock wait, and lock wait is optimized over network delay; then the time window sliding mean is input into the machine learning model, and the root cause and optimization suggestions corresponding to the performance problem are generated based on the priority strategy and the decision tree; at the same time, an FPGA (field programmable gate array) is used to accelerate the hash table, and FPGA co-processing is introduced in the kernel hash table operation to reduce the GET / SET operation delay.
[0068] Specifically, the parameter indicators including the number of lock waits per second, the I / O queue depth and the TCP retransmission rate are determined based on the performance data, and the parameter indicators are input into the machine learning model, and combined with the root cause location rule engine to generate the root cause of the problem and optimization suggestions corresponding to the performance problem, including: based on the performance data and using a preset statistical method to count the parameter indicators including the number of lock waits per second, the I / O queue depth and the TCP retransmission rate, and determine the average value corresponding to the parameter indicators in a preset sliding window to obtain a time window sliding mean; use the machine learning model and the root cause location rule engine to obtain a priority strategy, input the time window sliding mean into the machine learning model, and generate the root cause of the problem and optimization suggestions corresponding to the performance problem based on the priority strategy and the decision tree.
[0069] As can be seen from the above, this application uses the dynamic detection mechanism of eBPF to detect various performance data in the database kernel, and then based on the mapping relationship between the network traffic and the database session ID, it can quickly locate the specific session, and based on the mapping relationship and comprehensive performance data, it can accurately identify abnormal connections in the database. Early warning of potential security risks or performance risks to avoid abnormal connections from causing greater impact on the database. In this way, through parameter indicators such as the number of lock waits per second, I / O queue depth, and TCP retransmission rate, the performance status of the database in terms of locks, disks, and networks is intuitively reflected, so that database performance problems can be quantitatively evaluated, providing a clear basis for judging the severity of performance problems, and using machine learning models and root cause location rule engines to analyze performance data, intelligent diagnosis of performance problems and automatic generation of optimization suggestions are achieved, which reduces manual intervention and improves the level of system automation.
[0070] It can be seen from the above embodiments that the present application monitors the database based on the dynamic detection mechanism of eBPF to achieve database performance optimization. Therefore, the process of monitoring the database based on the dynamic detection mechanism of eBPF is described.
[0071] See also Figure 3 As shown, the embodiment of the present invention discloses a specific database performance optimization method based on eBPF, including:
[0072] In this embodiment, the PID of the database process is first obtained to identify and track database-related activities. Then, the kprobe / uprobe mechanism of eBPF is used to insert probes into the operating system kernel function and the database kernel function respectively. Based on the probe, the corresponding kernel state event is triggered, and it is checked whether the process ID triggering the kernel state event matches the PID collected previously. If so, various performance data are collected through the probe, and the mapping relationship between network traffic and database session ID is established using the first preset data structure of eBPF. Figure 4 A database performance problem analysis flowchart provided in this embodiment determines whether the kernel-mode event is a lock access event; if so, records the lock access time, aggregates the lock data in the performance data through the second preset data structure bpf_map of the eBPF to generate a lock wait relationship graph, and then determines whether the lock wait in the lock wait relationship graph exceeds the target lock wait threshold; updates the lock data; if not, records the unlock time, and then obtains the number of lock waits in the lock wait relationship graph.
[0073] It is understood that if the kernel-mode event is a disk access event, the file information corresponding to the target file involved in the disk read / write operation is determined based on the disk data and using the struct file pointer, and the target instance, target database, and target table corresponding to the target file are updated, and then the I / O queue depth is determined. If the kernel-mode event is an SQL execution event, the SQL statement corresponding to the SQL execution event is parsed to obtain the corresponding statement information. Figure 5 A security audit flowchart provided for this embodiment determines whether the statement information belongs to a sensitive operation based on a pre-loaded audit rule set; if so, generates corresponding alarm information and records the statement information; if not, determines whether the return value length corresponding to the execution result of the SQL statement exceeds the target length threshold; if so, generates an alarm information indicating the risk of data leakage.
[0074] Furthermore, if the kernel-mode event is a TCP message transmission / reception event, the message header corresponding to the TCP message transmission / reception event is parsed to obtain message header information. Based on the message header information, the length, frequency, and message type of the TCP message are updated, and the TCP retransmission rate corresponding to the TCP message transmission / reception event is determined. Parameter indicators for evaluating the database performance status are determined based on the number of lock waits, the I / O queue depth, and the TCP retransmission rate. These performance indicators are input into a machine learning model, and combined with the root cause location rule engine, the root cause of the problem and optimization suggestions corresponding to the performance issue are generated. For example, if the number of lock waits is found to be excessive, it will be recommended to optimize the lock strategy; if the I / O queue depth is too large, it will be recommended to adjust the I / O configuration. The analysis results are then displayed in a visual manner, including the execution time of functions at each layer, the call path of hot functions, a root cause analysis report of performance bottlenecks, and performance optimization suggestions, etc., to help database administrators quickly locate and resolve performance issues.
[0075] As can be seen from the above, the dynamic detection mechanism of eBPF in this application monitors the database to obtain performance data, and then based on the mapping relationship between the network traffic and the database session ID and the performance data, it can accurately identify abnormal situations in the database. Then, through parameter indicators such as the number of lock waits per second, I / O queue depth, and TCP retransmission rate, and using machine learning models and root cause location rule engines to analyze the performance data, it can identify and warn of potential security risks and performance issues, and provide corresponding optimization suggestions, thereby ensuring the security of the database and the stability of performance.
[0076] Accordingly, see Figure 6 As shown, the present application also provides a database performance optimization device based on eBPF, including:
[0077] A performance data acquisition module 11 is configured to deploy probes on the critical path of the database kernel using the dynamic detection mechanism of eBPF, obtain various performance data of the database during operation based on the probes, transfer the performance data to user mode, and establish a mapping relationship between network traffic and database session ID using a first preset data structure of eBPF; the performance data includes process data, connection data, SQL execution data, disk data, lock data, and network data;
[0078] An access situation analysis module 12 is used to identify abnormal connection situations of the database based on the mapping relationship and the performance data, and analyze the access situations of the locks and disks of the database during operation to identify corresponding performance issues;
[0079] The parameter indicator determination module 13 is used to determine parameter indicators including the number of lock waits per second, I / O queue depth and TCP retransmission rate based on the performance data, and input the parameter indicators into the machine learning model, and combine with the root cause location rule engine to generate the root cause and optimization suggestions corresponding to the performance problem.
[0080] As can be seen from the above, this application uses the dynamic detection mechanism of eBPF to obtain various types of performance data in the database kernel, and then can quickly locate the specific session based on the mapping relationship between the network traffic and the database session ID. Based on the mapping relationship and comprehensive performance data, it can accurately identify abnormal connections in the database. Early warning of potential security risks or performance risks to avoid abnormal connections from causing greater impact on the database. In this way, through parameter indicators such as the number of lock waits per second, I / O queue depth, and TCP retransmission rate, the performance status of the database in terms of locks, disks, and networks is intuitively reflected, so that database performance problems can be quantitatively evaluated, providing a clear basis for judging the severity of performance problems, and using machine learning models and root cause location rule engines to analyze performance data, intelligent diagnosis of performance problems and automatic generation of optimization suggestions are achieved, which reduces manual intervention and improves the level of system automation.
[0081] In some specific implementations, the performance data acquisition module 11 may specifically include:
[0082] A process data acquisition unit, configured to acquire process data including the database process PID using the uprobe mechanism of eBPF;
[0083] A connection data acquisition unit, configured to insert a probe into the network connection detection function of the database kernel based on the uprobe mechanism to obtain connection data including thread PID, connection timestamp and authentication status;
[0084] An execution data acquisition unit, configured to use the uprobe mechanism to add probes to the command scheduling function of the database kernel and each link function in the SQL execution process to extract SQL execution data including SQL text, execution session ID, and execution duration information of each link function;
[0085] A disk data acquisition unit is configured to monitor the disk flushing of the storage engine based on the uprobe mechanism to obtain a disk call frequency, insert probes into preset read and write functions of the database kernel virtual file system using the eBPF kprobe mechanism to obtain I / O operation information, and obtain disk data based on the disk call frequency and the I / O operation information;
[0086] A lock data acquisition unit, configured to insert a probe into the lock management function of the storage engine using the uprobe mechanism to obtain lock data including lock wait time, transaction ID, and thread context;
[0087] The network data acquisition unit is used to insert a probe into the data transmission function of the transmission control protocol layer in the operating system based on the kprobe mechanism to obtain network traffic, and filter the network traffic through the database process PID to obtain network data.
[0088] In some specific implementations, the access situation analysis module 12 may specifically include:
[0089] The client address storage unit is used to store the first client IP address that has passed the legality authentication into a preset whitelist, and store the second client IP address that has launched an attack or has frequently violated the rules into a preset blacklist;
[0090] a connection count judgment unit, configured to count the number of connections of the IP address to be detected within a preset time period if the IP address to be detected is not in the preset whitelist and the preset blacklist, and to judge whether the number of connections is greater than a target connection threshold;
[0091] a performance data judgment unit, configured to, if the number of connections is greater than the target connection threshold, indicate that the IP address to be detected has an abnormal connection situation, store the IP address to be detected in the preset blacklist, and then use the performance data to determine whether the abnormal connection situation will cause thread accumulation;
[0092] The abnormal connection termination unit is used to terminate the abnormal connection by using a preset fuse mechanism if the abnormal connection situation will cause thread accumulation.
[0093] In some specific implementations, the access situation analysis module 12 may specifically include:
[0094] a relationship graph generating unit, configured to aggregate the lock data using a second preset data structure of the eBPF to generate a lock wait relationship graph, and determine whether the lock wait in the lock wait relationship graph exceeds a target lock wait threshold;
[0095] The target position determining unit is used to determine the target position corresponding to the target lock wait that exceeds the target lock wait threshold using a database query statement if the lock wait in the lock wait relationship graph exceeds the target lock wait threshold, so as to obtain the corresponding performance problem.
[0096] In some specific implementations, the access situation analysis module 12 may specifically include:
[0097] a file information determining unit, configured to determine file information corresponding to a target file involved in a disk read / write operation based on the disk data, and record a single I / O data volume and a file offset during the disk access process;
[0098] A read / write volume determination unit is configured to perform data aggregation dimension statistics on the single I / O data volume and the file offset to obtain the read / write volume per second;
[0099] The first performance problem determining unit is configured to determine a performance problem based on the single I / O data volume, the file offset, and the read / write volume per second.
[0100] In some specific implementations, the access situation analysis module 12 may specifically include:
[0101] a performance data integration unit, configured to write the performance data into Prometheus and integrate the performance data based on Prometheus to construct a data view reflecting the operation status of the database;
[0102] A function hotspot function determination unit is used to generate a monitoring panel based on a preset visualization panel and the data view, and locate the target hotspot function and the corresponding function call relationship based on the monitoring panel and using the Flaming Mountain tool;
[0103] The second performance problem determining unit is configured to determine corresponding performance problems based on the target hotspot function and the function call relationship.
[0104] In some specific implementations, the parameter index determination module 13 may specifically include:
[0105] a sliding mean determining unit, configured to calculate, based on the performance data and using a preset statistical method, parameter indicators including the number of lock waits per second, the I / O queue depth, and the TCP retransmission rate, and determine an average value corresponding to the parameter indicators within a preset sliding window to obtain a time window sliding mean;
[0106] An optimization suggestion determination unit is used to obtain a priority strategy using a machine learning model and a root cause location rule engine, input the time window sliding mean into the machine learning model, and generate the root cause and optimization suggestion corresponding to the performance problem based on the priority strategy and the decision tree.
[0107] Furthermore, the embodiment of the present application also discloses an electronic device, Figure 7This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be considered as any limitation on the scope of use of this application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the database performance optimization method based on eBPF disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0108] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device. The communication protocol it follows is any communication protocol that can be applied to the technical solution of this application and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world. Its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0109] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or CD, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0110] The operating system 221 is used to manage and control the hardware devices on the electronic device 20 and the computer program 222, and can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of implementing the eBPF-based database performance optimization method executed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program capable of implementing other specific tasks.
[0111] Furthermore, this application discloses a computer-readable storage medium for storing a computer program; wherein, when executed by a processor, the computer program implements the aforementioned eBPF-based database performance optimization method. The specific steps of this method can be referred to the corresponding content disclosed in the aforementioned embodiments and will not be repeated here.
[0112] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from the other embodiments. Reference can be made to the descriptions of the identical or similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions of the methods.
[0113] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0114] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0115] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
[0116] The above is a detailed introduction to the technical solution provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the ideas of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A database performance optimization method based on eBPF, characterized in that: include: Probes are deployed on the critical path of the database kernel using the dynamic detection mechanism of eBPF to obtain various performance data of the database during operation based on the probes, the performance data is passed to user mode, and a mapping relationship between network traffic and database session ID is established using a first preset data structure of the eBPF; the performance data includes process data, connection data, SQL execution data, disk data, lock data, and network data; Identifying abnormal connection conditions of the database based on the mapping relationship and the performance data, and analyzing lock and disk access conditions of the database during operation to obtain corresponding performance issues; Based on the performance data, parameter indicators including the number of lock waits per second, I / O queue depth and TCP retransmission rate are determined, and the parameter indicators are input into the machine learning model, and combined with the root cause location rule engine to generate the root cause and optimization suggestions corresponding to the performance problem.
2. The database performance optimization method based on eBPF according to claim 1, characterized in that: The dynamic detection mechanism of eBPF is used to deploy probes on the critical path of the database kernel to obtain various performance data of the database during operation based on the probes, including: Use eBPF's uprobe mechanism to obtain process data including the database process PID; Inserting a probe into the network connection detection function of the database kernel based on the uprobe mechanism to obtain connection data including thread PID, connection timestamp and authentication status; Using the uprobe mechanism, probes are added to the command scheduling function of the database kernel and each link function in the SQL execution process to extract SQL execution data including SQL text, execution session ID and execution duration information of each link function; The storage engine's disk flushing is monitored based on the uprobe mechanism to obtain the disk call frequency, and the eBPF kprobe mechanism is used to insert probes into the preset read and write functions of the database kernel virtual file system to obtain I / O operation information. Disk data is obtained based on the disk call frequency and the I / O operation information. Using the uprobe mechanism, a probe is inserted into the lock management function of the storage engine to obtain lock data including lock wait time, transaction ID, and thread context; Based on the kprobe mechanism, a probe is inserted into the data transmission function of the transmission control protocol layer in the operating system to obtain network traffic, and the network traffic is filtered through the database process PID to obtain network data.
3. The database performance optimization method based on eBPF according to claim 1, characterized in that: The identifying an abnormal connection condition of the database based on the mapping relationship and the performance data includes: The first client IP address that has been authenticated is stored in a preset whitelist, and the second client IP address that has launched attacks or has frequent illegal operations is stored in a preset blacklist; If the IP address to be detected is not in the preset whitelist and the preset blacklist, the number of connections of the IP address to be detected within a preset time period is counted, and it is determined whether the number of connections is greater than the target connection threshold; If the number of connections is greater than the target connection threshold, it indicates that the IP address to be detected has an abnormal connection situation, and the IP address to be detected is stored in the preset blacklist, and then the performance data is used to determine whether the abnormal connection situation will cause thread accumulation; If the abnormal connection situation will cause thread accumulation, the preset fuse mechanism is used to terminate the abnormal connection.
4. The database performance optimization method based on eBPF according to claim 1, characterized in that: The analysis of the lock and disk access during the operation of the database to obtain corresponding performance issues includes: aggregating the lock data using a second preset data structure of the eBPF to generate a lock wait relationship graph, and determining whether the lock wait in the lock wait relationship graph exceeds a target lock wait threshold; If the lock wait in the lock wait relationship graph exceeds the target lock wait threshold, a database query statement is used to determine a target position corresponding to the target lock wait that exceeds the target lock wait threshold to obtain a corresponding performance problem.
5. The database performance optimization method based on eBPF according to claim 1, characterized in that: The analysis of the lock and disk access during the operation of the database to obtain corresponding performance issues includes: Determining file information corresponding to a target file involved in a disk read / write operation based on the disk data, and recording a single I / O data volume and a file offset during the disk access process; Performing data aggregation dimension statistics on the single I / O data volume and the file offset to obtain the read and write volume per second; A performance problem is determined based on the single I / O data volume, the file offset, and the read and write volume per second.
6. The database performance optimization method based on eBPF according to claim 1, characterized in that: The identifying of abnormal connection conditions of the database based on the mapping relationship and the performance data, and analyzing the access conditions of locks and disks of the database during operation to obtain corresponding performance issues, includes: Writing the performance data into Prometheus, and integrating the performance data based on Prometheus to construct a data view reflecting the database operation status; Generate a monitoring panel based on a preset visualization panel and the data view, and locate target hotspot functions and corresponding function call relationships based on the monitoring panel and using the Huoyanshan tool; Corresponding performance issues are determined based on the target hotspot function and the function call relationship.
7. The database performance optimization method based on eBPF according to any one of claims 1 to 6, characterized in that: Determining parameter indicators including the number of lock waits per second, I / O queue depth, and TCP retransmission rate based on the performance data, inputting the parameter indicators into a machine learning model, and combining the root cause location rule engine to generate the root cause and optimization suggestions corresponding to the performance issue, including: Based on the performance data and using a preset statistical method, statistically analyze parameter indicators including the number of lock waits per second, I / O queue depth, and TCP retransmission rate, and determine the average values corresponding to the parameter indicators within a preset sliding window to obtain a time window sliding mean; A priority strategy is obtained using a machine learning model and a root cause location rule engine, the time window sliding mean is input into the machine learning model, and the root cause and optimization suggestions corresponding to the performance problem are generated based on the priority strategy and the decision tree.
8. A database performance optimization device based on eBPF, characterized in that: include: A performance data acquisition module is configured to deploy probes on the critical path of the database kernel using the dynamic detection mechanism of eBPF, obtain various performance data of the database during operation based on the probes, pass the performance data to user mode, and establish a mapping relationship between network traffic and database session ID using a first preset data structure of eBPF; the performance data includes process data, connection data, SQL execution data, disk data, lock data, and network data; an access situation analysis module, configured to identify abnormal connection situations of the database based on the mapping relationship and the performance data, and to analyze access situations of locks and disks during operation of the database to identify corresponding performance issues; A parameter indicator determination module is used to determine parameter indicators including the number of lock waits per second, I / O queue depth and TCP retransmission rate based on the performance data, input the parameter indicators into the machine learning model, and combine with the root cause location rule engine to generate the root cause and optimization suggestions corresponding to the performance problem.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the database performance optimization method based on eBPF according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that Used to store a computer program, wherein when the computer program is executed by a processor, the eBPF-based database performance optimization method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Kernel audit message transmission method based on structural body
CN121116899A