Fault detection method and device based on DRDS, equipment and storage medium

By acquiring the operating parameter information of the distributed relational database for status and performance testing, the problem of inaccurate detection in existing technologies is solved, the accuracy of fault detection is improved, the cost and error rate of manual operation are reduced, and the stability and accuracy of database performance testing are enhanced.

CN120929440APending Publication Date: 2025-11-11CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410579529.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-10
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing distributed relational databases are complex in troubleshooting, resulting in time-consuming and labor-intensive fault detection, which affects the stable operation of production applications.

Method used

By acquiring the operating parameter information of the distributed relational database, including status information and performance information, status detection and performance detection are performed sequentially to generate a fault detection report.

Benefits of technology

It enables fault detection in distributed databases, improves detection accuracy, reduces manual operation costs and error rates, and enhances database stability and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929440A_ABST
    Figure CN120929440A_ABST
Patent Text Reader

Abstract

The invention provides a DRDS-based fault detection method and device, equipment and a storage medium, and is applied to the technical field of databases. When a plurality of pieces of state information and a plurality of pieces of performance information corresponding to the to-be-detected distributed relational database are obtained, the plurality of pieces of state information and the plurality of pieces of performance information are detected in sequence to obtain a plurality of state detection results corresponding to the plurality of pieces of state information and a plurality of performance detection results corresponding to the plurality of pieces of performance information, and generating a fault detection report according to the plurality of state detection results and the plurality of performance detection results. According to the method, the problem that the operation fault of the distributed relational database cannot be accurately detected is solved, the operation state and performance of the database can be more accurately judged, the cost and the error rate of manual operation can be reduced, the fault detection accuracy is further improved, system resources can be more reasonably distributed, and the fault detection efficiency is improved. And the performance and the reliability of the distributed relational database are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of database technology, and in particular to a fault detection method, apparatus, device and storage medium based on DRDS. Background Technology

[0002] With the rapid development of the internet and the increasing demand for data processing, single-machine relational databases have gradually encountered problems such as capacity bottlenecks, difficulties in expansion, and high usage costs. These challenges limit enterprises' data processing capabilities and efficiency. To solve these problems, distributed relational database services have emerged. These services adopt a distributed architecture, possessing characteristics such as lightweight, flexible, stable, and efficient performance, and are highly compatible with MySQL protocols and syntax. They achieve distributed data storage and processing, thereby improving database capacity and scalability.

[0003] However, existing distributed relational databases present certain complexities in troubleshooting. When problems arise, relevant personnel typically need to manually check and test various logs and processes within the current environment to pinpoint the issue. Due to the architectural characteristics of distributed databases, the numerous and complex metrics requiring troubleshooting make this process extremely time-consuming and labor-intensive, severely impacting the stable operation of production applications.

[0004] Therefore, the urgent technical problem to be solved is how to accurately detect operational failures in distributed relational databases so that problems can be quickly found and resolved when they occur, thereby improving the stability and operational efficiency of distributed relational databases. Summary of the Invention

[0005] This application provides a fault detection method, apparatus, device, and storage medium based on DRDS to solve the problem of the inability to perform fault detection on distributed relational databases.

[0006] Firstly, this application provides a fault detection method based on DRDS, which includes:

[0007] Obtain the distributed relational database to be tested, which is used to indicate that there is an anomaly in the distributed relational database during operation;

[0008] Based on the distributed relational database to be detected, the operating parameter information of the distributed relational database to be detected is determined. The operating parameter information includes: status information and performance information. The status information includes multiple sub-candidate status information, and the performance information includes multiple sub-candidate performance information.

[0009] The state detection is performed on the multiple sub-candidate state information in sequence to obtain multiple state detection results, and the performance detection is performed on the multiple sub-candidate performance information in sequence to obtain multiple performance detection results;

[0010] A fault detection report is generated based on the multiple status detection results and the multiple performance detection results.

[0011] Optionally, the sub-candidate state information includes: link information, lock management information, long-connection information, and transaction information. The process of sequentially performing state detection on the multiple sub-candidate state information yields multiple state detection results, including:

[0012] If a status detection is performed on the link information, then the status detection result corresponding to the link information is determined;

[0013] and / or;

[0014] If a status detection is performed on the lock management information, a status detection result corresponding to the lock management information is determined. The lock management information includes: lock information and lock waiting information.

[0015] and / or;

[0016] If a status detection is performed on the long link information, a status detection result corresponding to the long link information is determined. The long link information includes: execution time and execution status.

[0017] and / or;

[0018] If a status check is performed on the transaction information, the status check result corresponding to the transaction information is determined.

[0019] Optionally, the sub-candidate performance information includes: operating information of multiple RDS instances, operating information of DRDS-Server, master-slave failover information, and slow SQL query information. The performance of the multiple sub-candidate performance information is then sequentially tested to obtain multiple performance test results, including:

[0020] If performance testing is performed on the operation information of the multiple RDS, then based on the first anomaly detection model and the operation information of each RDS, the performance testing result corresponding to the operation information of each RDS is determined.

[0021] and / or;

[0022] If the DRDS-Server's operating information is to be tested for performance, a second anomaly detection model is used to test the DRDS-Server's operating information to obtain the performance test result corresponding to the DRDS-Server's operating information. The second anomaly detection model is different from the first anomaly detection model.

[0023] and / or;

[0024] If performance testing is performed on the primary / backup switchover information, then the performance testing result corresponding to the primary / backup switchover information is determined based on the primary / backup switchover information.

[0025] and / or;

[0026] If performance testing is performed on the slow SQL query information, then 301 determines the performance testing result corresponding to the slow SQL query information based on the slow SQL query information.

[0027] Optionally, before determining the performance detection result corresponding to the operating information of each RDS based on the first anomaly detection model and the operating information of each RDS, the method further includes:

[0028] Based on the operational information of the multiple RDSs, determine the first historical monitoring indicator dataset corresponding to the operational information of each RDS;

[0029] The multiple first historical monitoring indicator datasets are classified and summarized to obtain multiple candidate monitoring indicator datasets with the same category;

[0030] The first anomaly detection model is obtained by training the multiple candidate monitoring indicator datasets using the isolated forest algorithm.

[0031] Optionally, before performing performance testing on the DRDS-Server's operational information using the second anomaly detection model to obtain the performance testing result corresponding to the DRDS-Server's operational information, the method further includes:

[0032] Based on the DRDS-Server's operating information, determine the second historical monitoring indicator dataset corresponding to the DRDS-Server's operating information;

[0033] The isolated forest algorithm is used to train the second historical monitoring index dataset to obtain the second anomaly detection model.

[0034] Optionally, determining the performance test result corresponding to the slow SQL query information based on the slow SQL query information includes:

[0035] The slow SQL query information is parsed and processed to obtain multiple characteristic slow SQL query statements;

[0036] The multiple slow SQL query statements with the same characteristics are subjected to feature merging processing to obtain a set of multiple slow SQL query statements with the same characteristics;

[0037] Based on the multiple sets of slow SQL query statements, determine the sub-performance test result corresponding to each set of slow SQL query statements;

[0038] Based on the sub-performance test results corresponding to each set of slow SQL query statements, determine the performance test results corresponding to the slow SQL query information.

[0039] Optionally, the method further includes:

[0040] The multiple sets of slow SQL query statements are parsed and processed to obtain the problematic slow SQL query statements;

[0041] The problematic slow SQL query is diagnosed and processed using a set of diagnostic processing rules to obtain the target solution corresponding to the problematic slow SQL query. The set of diagnostic processing rules includes: heuristic diagnostic suggestion rules, index optimization suggestion rules, and execution plan interpretation rules.

[0042] Secondly, this application provides a fault detection device based on DRDS, the device comprising:

[0043] The acquisition module is used to acquire the distributed relational database to be detected, wherein the distributed relational database to be detected is used to indicate that there is an anomaly in the distributed relational database during operation;

[0044] The determination module is used to determine the operating parameter information of the distributed relational database to be detected based on the distributed relational database to be detected. The operating parameter information includes: status information and performance information. The status information includes multiple sub-candidate status information, and the performance information includes multiple sub-candidate performance information.

[0045] The processing module is used to sequentially perform state detection on the multiple sub-candidate state information to obtain multiple state detection results, and sequentially perform performance detection on the multiple sub-candidate performance information to obtain multiple performance detection results.

[0046] The generation module is used to generate a fault detection report based on the multiple state detection results and the multiple performance detection results.

[0047] Optionally, the determining module is further configured to determine the state detection result corresponding to the link information when performing state detection on the link information;

[0048] and / or;

[0049] The determining module is further configured to determine the state detection result corresponding to the lock management information when performing state detection on the lock management information, wherein the lock management information includes: lock information and lock waiting information;

[0050] and / or;

[0051] The determining module is further configured to determine the status detection result corresponding to the long link information when performing status detection on the long link information, wherein the long link information includes: execution time and execution status;

[0052] and / or;

[0053] The determining module is further configured to determine the status detection result corresponding to the transaction information when performing status detection on the transaction information.

[0054] Optionally, the determining module is further configured to, when performing performance testing on the running information of the plurality of RDS, determine the performance testing result corresponding to the running information of each RDS based on the first anomaly detection model and the running information of each RDS;

[0055] and / or;

[0056] The determining module is further configured to, when performing performance testing on the DRDS-Server's operating information, use a second anomaly detection model to perform performance testing on the DRDS-Server's operating information, and obtain a performance testing result corresponding to the DRDS-Server's operating information. The second anomaly detection model is different from the first anomaly detection model.

[0057] and / or;

[0058] The determining module is further configured to, when performing performance testing on the primary / backup switching information, determine the performance testing result corresponding to the primary / backup switching information based on the primary / backup switching information;

[0059] and / or;

[0060] The determining module is further configured to, when performing performance testing on the slow SQL query information, determine the performance testing result corresponding to the slow SQL query information based on the slow SQL query information.

[0061] Optionally, the determining module is further configured to determine a first historical monitoring indicator dataset corresponding to each RDS operation information based on the operation information of the plurality of RDSs;

[0062] The processing module is also used to classify and summarize the multiple first historical monitoring indicator datasets to obtain multiple candidate monitoring indicator datasets with the same category.

[0063] The device further includes: a training module;

[0064] The training module is used to train the first anomaly detection model on the multiple candidate monitoring indicator datasets using the isolated forest algorithm.

[0065] Optionally, the determining module is further configured to determine a second historical monitoring indicator dataset corresponding to the operation information of the DRDS-Server based on the operation information of the DRDS-Server;

[0066] The training module is also used to train the second historical monitoring indicator dataset using the isolated forest algorithm to obtain the second anomaly detection model.

[0067] Optionally, the processing module is further configured to parse the slow SQL query information to obtain multiple characteristic slow SQL query statements;

[0068] The processing module is also used to perform feature merging processing on the multiple slow SQL query statements with the same characteristics to obtain a set of multiple slow SQL query statements with the same characteristics.

[0069] The determining module is further configured to determine a sub-performance detection result corresponding to each slow SQL query statement set based on the plurality of slow SQL query statement sets;

[0070] The determining module is specifically used to determine the performance test result corresponding to the slow SQL query information based on the sub-performance test result corresponding to each slow SQL query statement set.

[0071] Optionally, the processing module is further configured to parse the multiple sets of slow SQL query statements to obtain the problematic slow SQL query statements;

[0072] The processing module is further configured to perform diagnostic processing on the problematic slow SQL query statement using a diagnostic processing rule set to obtain a target solution corresponding to the problematic slow SQL query statement. The diagnostic processing rule set includes: heuristic diagnostic suggestion rules, index optimization suggestion rules, and execution plan interpretation rules.

[0073] Thirdly, this application provides a fault detection device based on DRDS, comprising:

[0074] Memory;

[0075] processor;

[0076] The memory stores computer-executed instructions;

[0077] The processor executes computer execution instructions stored in the memory to implement the DRDS-based fault detection method as described in the first aspect and various possible implementations of the first aspect above.

[0078] Fourthly, this application provides a computer storage medium storing computer execution instructions thereon, which are executed by a processor to implement the DRDS-based fault detection method as described in the first aspect and various possible implementations of the first aspect above.

[0079] The fault detection method based on DRDS provided in this application, upon obtaining multiple state information and multiple performance information corresponding to the distributed relational database to be detected, sequentially detects these multiple state information and multiple performance information to obtain multiple state detection results corresponding to the multiple state information and multiple performance detection results corresponding to the multiple performance information. Then, based on the multiple state detection results and multiple performance detection results, a fault detection report is generated. This method solves the problem of inaccurate detection of operational faults in distributed relational databases. It can not only more accurately determine the database's operational status and performance, but also reduce the cost and error rate of manual operation, thereby improving the accuracy of fault detection. This allows for more rational allocation of system resources, improving the performance and reliability of the distributed relational database. Attached Figure Description

[0080] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0081] Figure 1 This is the flowchart of the DRDS-based fault detection method provided in this application. Figure 1 ;

[0082] Figure 2 This is the flowchart of the DRDS-based fault detection method provided in this application. Figure 2 ;

[0083] Figure 3 This is the flowchart of the DRDS-based fault detection method provided in this application. Figure 3 ;

[0084] Figure 4 This is a schematic diagram of the DRDS-based fault detection device provided in this application;

[0085] Figure 5 This is a schematic diagram of the fault detection device based on DRDS provided in this application.

[0086] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0087] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0088] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented, for example, in orders other than those illustrated or described herein.

[0089] In this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0090] First, the terms used in this application will be explained:

[0091] 1. DRDS: DRDS stands for Distributed Relational Database Service, a horizontally scalable, read-write separated online distributed database service. It's a distributed OLTP database service product based on MySQL storage and employing database sharding and partitioning techniques for horizontal scaling. This database, based on RDS, achieves high availability, high performance, and scalability through distributed storage and database sharding / partitioning. DRDS is suitable for large-scale data storage and computation scenarios, capable of handling ultra-large-scale data storage and complex computational tasks.

[0092] 2. RDS: RDS stands for Relational Database Service, a ready-to-use, stable, reliable, and elastically scalable online database service. It supports engines such as MySQL, SQL Server, PostgreSQL, PPAS (PostgreSQL Plus Advanced Server, highly compatible with Oracle databases), and MariaDB, and provides a complete solution for disaster recovery, backup, restoration, monitoring, and migration.

[0093] 3. DRDS-Server: The DRDS-Server is a DRDS service node and a core component of DRDS, providing SQL parsing, optimization, routing, and result merging. It handles SQL requests, including parsing SQL statements, optimizing execution plans, and routing to appropriate database nodes. Each service node is stateless and processes SQL requests simultaneously. In distributed database services, the DRDS-Server is a crucial component, improving the performance and availability of the database service.

[0094] 4. Connection Information: DRDS connection information includes the database domain name, port, username, and password. This information is used to establish a connection with the DRDS database server for data access and manipulation. Correct connection information, including the correct database domain name, port, username, and password, is required when connecting. Incorrect or invalid connection information will prevent the establishment of a connection and access to the database service.

[0095] 5. Lock Information: DRDS lock information consists of parameters related to locks, used to control the behavior of locks in the database. There are four lock-related parameters in DRDS, with the following meanings:

[0096] (1) deadlock_timeout: indicates the deadlock detection time. Deadlock detection will be performed after this time. The default is 1 second.

[0097] (2) lockwait_timeout: This takes effect when a table lock conflict occurs. If the waiting time for the table lock exceeds the configured time, an error will be thrown and returned. The default is 20 minutes.

[0098] (3) update_lockwait_timeout: This takes effect when a record lock conflict occurs. If the waiting time for the record lock exceeds update_lockwait_timeout, an error is thrown and returned. The default is 2 minutes.

[0099] (4)ddl_lock_timeout: This function takes effect when there is a conflict with the level 8 table lock. If the time to wait for the level 8 lock exceeds the configured time, an error will be thrown and returned. The default value is 0, which means that it is not effective and needs to be enabled manually by the user.

[0100] These parameters can be used to control lock wait timeout behavior in the database to avoid performance issues caused by long wait times. They can also be adjusted and optimized according to actual business needs and system performance.

[0101] 6. Lock Waiting Information: In DRDS, lock waiting information refers to a situation where a transaction needs to acquire and hold a lock during execution, but that lock is already held by another transaction. This waiting can have the following consequences:

[0102] (1) The party holding the lock releases the lock (the corresponding action is usually the transaction holding the lock submits), waits for the party holding the lock to acquire the lock, and then continues execution.

[0103] (2) The waiting party for the lock causes an error due to lock waiting timeout.

[0104] 7. Long-lived connection information: DRDS long-lived connection information refers to information related to data access and operations performed using long-lived connections when using the DRDS database service. Long-lived connections can reduce the overhead of frequently establishing and closing connections, improving the efficiency and performance of data access. Long-lived connection information includes both execution time and execution status.

[0105] (1) Execution time refers to the time spent by the database to perform a specific operation. By monitoring and recording execution time, performance bottlenecks and problems of database operations can be analyzed, and performance problems can be identified and resolved in a timely manner.

[0106] (2) Execution status refers to the state of the database at a specific moment, including connection status, locking status, transaction status, etc. By monitoring and recording the execution status, the execution process and results of database operations can be analyzed, abnormal situations can be detected and resolved in a timely manner, and the stability and reliability of the database can be guaranteed.

[0107] 8. Transaction Information: DRDS transaction information refers to information related to transaction processing when using the DRDS database service. A transaction is a collection of commands that can be executed sequentially and exclusively in a queue. During transaction execution, commands in the queue are executed serially in order, and command requests submitted by other clients will not be inserted into the transaction's command execution sequence.

[0108] In DRDS, transaction information includes details about transaction initiation, commit, and rollback. This information allows you to understand the execution process and results of transactions, analyzing whether they were successfully executed or if any exceptions occurred. Furthermore, transaction information can be used to monitor and optimize database performance and transaction processing capabilities, helping database administrators to promptly identify and resolve issues, ensuring the normal operation and performance optimization of database services.

[0109] 9. Master-Slave Switchover Information: DRDS master-slave switchover information refers to the operations and status information related to master-slave switching in the DRDS system. Master-slave switchover means that when the master node fails, the slave node takes over the work of the master node to ensure the normal operation of the distributed relational database service.

[0110] In a DRDS system, master-slave failover information includes failover status, progress, and time. This information helps database administrators identify and resolve issues promptly, ensuring the normal operation and performance optimization of the distributed relational database service. Furthermore, master-slave failover information can be used to analyze and optimize the performance and architecture design of the distributed relational database, improving its overall performance and reliability.

[0111] 10. Slow SQL Query Information: DRDS slow SQL query information refers to information about SQL statements whose execution time exceeds a certain threshold when using the DRDS database service. This slow SQL query information can be caused by various factors, such as network latency, high CPU load, and disk I / O bottlenecks. In DRDS, slow SQL query information typically includes the following:

[0112] (1) SQL statement itself: Slow SQL query information will record SQL statements whose execution time exceeds the threshold, including the text content of the SQL statement.

[0113] (2) Execution time: Slow SQL query information records the execution time of each slow SQL statement, that is, the time required from the start of the SQL statement to the completion of the execution.

[0114] (3) Execution plan: Slow SQL query information may also record the execution plan related to the SQL statement execution, including the tables accessed, indexes, and query conditions executed.

[0115] (4) Other information: Depending on the specific implementation and configuration, slow SQL query information may also include other relevant information, such as client IP address, user identifier, etc.

[0116] 10. Isolation Forest Algorithm: Isolation Forest is a fast anomaly detection method based on ensemble algorithms. It features linear time complexity and high accuracy, making it a state-of-the-art algorithm suitable for big data processing. It posits that anomalous samples are typically few and different: compared to normal samples, they are fewer in number and have significantly different feature values. Therefore, anomalous samples are more easily isolated.

[0117] Isolation Forest is suitable for continuous data and uses an isolation tree structure to isolate samples. Specifically, the Isolation Forest algorithm detects outliers by isolating sample points. It uses a random hyperplane to cut the data space, cutting it multiple times until each subspace contains only one data point. In this process, high-density clusters can be cut many times before stopping, while low-density points are easily isolated into a single subspace. Therefore, outliers (i.e., sparse and high-density points) are isolated earlier, while normal values ​​are isolated later.

[0118] With the rapid development of the internet, data volume has exploded, and the demand for data processing has continuously increased. Against this backdrop, single-machine relational databases have gradually revealed problems such as capacity bottlenecks, difficulties in expansion, and high usage costs. These issues have become major challenges to enterprises' data processing capabilities and efficiency. To address these problems, distributed relational database services have emerged.

[0119] Distributed relational database services employ a distributed architecture, distributing data across multiple nodes, and are characterized by lightweightness, flexibility, stability, and efficiency. This service is highly compatible with the MySQL protocol and syntax, allowing developers to continue using familiar MySQL syntax for data manipulation without requiring additional adaptation or learning. Simultaneously, distributed relational database services achieve distributed data storage and processing, effectively improving database capacity and scalability, and meeting enterprises' needs for large-scale data processing.

[0120] Compared to traditional single-machine relational databases, distributed relational database services offer significant advantages. First, they can increase data processing capacity and volume by adding nodes, avoiding the capacity and performance bottlenecks of traditional databases. Second, distributed relational database services offer better scalability, easily supporting data synchronization and replication across multiple database servers, thus achieving high availability and disaster recovery capabilities. Finally, the cost of using distributed relational database services is relatively low because multiple nodes can share hardware resources and network bandwidth, resulting in a lower overall cost of ownership.

[0121] However, existing distributed relational databases present certain complexities in troubleshooting. When problems arise, relevant personnel typically need to manually check and test various logs and processes within the current environment to pinpoint the issue. Due to the architectural characteristics of distributed databases, the numerous and complex metrics requiring troubleshooting make this process extremely time-consuming and labor-intensive, severely impacting the stable operation of production applications.

[0122] Therefore, the urgent technical problem to be solved is how to accurately detect operational failures in distributed relational databases so that problems can be quickly found and resolved when they occur, thereby improving the stability and operational efficiency of distributed relational databases.

[0123] The fault detection method based on DRDS provided in this application, upon obtaining multiple state information and multiple performance information corresponding to the distributed relational database to be detected, sequentially detects these multiple state information and multiple performance information to obtain multiple state detection results corresponding to the multiple state information and multiple performance detection results corresponding to the multiple performance information. Then, based on the multiple state detection results and multiple performance detection results, a fault detection report is generated. This method solves the problem of inaccurate detection of operational faults in distributed relational databases. It can not only more accurately determine the database's operational status and performance, but also reduce the cost and error rate of manual operation, thereby improving the accuracy of fault detection. This allows for more rational allocation of system resources, improving the performance and reliability of the distributed relational database.

[0124] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0125] Figure 1 This embodiment provides a flowchart of a fault detection method based on DRDS. Figure 1 .like Figure 1 As shown in the figure, this embodiment illustrates a fault detection method based on DRDS, including:

[0126] S101: Obtain the distributed relational database to be tested, which is used to indicate that there is an anomaly in the distributed relational database during operation.

[0127] In this case, obtaining the distributed relational database to be tested means that the current distributed relational database has encountered an anomaly during operation, which leads to a decrease in the stability and operating efficiency of the distributed relational database.

[0128] Understandably, distributed relational databases consist of multiple nodes (i.e., database servers), which means the probability of failure is relatively high. Each node may encounter hardware, software, or network failures during operation, all of which can affect the overall performance and availability of the distributed relational database.

[0129] Meanwhile, distributed relational databases need to handle complex distributed transactions and data consistency issues during operation, which brings many potential problems to data replication, sharding and synchronization in multi-node environments. These problems will also affect the overall performance and availability of distributed relational databases.

[0130] Therefore, in order to accurately locate the abnormal problems that occur, it is necessary to obtain the distributed relational database that malfunctions during operation, so as to accurately locate the fault in the distributed relational database.

[0131] The methods for obtaining the distributed relational database to be tested include, but are not limited to, obtaining it through JDBC.

[0132] S102: Based on the distributed relational database to be detected, determine the operating parameter information of the distributed relational database to be detected. The operating parameter information includes: status information and performance information. The status information includes multiple sub-candidate status information, and the performance information includes multiple sub-candidate performance information.

[0133] The status information refers to the state of the distributed relational database during operation, including multiple sub-candidate status information. These sub-candidate status information include, for example, whether the distributed relational database is properly connected, whether the distributed relational database is running normally, whether there are any faults, and the status of transaction processing.

[0134] Performance information refers to the performance of a distributed relational database during operation, including multiple sub-candidate performance information. These sub-candidate performance information include, for example, the response time, throughput, CPU usage, and memory usage of the distributed relational database.

[0135] To accurately pinpoint the cause of a failure in a distributed relational database, it is necessary to determine the corresponding operational parameters, including status and performance information. These parameters provide crucial information about the database's status and performance during operation, facilitating accurate analysis and troubleshooting of anomalies.

[0136] As is understandable, distributed relational databases are composed of multiple nodes. Therefore, when a fault is found in the operation of a distributed relational database, in order to accurately find the fault, it is necessary to determine the corresponding operating parameter information of the distributed relational database, such as node failure, network communication failure, data inconsistency, etc., so as to accurately analyze these parameter information.

[0137] For example, by acquiring and analyzing the connection parameters of a distributed relational database, information such as the number of connections and connection status can be obtained, thus enabling an accurate assessment of the database's availability and performance. Similarly, by acquiring and analyzing transaction processing parameters, information such as the database's transaction processing capabilities, response time, and throughput can be obtained, thus enabling an accurate assessment of the database's performance and data processing capabilities.

[0138] S103: Perform state detection on the multiple sub-candidate state information in sequence to obtain multiple state detection results, and perform performance detection on the multiple sub-candidate performance information in sequence to obtain multiple performance detection results.

[0139] By sequentially detecting multiple sub-candidate state information and multiple sub-candidate performance information, the corresponding state detection results and performance detection results can be obtained. Based on the multiple state detection results and multiple performance detection results, the state performance and performance of the distributed relational database during operation can be determined.

[0140] Understandably, distributed relational databases, being database systems composed of multiple nodes, possess characteristics such as high availability, scalability, and fault tolerance. However, various problems may arise during the operation of distributed relational databases, such as node failures, network communication failures, and data inconsistencies. These problems can affect the performance and stability of the distributed relational database.

[0141] Therefore, by sequentially detecting multiple sub-candidate status information and multiple sub-candidate performance information, it means that various parameters of the distributed relational database during operation can be detected, such as connection information, lock management information, long connection information, transaction information, RDS operation information, DRDS-Server operation information, master-slave switchover information, slow SQL query information, etc., thereby determining the detection results corresponding to these parameter information.

[0142] S104: Generate a fault detection report based on the multiple state detection results and the multiple performance detection results.

[0143] The fault detection report is used to organize and display multiple sub-candidate status information, multiple status detection results corresponding to the multiple sub-candidate status information, multiple sub-candidate performance information, and multiple performance detection results corresponding to the multiple sub-candidate performance information, so that the database administrator can comprehensively evaluate and analyze the status and performance of the database during operation.

[0144] By generating fault detection reports, a comprehensive understanding of the status and performance of a distributed relational database during operation can be achieved, enabling timely identification and resolution of potential problems. This provides crucial reference and basis for the optimization and management of the distributed relational database. Furthermore, fault detection reports can serve as important documents for the maintenance and management of the distributed relational database, allowing for future review and analysis.

[0145] Optionally, once the distributed relational database to be tested is obtained, it can be periodically tested according to a preset cycle, and a fault detection report can be generated.

[0146] Understandably, since distributed relational databases are the core data storage and management platform of business systems, regular testing of distributed relational databases can promptly detect faults that occur during database operation, diagnose the causes and locations of these faults, and take appropriate remedial measures to reduce the impact of faults on business continuity.

[0147] Meanwhile, by regularly generating a fault detection report on the operational failures of the distributed relational database, database administrators can understand the performance status of the distributed relational database, enabling them to perform targeted system optimization and adjustments, thereby improving the overall performance and response speed of the distributed relational database.

[0148] The DRDS-based fault detection method provided in this embodiment first acquires the distributed relational database to be detected. Then, based on the distributed relational database, it determines multiple state and performance information of the database. Next, it sequentially detects the multiple state and performance information, obtaining state detection results corresponding to the multiple state information and performance detection results corresponding to the multiple performance information. Finally, it generates a fault detection report based on the multiple state and performance detection results. This method solves the problem of inaccurate detection of operational faults in distributed relational databases. It can not only more accurately determine the database's operational status and performance but also reduce the cost and error rate of manual operation, thereby improving the accuracy of fault detection. This allows for more rational allocation of system resources, improving the performance and reliability of the distributed relational database.

[0149] Figure 2This is the flowchart of the DRDS-based fault detection method provided in this embodiment. Figure 2 This embodiment is... Figure 1 Based on the embodiments, the specific implementation process of the DRDS-based fault detection method is described in detail. Figure 2 As shown in the figure, the fault detection method based on DRDS illustrated in this embodiment includes:

[0150] S201: Obtain the distributed relational database to be tested, which is used to indicate that there is an anomaly in the distributed relational database during operation.

[0151] The explanation of step S201 is similar to that of step S101 above, and will not be repeated here.

[0152] S202: Based on the distributed relational database to be detected, determine the operating parameter information of the distributed relational database to be detected. The operating parameter information includes: status information and performance information. The status information includes multiple sub-candidate status information, and the performance information includes multiple sub-candidate performance information.

[0153] The explanation of step S202 is similar to that of step S102 above, and will not be repeated here.

[0154] S203: The sub-candidate state information includes: link information, lock management information, long link information, and transaction information. If the link information is subjected to state detection, the state detection result corresponding to the link information is determined.

[0155] By checking the status of the connection information, the connection status can be determined, including whether the connection was successful or failed. For successful connections, the status check result will typically show a normal connection status, indicating that the database server can receive and process requests normally. For failed connections, the status check result will typically show an abnormal connection status, indicating that the database server cannot receive and process requests normally.

[0156] Understandably, if the connection status shows "connection successful," it confirms that the distributed relational database server can receive and process requests normally. At this point, further status checks can be performed on this successful connection status to obtain the corresponding results. For example, metrics such as the duration of the successful connection and the number of connections can be checked to assess the database server's operating status and load.

[0157] Conversely, if the connection status indicates a connection failure, it means the distributed relational database server is unable to receive and process requests normally. In this case, status detection is needed to determine the corresponding detection result for the connection failure status. Connection failures can be caused by various reasons, such as network failures, database server failures, and user permission issues. By detecting the connection failure status, these problems can be further investigated and resolved.

[0158] Therefore, when an anomaly is detected in a distributed relational database, it is necessary to first check the connection information and perform corresponding status checks in order to quickly locate the cause of the problem and prevent the failure from spreading, thereby improving the availability and stability of the distributed relational database.

[0159] S204: If a status detection is performed on the lock management information, a status detection result corresponding to the lock management information is determined. The lock management information includes: lock information and lock waiting information.

[0160] By performing status detection on lock management information, abnormal situations in distributed relational databases can be detected and resolved in a timely manner, thereby preventing the occurrence and spread of faults and improving the availability and stability of the entire distributed relational database.

[0161] Understandably, lock management in distributed relational databases is a critical component of database operation. It coordinates access to shared resources by multiple transactions to ensure data consistency and integrity. Abnormalities in lock management can lead to transaction errors, data inconsistencies, or deadlocks, thus affecting the normal operation of the entire system.

[0162] Therefore, by performing status checks on lock management information, the corresponding status check results can be quickly obtained. For example, if the status check result indicates that the current distributed relational database may have locks or abnormal transactions, then the problem can be investigated specifically, the specific cause of lock waiting can be identified, and the problem can be resolved in a timely manner, thereby preventing the occurrence and spread of faults and improving the availability and stability of the entire system.

[0163] S205: If the long link information is subjected to status detection, the status detection result corresponding to the long link information is determined, wherein the long link information includes: execution time and execution status.

[0164] Execution time refers to the time required to execute operations on a long-lived connection, reflecting its execution efficiency. By monitoring execution time, the performance of long-lived connections can be understood, and potential performance issues can be identified and resolved promptly. For example, if excessively long execution times are detected, it may be necessary to optimize the long-lived connection execution process or upgrade the database server to improve its processing capacity.

[0165] Execution status refers to the current state of a long-lived link, such as whether it is active or has encountered an error. By monitoring the execution status, we can understand the operation of long-lived links and promptly identify and resolve potential errors. For example, if an error is detected in a long-lived link, the cause of the error can be analyzed to ensure its normal operation.

[0166] By performing state checks on long-lived connection information, we can obtain corresponding state check results. These results can be used to evaluate the performance and operation of long-lived connections, thereby improving the performance of distributed relational databases and ultimately enhancing the availability and stability of the entire distributed relational database system.

[0167] Understandably, since long-lived connections are a key component in distributed relational databases for coordinating multiple transactions to access shared resources, their operation and performance have a significant impact on the stability and performance of the entire system.

[0168] Therefore, by monitoring the execution time of long-lived connection information, the execution efficiency of long-lived connections can be assessed. If the execution time is too long, it may indicate a bottleneck in the execution process. Therefore, optimization of the current long-lived connections is necessary to improve their execution efficiency, thereby reducing client waiting time and enhancing user experience.

[0169] By performing status checks on the execution status of long-lived connections, the execution status of those connections can be monitored. If a long-lived connection is in an error state, it may indicate an anomaly or error. Therefore, the current long-lived connection needs to be optimized to avoid service interruptions or other problems caused by errors.

[0170] S206: If a status detection is performed on the transaction information, then the status detection result corresponding to the transaction information is determined.

[0171] By performing status checks on transaction information, information such as the processing status and execution status of transactions can be obtained, thus avoiding problems such as deadlocks and timeouts in the transaction processing of distributed relational databases.

[0172] Understandably, in distributed relational databases, a transaction is a process of executing a series of database operations, which either all succeed or all are rolled back. However, during transaction processing, problems such as deadlocks and timeouts may occur, preventing transactions from executing normally and thus affecting the availability and consistency of the entire system.

[0173] Therefore, by performing status checks on transaction information, these problems can be identified and resolved in a timely manner, avoiding any impact on the normal operation of the distributed relational database and improving the availability and consistency of the entire distributed relational database system.

[0174] For example, suppose there is an e-commerce website where users can buy goods and submit orders. When a user submits an order, the website creates a transaction to perform a series of operations, including writing the order information to the database, updating inventory, and sending a confirmation email.

[0175] In this transaction, the order information is first written, then the inventory quantity is reduced, and finally a confirmation email is sent to the user. If all these operations are executed successfully, the transaction is committed, and the order is confirmed. However, if a deadlock or timeout occurs during the operation, the transaction needs to be rolled back, the order information is undone, and the inventory quantity is restored to its original state.

[0176] Optionally, after obtaining the transaction information, deduplication should also be performed on the transaction information. The purpose is to ensure the consistency of data in the distributed relational database and improve the performance and efficiency of the distributed relational database.

[0177] Understandably, when it is determined that an anomaly has occurred in the operation of a distributed relational database, and transaction information is obtained, the troubleshooting process becomes complex and time-consuming because each transaction's information is recorded separately. Therefore, by deduplicating transaction information, similar transactions can be grouped into the same category, so as to find the potentially problematic areas more quickly and thus find and solve the problem faster.

[0178] S207: The sub-candidate performance information includes: the operation information of multiple RDS, the operation information of DRDS-Server, the master-slave switchover information, and the slow SQL query information. If performance testing is performed on the operation information of the multiple RDS, the performance testing result corresponding to the operation information of each RDS is determined according to the first anomaly detection model and the operation information of each RDS.

[0179] The first anomaly detection model is used to detect anomalies in each RDS instance. For example, if the response time of an RDS system suddenly slows down, the first anomaly detection model can detect this anomaly and assess and address it.

[0180] By performing performance testing on the operational information of multiple RDSs and combining it with the first anomaly detection model to determine the performance test results of each RDS, we can gain a more accurate understanding of the performance and bottleneck information of each RDS.

[0181] Understandably, the first anomaly detection model is used to monitor and evaluate various performance metrics of each RDS system during operation. These metrics include, but are not limited to, response time, throughput, CPU utilization, memory utilization, and disk utilization. Therefore, by detecting and analyzing these metrics, we can understand the performance characteristics and bottlenecks of the RDS system, enabling timely detection and resolution of performance issues and ensuring the stability and availability of the distributed relational database system.

[0182] Optionally, the specific implementation process of the first anomaly detection model may be as follows: based on the operation information of the multiple RDS, determine the first historical monitoring indicator dataset corresponding to the operation information of each RDS; classify and summarize the multiple first historical monitoring indicator datasets to obtain multiple candidate monitoring indicator datasets with the same category; and use the isolated forest algorithm to train the multiple candidate monitoring indicator datasets to obtain the first anomaly detection model.

[0183] Each RDS instance generates various operational data during its operation, such as response time, throughput, CPU utilization, and memory utilization. This data is crucial for evaluating RDS performance and monitoring system status. Therefore, to dynamically monitor this data, it's necessary to construct a historical monitoring metric dataset corresponding to each RDS instance's operational data. These historical monitoring metric datasets can be used to analyze RDS performance and behavioral patterns, thereby enabling timely identification and resolution of potential problems and optimization of system configuration and capacity planning.

[0184] By classifying and summarizing the historical monitoring indicator datasets of each RDS, datasets with the same characteristics can be grouped into a unified category, so as to better identify the differences between different categories.

[0185] When multiple candidate monitoring indicator datasets are obtained, the Isolation Forest algorithm is used to train each candidate monitoring indicator dataset to obtain an anomaly detection model with multiple monitoring indicators. This model can process various monitoring indicator data, such as response time, throughput, CPU utilization, and memory utilization, so as to more comprehensively evaluate the performance and behavior patterns of RDS.

[0186] S208: If performance testing is performed on the DRDS-Server's operating information, a second anomaly detection model is used to perform performance testing on the DRDS-Server's operating information to obtain a performance testing result corresponding to the DRDS-Server's operating information. The second anomaly detection model is different from the first anomaly detection model.

[0187] By performing performance testing on the DRDS-Server's operational information, metrics such as CPU utilization and memory utilization corresponding to the DRDS-Server can be obtained, along with performance test results. These results can then be used to evaluate the DRDS-Server's performance and monitor the system's operational status.

[0188] Understandably, by obtaining the corresponding metrics for the DRDS-Server, one can understand its performance and resource utilization under different load conditions. This information is used to evaluate the performance level of the DRDS-Server and monitor the system's operational status. For example, excessively high CPU or memory usage may lead to slow system response or resource exhaustion, impacting the database system's performance.

[0189] Therefore, by performing performance testing on the DRDS-Server's operational information, we can better evaluate the DRDS-Server's performance and monitor the operational status of the distributed relational data system.

[0190] Optionally, the specific implementation process of the second anomaly detection model may be as follows: based on the operation information of the DRDS-Server, determine the second historical monitoring indicator dataset corresponding to the operation information of the DRDS-Server; use the isolated forest algorithm to train the second historical monitoring indicator dataset to obtain the second anomaly detection model.

[0191] The second historical monitoring metrics dataset is used to indicate the historical performance data of the DRDS-Server. This performance data includes several historical metrics such as response time, throughput, CPU utilization, and memory utilization.

[0192] By analyzing the DRDS-Server's operational information, you can obtain a historical monitoring metrics dataset corresponding to this operational information. This dataset contains various performance metrics generated by the DRDS-Server during its past operation, such as response time, throughput, CPU utilization, and memory utilization.

[0193] By using the isolated forest algorithm to train on a historical monitoring indicator dataset, an anomaly detection model with multiple monitoring indicators can be obtained. This model can process various monitoring indicator data, such as response time, throughput, CPU utilization, and memory utilization, in order to more comprehensively evaluate the performance and behavior patterns of the DRDS-Server.

[0194] Understandably, since Isolation Forest is an unsupervised learning algorithm that can be used for anomaly detection and classification, when a historical monitoring index dataset corresponding to the DRDS-Server's operational information is obtained, the Isolation Forest algorithm can be used to train on this dataset to construct an algorithm for detecting anomalies in the current operational information of the DRDS-Server. For example, when operational information that does not match historical data is detected, the anomaly detection model can identify it as an anomaly and record it.

[0195] S209: If performance testing is performed on the primary / backup switching information, then the performance testing result corresponding to the primary / backup switching information is determined based on the primary / backup switching information.

[0196] By performing performance testing on the primary / standby switchover information, we can obtain switchover information during the switchover process, such as switchover time, response time, data replication, and transmission status. Based on this switchover information, we can determine whether the database has undergone a fault migration operation.

[0197] Understandably, in distributed relational database systems, master-slave failover is used to ensure high availability and continuity of the database. When the master node fails, the system automatically promotes the slave node to master to ensure service continuity. During master-slave failover, fault migration is used to migrate the data and state of the failed node to other healthy nodes.

[0198] Therefore, by querying the primary / standby switchover status of the database, switchover information related to the switchover process can be obtained, such as switchover time, response time, data replication and transmission status, etc. In order to determine whether the database has undergone a fault migration operation by analyzing this switchover information.

[0199] S210: If performance testing is performed on the slow SQL query information, then the performance testing result corresponding to the slow SQL query information is determined based on the slow SQL query information.

[0200] By performing performance testing on slow SQL queries, it can be determined that the distributed relational database has multiple time-consuming SQL statements during operation, and thus the performance test results corresponding to these multiple time-consuming SQL statements can be identified.

[0201] Understandably, since data in a distributed relational database is scattered across multiple nodes, queries require cross-node communication and data aggregation. If the data table does not have appropriate indexes, or the index design is unreasonable, or even if the SQL query data design is unreasonable, these factors can lead to the SQL query statement requiring a full table scan or a large number of disk I / O operations, resulting in excessively long SQL query execution times.

[0202] Therefore, by performing performance testing on slow SQL queries, it is possible to identify multiple time-consuming SQL statements in the distributed relational database during its operation and determine the corresponding test results. These results can include performance metrics such as execution time, response time, and resource utilization. Based on the test results, the operational status of the distributed relational database can be determined.

[0203] S211: Generate a fault detection report based on the multiple state detection results and the multiple performance detection results.

[0204] The explanation of step S211 is similar to that of step S211 above, and will not be repeated here.

[0205] The DRDS-based fault detection method provided in this application, after obtaining the distributed relational database to be detected, firstly determines multiple sub-candidate state information and multiple sub-candidate performance information of the distributed relational database to be detected. Secondly, it sequentially detects the multiple sub-candidate state information and multiple sub-candidate performance information until it obtains the state detection result corresponding to each sub-candidate state information and the performance detection result corresponding to each sub-candidate performance information. Finally, it jointly determines the fault monitoring report of the distributed relational database based on the multiple state detection results and multiple performance detection results. This method can detect and analyze distributed relational databases in an automated manner, thereby reducing the influence of manual intervention and subjective judgment, and improving the objectivity and accuracy of detection. At the same time, by providing a detailed detection report, this method can provide database administrators with more comprehensive information, helping them to better understand the performance status and fault conditions of the distributed relational database, thereby better optimizing the distributed relational database.

[0206] Figure 3 This is the flowchart of the DRDS-based fault detection method provided in this embodiment. Figure 3 This embodiment is... Figure 2Based on the embodiments, the specific implementation process of determining the performance test result corresponding to the slow SQL query information based on the slow SQL query information, and the specific implementation process after obtaining the performance test result corresponding to the slow SQL query information, will be described in detail together. For example... Figure 3 As shown in the figure, the fault detection method based on DRDS illustrated in this embodiment includes:

[0207] S301: The slow SQL query information is parsed and processed to obtain multiple characteristic slow SQL query statements.

[0208] In order to obtain the data characteristics of each SQL data in the slow SQL query information, the slow SQL query information needs to be parsed and processed until multiple characteristic slow SQL query statements are obtained.

[0209] Understandably, because data in a distributed relational database is scattered across multiple nodes, queries require cross-node communication and data aggregation, which leads to additional network latency and computational overhead. Especially when processing large amounts of data, the execution time of SQL queries can become quite long.

[0210] At the same time, the complexity of SQL queries can also affect query performance. For example, queries involving multiple table joins, subqueries, or aggregate functions usually require more computation and communication operations, so the execution time may be longer.

[0211] Therefore, in order to better understand and optimize the performance of slow SQL queries, it is necessary to parse and process the slow SQL query information until the data characteristics of each SQL query statement are extracted, so as to quickly identify which SQL query statements are the main causes of database performance degradation and the bottlenecks of these SQL query statements.

[0212] S302: Perform feature merging processing on the multiple slow SQL query statements with the same characteristics to obtain a set of multiple slow SQL query statements with the same characteristics.

[0213] In order to simplify a large number of slow SQL queries into a few sets of statements with the same characteristics, feature merging is required for slow SQL queries with multiple characteristics after obtaining the data characteristics of each SQL data.

[0214] Understandably, the slow SQL query sets obtained after feature merging usually have similar or identical characteristics, which means that these slow SQL query sets with the same characteristics may have similar problems in terms of query conditions, table join methods, or index usage.

[0215] For example, if after feature merging it is found that multiple slow SQL queries use the same query conditions and join methods, but the indexes are used improperly, then the indexes can be optimized to speed up the query.

[0216] S303: Based on the multiple sets of slow SQL query statements, determine the sub-performance test result corresponding to each set of slow SQL query statements.

[0217] Since multiple slow SQL query sets typically share the same problem, multiple slow SQL query sets can be analyzed directly to more comprehensively evaluate the performance of distributed relational databases.

[0218] Understandably, in distributed relational databases, slow SQL queries can be caused by factors such as query optimization, data distribution, or network latency. Therefore, when multiple sets of slow SQL queries are obtained, they can be analyzed together to more quickly and comprehensively determine the performance status of the distributed relational database.

[0219] S304: Based on the sub-performance test results corresponding to each slow SQL query statement set, determine the performance test result corresponding to the slow SQL query information.

[0220] After obtaining the sub-performance test results for each set of slow SQL queries, these results need to be organized and analyzed to assess the overall performance impact of the current slow SQL queries on the distributed relational database.

[0221] Understandably, since the performance test results for each set of slow SQL queries may include performance metrics such as execution time, CPU utilization, and memory usage, organizing and analyzing these results allows us to determine the performance impact of each set of slow SQL queries on the entire distributed relational database.

[0222] S305: Parse and process the multiple slow SQL query statement sets to obtain the problematic slow SQL query statements.

[0223] Slow SQL statements with similar characteristics may share common performance bottlenecks. Therefore, by parsing and processing these sets of slow SQL statements, the bottlenecks in database performance can be identified more quickly, allowing for targeted optimization.

[0224] Understandably, slow SQL statements with similar characteristics may encounter the same problems during execution, such as using inappropriate indexes, overly complex query conditions, or excessive data volume. Therefore, by parsing and processing multiple slow SQL query sets, the most representative problematic slow SQL query statements can be obtained, such as those with many executions or long execution times, so that these representative problematic slow SQL query statements can be optimized.

[0225] S306: Use the diagnostic processing rule set to diagnose and process the problematic slow SQL query statement to obtain the target solution corresponding to the problematic slow SQL query statement.

[0226] The diagnostic processing rule set includes: heuristic diagnostic suggestion rules, index optimization suggestion rules, and execution plan interpretation rules. This rule set is used to identify and analyze problematic slow SQL queries and provide corresponding solutions. These rule sets may include heuristic diagnostic suggestion rules, index optimization suggestion rules, and execution plan interpretation rules, etc.

[0227] Heuristic diagnostic recommendations include, for example, performing selection operations as early as possible, performing projection and selection operations simultaneously, combining projection with the preceding and following binocular operations, combining certain selections with the Cartesian product to be performed before them into a single join operation, whether to shard the database and tables, and whether the use of sharding sequences is reasonable.

[0228] Index optimization suggestions include, for example, whether there is an index, whether the index is invalid, and whether the index is reasonable.

[0229] Execution plan interpretation refers to using MySQL's built-in tools to execute the plan and simulate the execution of SQL query statements in order to analyze the execution information and performance bottlenecks of the SQL query statements.

[0230] In order to quickly repair slow SQL queries and prevent abnormalities from occurring during the operation of distributed relational databases, it is necessary to use a diagnostic processing rule set to diagnose and process slow SQL queries until a method that can solve the problem is found.

[0231] Understandably, by using diagnostic processing rule sets to diagnose and process slow SQL queries, database performance issues can be identified and resolved more quickly, preventing anomalies in distributed relational databases during operation and thus improving system reliability and stability.

[0232] The fault detection method based on DRDS provided in this application involves parsing the obtained slow SQL query statements to obtain multiple characteristic slow SQL query statements. First, the characteristics of these multiple characteristic slow SQL query statements are merged to obtain multiple sets of slow SQL query statements with the same characteristics. Second, based on these sets, the performance detection result corresponding to each set of slow SQL query statements is determined. Simultaneously, these sets are parsed to obtain problematic slow SQL query statements. Finally, a diagnostic processing rule set is used to diagnose the problematic slow SQL query statements, resulting in multiple target solutions corresponding to them. This method not only provides a more comprehensive understanding of the performance status of distributed relational databases but also quickly locates problems, thereby shortening the time for diagnosis and problem-solving, and ultimately improving the reliability and stability of distributed relational databases.

[0233] Figure 4 This is a schematic diagram of the DRDS-based fault detection device provided in this application. Figure 4 As shown, this application provides a DRDS-based fault detection device, the DRDS-based fault detection device 400 comprising:

[0234] The acquisition module 401 is used to acquire the distributed relational database to be detected, wherein the distributed relational database to be detected is used to indicate that there is an anomaly in the distributed relational database during operation;

[0235] The determining module 402 is used to determine the operating parameter information of the distributed relational database to be detected based on the distributed relational database to be detected. The operating parameter information includes: status information and performance information. The status information includes multiple sub-candidate status information, and the performance information includes multiple sub-candidate performance information.

[0236] The processing module 403 is used to sequentially perform state detection on the plurality of sub-candidate state information to obtain a plurality of state detection results, and sequentially perform performance detection on the plurality of sub-candidate performance information to obtain a plurality of performance detection results;

[0237] The generation module 404 is used to generate a fault detection report based on the multiple state detection results and the multiple performance detection results.

[0238] Optionally, the determining module 402 is further configured to determine the state detection result corresponding to the link information when performing state detection on the link information;

[0239] and / or;

[0240] The determining module 402 is further configured to determine the state detection result corresponding to the lock management information when performing state detection on the lock management information, wherein the lock management information includes: lock information and lock waiting information;

[0241] and / or;

[0242] The determining module 402 is further configured to determine the status detection result corresponding to the long link information when performing status detection on the long link information, wherein the long link information includes: execution time and execution status;

[0243] and / or;

[0244] The determining module 402 is further configured to determine the status detection result corresponding to the transaction information when performing status detection on the transaction information.

[0245] Optionally, the determining module 402 is further configured to, when performing performance testing on the running information of the plurality of RDS, determine the performance testing result corresponding to the running information of each RDS based on the first anomaly detection model and the running information of each RDS;

[0246] and / or;

[0247] The determining module 402 is further configured to, when performing performance testing on the DRDS-Server's operating information, use a second anomaly detection model to perform performance testing on the DRDS-Server's operating information, and obtain a performance testing result corresponding to the DRDS-Server's operating information, wherein the second anomaly detection model is different from the first anomaly detection model;

[0248] and / or;

[0249] The determining module 402 is further configured to, when performing performance testing on the primary / backup switching information, determine the performance testing result corresponding to the primary / backup switching information based on the primary / backup switching information;

[0250] and / or;

[0251] The determining module 402 is further configured to, when performing performance testing on the slow SQL query information, determine the performance testing result corresponding to the slow SQL query information based on the slow SQL query information.

[0252] Optionally, the determining module 402 is further configured to determine a first historical monitoring indicator dataset corresponding to each RDS operation information based on the operation information of the plurality of RDSs;

[0253] The processing module 403 is also used to classify and summarize the multiple first historical monitoring indicator datasets to obtain multiple candidate monitoring indicator datasets with the same category.

[0254] The device further includes: a training module 405;

[0255] The training module 405 is used to train the multiple candidate monitoring indicator datasets using the isolated forest algorithm to obtain the first anomaly detection model.

[0256] Optionally, the determining module 402 is further configured to determine a second historical monitoring indicator dataset corresponding to the operation information of the DRDS-Server based on the operation information of the DRDS-Server;

[0257] The training module 405 is also used to train the second historical monitoring index dataset using the isolated forest algorithm to obtain the second anomaly detection model.

[0258] Optionally, the processing module 403 is further configured to parse the slow SQL query information to obtain multiple characteristic slow SQL query statements;

[0259] The processing module 403 is also used to perform feature merging processing on the multiple slow SQL query statements with the same features to obtain a set of multiple slow SQL query statements with the same features.

[0260] The determining module 402 is further configured to determine a sub-performance detection result corresponding to each slow SQL query statement set based on the plurality of slow SQL query statement sets;

[0261] The determining module 402 is specifically used to determine the performance test result corresponding to the slow SQL query information based on the sub-performance test result corresponding to each slow SQL query statement set.

[0262] Optionally, the processing module 403 is further configured to parse the multiple slow SQL query statement sets to obtain the problematic slow SQL query statements;

[0263] The processing module 403 is further configured to perform diagnostic processing on the problematic slow SQL query statement using a diagnostic processing rule set to obtain a target solution corresponding to the problematic slow SQL query statement. The diagnostic processing rule set includes: heuristic diagnostic suggestion rules, index optimization suggestion rules, and execution plan interpretation rules.

[0264] Figure 5 This is a schematic diagram of the DRDS-based fault detection device provided in this application. Figure 5As shown, this application provides a DRDS-based fault detection device 500, which includes a receiver 501, a transmitter 502, a processor 503, and a memory 504.

[0265] Receiver 501 is used to receive instructions and data;

[0266] Transmitter 502 is used to send commands and data;

[0267] Memory 504 is used to store instructions executed by the computer;

[0268] The processor 503 is used to execute computer execution instructions stored in the memory 504 to implement the various steps performed by the multimodal pre-trained model training method in the above embodiments. For details, please refer to the relevant descriptions in the foregoing embodiments of the multimodal pre-trained model training method.

[0269] Alternatively, the memory 504 can be either standalone or integrated with the processor 503.

[0270] When the memory 504 is set up independently, the electronic device also includes a bus for connecting the memory 504 and the processor 503.

[0271] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the training method of the multimodal pre-trained model as described above, executed by the training device for the multimodal pre-trained model.

[0272] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0273] The technical solutions of this application have been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it is readily understood by those skilled in the art that the scope of protection of this application is obviously not limited to these specific embodiments. The above embodiments are only used to illustrate the technical solutions of this application and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A fault detection method based on DRDS, characterized in that, The method includes: Obtain the distributed relational database to be tested, which is used to indicate that there is an anomaly in the distributed relational database during operation; Based on the distributed relational database to be detected, the operating parameter information of the distributed relational database to be detected is determined. The operating parameter information includes: status information and performance information. The status information includes multiple sub-candidate status information, and the performance information includes multiple sub-candidate performance information. The state detection is performed on the multiple sub-candidate state information in sequence to obtain multiple state detection results, and the performance detection is performed on the multiple sub-candidate performance information in sequence to obtain multiple performance detection results; A fault detection report is generated based on the multiple status detection results and the multiple performance detection results.

2. The method according to claim 1, characterized in that, The sub-candidate state information includes: link information, lock management information, long-connection information, and transaction information. The process of sequentially performing state detection on the multiple sub-candidate state information yields multiple state detection results, including: If a status detection is performed on the link information, then the status detection result corresponding to the link information is determined; and / or; If a status detection is performed on the lock management information, a status detection result corresponding to the lock management information is determined. The lock management information includes: lock information and lock waiting information. and / or; If a status detection is performed on the long link information, a status detection result corresponding to the long link information is determined. The long link information includes: execution time and execution status. and / or; If a status check is performed on the transaction information, the status check result corresponding to the transaction information is determined.

3. The method according to claim 1, characterized in that, The sub-candidate performance information includes: operational information of multiple RDS instances, operational information of the DRDS-Server, master-slave failover information, and slow SQL query information. The performance of these multiple sub-candidate performance information is then sequentially tested to obtain multiple performance test results, including: If performance testing is performed on the operation information of the multiple RDS, then based on the first anomaly detection model and the operation information of each RDS, the performance testing result corresponding to the operation information of each RDS is determined. and / or; If the DRDS-Server's operating information is to be tested for performance, a second anomaly detection model is used to test the DRDS-Server's operating information to obtain the performance test result corresponding to the DRDS-Server's operating information. The second anomaly detection model is different from the first anomaly detection model. and / or; If performance testing is performed on the primary / backup switchover information, then the performance testing result corresponding to the primary / backup switchover information is determined based on the primary / backup switchover information. and / or; If performance testing is performed on the slow SQL query information, then 301 determines the performance testing result corresponding to the slow SQL query information based on the slow SQL query information.

4. The method according to claim 3, characterized in that, Before determining the performance detection result corresponding to the operating information of each RDS based on the first anomaly detection model and the operating information of each RDS, the method further includes: Based on the operational information of the multiple RDSs, determine the first historical monitoring indicator dataset corresponding to the operational information of each RDS; The multiple first historical monitoring indicator datasets are classified and summarized to obtain multiple candidate monitoring indicator datasets with the same category; The first anomaly detection model is obtained by training the multiple candidate monitoring indicator datasets using the isolated forest algorithm.

5. The method according to claim 3, characterized in that, Before performing performance testing on the DRDS-Server's operational information using the second anomaly detection model to obtain the performance testing results corresponding to the DRDS-Server's operational information, the method further includes: Based on the DRDS-Server's operating information, determine the second historical monitoring indicator dataset corresponding to the DRDS-Server's operating information; The isolated forest algorithm is used to train the second historical monitoring index dataset to obtain the second anomaly detection model.

6. The method according to claim 3, characterized in that, The step of determining the performance test result corresponding to the slow SQL query information based on the slow SQL query information includes: The slow SQL query information is parsed and processed to obtain multiple characteristic slow SQL query statements; The multiple slow SQL query statements with the same characteristics are subjected to feature merging processing to obtain a set of multiple slow SQL query statements with the same characteristics; Based on the multiple sets of slow SQL query statements, determine the sub-performance test result corresponding to each set of slow SQL query statements; Based on the sub-performance test results corresponding to each set of slow SQL query statements, determine the performance test results corresponding to the slow SQL query information.

7. The method according to claim 6, characterized in that, The method further includes: The multiple sets of slow SQL query statements are parsed and processed to obtain the problematic slow SQL query statements; The problematic slow SQL query is diagnosed and processed using a set of diagnostic processing rules to obtain the target solution corresponding to the problematic slow SQL query. The set of diagnostic processing rules includes: heuristic diagnostic suggestion rules, index optimization suggestion rules, and execution plan interpretation rules.

8. A fault detection device based on DRDS, characterized in that, include: The acquisition module is used to acquire the distributed relational database to be detected, wherein the distributed relational database to be detected is used to indicate that there is an anomaly in the distributed relational database during operation; The determination module is used to determine the operating parameter information of the distributed relational database to be detected based on the distributed relational database to be detected. The operating parameter information includes: status information and performance information. The status information includes multiple sub-candidate status information, and the performance information includes multiple sub-candidate performance information. The processing module is used to sequentially perform state detection on the multiple sub-candidate state information to obtain multiple state detection results, and sequentially perform performance detection on the multiple sub-candidate performance information to obtain multiple performance detection results. The generation module is used to generate a fault detection report based on the multiple state detection results and the multiple performance detection results.

9. A fault detection device based on DRDS, characterized in that, include: Memory; processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the DRDS-based fault detection method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the DRDS-based fault detection method as described in any one of claims 1-7.