Automatic database operation and maintenance system based on intelligent monitoring and self-healing mechanism
The database automation operation and maintenance system with intelligent monitoring and self-healing mechanism solves the problems of real-time monitoring lag, low efficiency of manual fault handling and imperfect data backup in traditional database operation and maintenance methods, realizes efficient, stable and automated operation and maintenance of database operation, and reduces the risk of business interruption and data loss.
Patent Information
- Application Number
- CN202510769280.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-05
AI Technical Summary
Traditional database operation and maintenance methods have problems such as delayed real-time monitoring, inefficient manual troubleshooting, lack of intelligent performance optimization, and incomplete data backup, resulting in unstable database operation and high risk of business interruption.
A database automation operation and maintenance system based on intelligent monitoring and self-healing mechanisms is adopted, including real-time performance monitoring, anomaly detection and alarm, automated fault handling, intelligent indexing and query optimization, and backup and recovery management modules, to achieve real-time monitoring, automated fault repair and intelligent optimization of the database.
It achieves timely perception of database operation status and efficient fault repair, improves database performance and data security, reduces the risk of business interruption and data loss, and realizes full-process automation and standardization of database operation and maintenance.
Smart Images

Figure CN120596461A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of database management, and more specifically, to a database automation operation and maintenance system based on intelligent monitoring and self-healing mechanisms. Background Art
[0002] In today's digital age, databases are core components for data storage and management, and their stable operation is crucial. However, traditional database operation and maintenance methods have many drawbacks: Manual monitoring struggles to comprehensively and comprehensively capture database performance metrics in real time, and can hinder timely detection of potential risks. For example, when key metrics like the number of queries per second (QPS) and transactions per second (TPS) fluctuate abnormally, manual monitoring can lag behind and prevent timely detection.
[0003] Troubleshooting relies on manual intervention, which is inefficient and error-prone. Once common database failures such as deadlocks, delayed master-slave connections, and leaks occur, manual troubleshooting and repair often takes a significant amount of time, potentially leading to business interruptions and losses for the enterprise.
[0004] Performance optimization lacks intelligent methods. Traditionally, database index optimization and SQL query optimization rely on the experience of operations and maintenance personnel, making it difficult to achieve accuracy and efficiency, resulting in inadequate database performance.
[0005] Inadequate data backup and recovery management. Manual backups can lead to issues like untimely backups and unreasonable backup strategies. In exceptional circumstances, data recovery is complex and time-consuming, making it difficult to guarantee data security and integrity.
[0006] In order to solve the above problems, a technical solution is now provided. Summary of the Invention
[0007] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a database automatic operation and maintenance system based on intelligent monitoring and self-healing mechanism to solve the problems raised in the above-mentioned background technology.
[0008] To achieve the above object, the present invention provides the following technical solutions: An automated database operation and maintenance system based on intelligent monitoring and self-healing mechanisms, including a real-time performance monitoring module, an anomaly detection and alarm module, an automated fault handling module, an intelligent indexing and query optimization module, and a backup and recovery management module; The real-time performance monitoring module is used to collect the number of queries per second, number of transactions per second, number of connections, and slow query rate of the database. Through real-time collection and dynamic analysis of the number of queries per second, number of transactions per second, number of connections, and slow query rate, the health status of the database is evaluated. The anomaly detection and alerting module monitors various situations during database operation. For example, it uses the ARIMA time series forecasting model to predict normal fluctuations in key indicators and uses Isolation Forest to detect anomalies. When abnormal situations such as a sudden increase in the number of connections, lock wait timeouts, and insufficient disk space are identified, multi-channel alerts are immediately triggered. The automated fault handling module has a preset self-healing strategy library, which stores repair plans and execution processes by fault type and formulates corresponding repair strategies for common faults. The intelligent indexing and query optimization module is used to analyze the execution of SQL statements in the query log, automatically recommend index optimization solutions, and optimize inefficient SQL using syntax rewriting rules; The backup and recovery management module is used to automatically back up data regularly. It sets a reasonable backup cycle based on the size of the database and business needs, and triggers a rapid recovery mechanism in abnormal situations.
[0009] In a preferred embodiment, the real-time performance monitoring module operation specifically includes the following: Deploy a data collection agent on the database server to establish a connection with the database and collect the database's queries per second, transactions per second, number of connections, and slow query rate data at specific time intervals. The collected data is transmitted to the data processing center, and the sliding window algorithm is used to calculate the average value and standard deviation within a certain time window. If the current number of queries per second, number of transactions per second, number of connections, and slow query rate data deviate from the average value μ by more than a certain multiple of the standard deviation σ, it is determined that an abnormal fluctuation has occurred.
[0010] In a preferred embodiment, the operation of the anomaly detection and alarm module specifically includes the following: Obtain connection count, lock wait time, and disk space usage data from the real-time performance monitoring module and other relevant data sources of the database system, and perform integration and preprocessing; Analyze its historical data, learn the changing patterns of the number of connections in different time periods, establish a corresponding time series prediction model, and predict the fluctuation range of the number of connections under normal circumstances.
[0011] In a preferred embodiment, the operation of the anomaly detection and alarm module specifically includes the following: Based on historical data and normal patterns, we train ARIMA time series forecasting models and Isolation Forest anomaly detection models to identify abnormal patterns. We also set thresholds for connection counts, lock wait times, and disk space usage based on database configuration, business requirements, and historical operational data. Continuously monitor various indicator data in real time, comparing current data with set thresholds and normal fluctuation ranges obtained through time series analysis and anomaly detection models; When the number of connections exceeds the set threshold, or the lock waiting time exceeds the maximum allowed time, or the disk space usage reaches the set upper limit, it is determined to be an abnormal situation; Multiple alarm channels are pre-configured, including SMS interface, email sending module and in-site messaging system. Once an abnormal situation is detected, the alarm logic is immediately triggered to generate an alarm message containing the abnormal indicator name, current value, normal threshold range and the time when the abnormality occurred.
[0012] In a preferred embodiment, the operation of the automated fault handling module specifically includes the following: Detailed classification of common faults that may occur during database operation, such as deadlocks, connection leaks, and long transaction blocking, and development of detailed repair strategies for each type of fault; Regularly evaluate and update the self-healing strategy library, and promptly adjust and improve the corresponding repair strategies as the database system is upgraded, business needs change, and new fault types emerge.
[0013] In a preferred embodiment, the operation of the automated fault handling module specifically further includes the following: Works closely with the anomaly detection and alarm module, receives fault information sent by it in real time, and retrieves the corresponding repair strategy from the self-healing strategy library based on the received fault information; Automatically execute the corresponding repair operation according to the retrieved repair strategy. During the execution process, the repair progress and effect are monitored in real time. If the repair operation fails, the failure information is automatically recorded, including the failure reason and the steps executed. After the repair operation is completed, the relevant performance indicators and operating status of the database are checked and evaluated to determine whether the fault has been effectively resolved. The repair results and related data are fed back to the system to provide a basis for subsequent strategy optimization.
[0014] In a preferred embodiment, the operation of the intelligent indexing and query optimization module specifically includes the following: Use the SQL syntax analyzer to parse the extracted SQL statements. Based on the execution time, number of scanned rows, and number of returned rows in the query log, combined with the performance data of the database system, evaluate the execution efficiency of each SQL statement. Based on the structure of SQL statements and data distribution, it analyzes whether the existing indexes in the current database can meet the query requirements and automatically recommends appropriate index optimization solutions for each inefficient SQL statement. According to the characteristics of inefficient SQL statements and the identified problems, applicable semantically equivalent rewriting rules such as projection pushdown, predicate pushdown, subquery elimination, etc. are matched from the rule base to automatically rewrite the inefficient SQL.
[0015] In a preferred embodiment, the backup and recovery management module operation specifically includes the following: Comprehensively assess backup needs based on database size, business importance, and data update frequency, automatically execute backup tasks according to pre-set backup strategies, and monitor backup progress in real time, recording backup start and end times and backup data volume. Linked with the database monitoring system, it monitors the database operation status in real time and immediately triggers the recovery mechanism when database crash, accidental data deletion, hardware failure, or virus attack is detected; After the recovery process is started, the recovery operation is performed automatically according to the established plan or with the intervention of the operation and maintenance personnel. The database is quickly rebuilt using the backup data, and the data integrity and system consistency are verified after the recovery. The backup data is restored periodically to simulate real failure scenarios and verify the effectiveness of the backup data and recovery process.
[0016] The technical effects and advantages of the database automated operation and maintenance system based on intelligent monitoring and self-healing mechanism of the present invention are as follows: 1. The real-time performance monitoring module collects key performance indicators at high frequency and uses a sliding window algorithm to deeply analyze the data. It can capture subtle abnormal fluctuations in the database in milliseconds. Compared with traditional manual monitoring, it can not only detect potential risks hours or even days in advance, but also avoid monitoring blind spots caused by human negligence. This greatly improves the timeliness and accuracy of database operation status perception, saving valuable time for risk prevention and control. Combining threshold judgment and model prediction, it accurately identifies abnormal situations from multiple angles. The multi-channel alert mechanism ensures that operation and maintenance personnel, regardless of their location, can obtain alert information immediately, significantly shortening fault response time and reducing the duration of business interruption and the risk of data loss caused by faults. 2. Based on a pre-set self-healing strategy library, it can automatically execute repair operations the moment a fault occurs, eliminating the need for manual layer-by-layer troubleshooting and processing. For example, in the case of a deadlock fault, the system can complete transaction rollback within seconds and restore normal database operation. Through comprehensive analysis and in-depth optimization of SQL statements, database query performance can be significantly improved. 3. The differentiated backup strategies developed by the backup and recovery management module, combined with encrypted storage and off-site disaster recovery, provide a solid defense for data security. In the event of catastrophic events such as database crashes and accidental data deletion, the rapid recovery mechanism can restore the database to normal status in a short period of time, minimizing data loss. Regular recovery drills ensure the availability of backup data and the reliability of the recovery process, guaranteeing business continuity and avoiding business stagnation and financial losses caused by data issues. 4. The various modules of the system work closely together to achieve full-process automation of database operations and maintenance, from monitoring, alarms, troubleshooting, performance optimization, and data protection. This intelligent operation and maintenance model meets the needs of enterprise digital transformation, reduces reliance on manual experience, improves the standardization and regularization of operations and maintenance, and builds an efficient, stable, and secure database operating environment for enterprises, helping them enhance their core competitiveness in the digital age. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a schematic diagram of the structure of the database automatic operation and maintenance system based on intelligent monitoring and self-healing mechanism of the present invention. DETAILED DESCRIPTION
[0018] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0019] Example 1 Figure 1 The present invention provides a database automated operation and maintenance system based on intelligent monitoring and self-healing mechanism, including a real-time performance monitoring module, an anomaly detection and alarm module, an automated fault handling module, an intelligent indexing and query optimization module, and a backup and recovery management module; The real-time performance monitoring module is used to collect the number of queries per second, number of transactions per second, number of connections, and slow query rate of the database. Through real-time collection and dynamic analysis of the number of queries per second, number of transactions per second, number of connections, and slow query rate, the health status of the database is evaluated. The anomaly detection and alerting module monitors various situations during database operation. For example, it uses the ARIMA time series forecasting model to predict normal fluctuations in key indicators and uses Isolation Forest to detect anomalies. When abnormal situations such as a sudden increase in the number of connections, lock wait timeouts, and insufficient disk space are identified, multi-channel alerts are immediately triggered. The automated fault handling module has a preset self-healing strategy library, which stores repair plans and execution processes by fault type and formulates corresponding repair strategies for common faults. The intelligent indexing and query optimization module is used to analyze the execution of SQL statements in the query log, automatically recommend index optimization solutions, and optimize inefficient SQL using syntax rewriting rules; The backup and recovery management module is used to automatically back up data regularly. It sets a reasonable backup cycle based on the size of the database and business needs, and triggers a rapid recovery mechanism in abnormal situations.
[0020] The real-time performance monitoring module operation specifically includes the following: Deploy a data collection agent on the database server to establish a connection with the database and collect the database's queries per second, transactions per second, number of connections, and slow query rate data at specific time intervals. The collected data is transmitted to the data processing center, and the sliding window algorithm is used to calculate the average value and standard deviation within a certain time window. If the current number of queries per second, number of transactions per second, number of connections, and slow query rate data deviate from the average value μ by more than a certain multiple of the standard deviation σ, it is determined to be an abnormal fluctuation. The calculation formula is as follows: ; .
[0021] The real-time performance monitoring module serves as the "perception layer" of the system's operations. It first deploys a data collection agent on the database server, establishing a stable connection with the database. This agent then collects key metrics such as queries per second (QPS), transactions per second (TPS), number of connections, and slow query rate at high frequencies, at intervals of seconds or less. After the collected data is transmitted to the data processing center, it undergoes in-depth analysis using a sliding window algorithm. For example, the algorithm calculates the average and standard deviation of the QPS metric within a specific time window. If the current QPS value deviates from the average by more than a set multiple of the standard deviation, the system determines that the database has experienced abnormal fluctuations. The system then combines this with other metrics to comprehensively assess the database's health, providing the underlying data for subsequent anomaly detection.
[0022] The operation of the anomaly detection and alarm module specifically includes the following: Obtain connection count, lock wait time, and disk space usage data from the real-time performance monitoring module and other relevant data sources of the database system, and perform integration and preprocessing; Analyze its historical data, learn the changing patterns of the number of connections in different time periods, establish a corresponding time series prediction model, and predict the fluctuation range of the number of connections under normal circumstances.
[0023] The operation of the anomaly detection and alarm module specifically includes the following: Based on historical data and normal patterns, we train ARIMA time series forecasting models and Isolation Forest anomaly detection models to identify abnormal patterns. We also set thresholds for connection counts, lock wait times, and disk space usage based on database configuration, business requirements, and historical operational data. Continuously monitor various indicator data in real time, comparing current data with set thresholds and normal fluctuation ranges obtained through time series analysis and anomaly detection models; When the number of connections exceeds the set threshold, or the lock waiting time exceeds the maximum allowed time, or the disk space usage reaches the set upper limit, it is determined to be an abnormal situation; Multiple alarm channels are pre-configured, including SMS interface, email sending module and in-site messaging system. Once an abnormal situation is detected, the alarm logic is immediately triggered to generate an alarm message containing the abnormal indicator name, current value, normal threshold range and the time when the abnormality occurred.
[0024] The anomaly detection and alerting module, based on data provided by the real-time performance monitoring module, integrates and preprocesses information such as connection counts, lock wait times, and disk space usage from multiple data sources. The system analyzes historical data and uses time series analysis algorithms (such as the ARIMA model) to learn how each indicator changes over different time periods (weekdays, weekends, and business peaks and troughs). It then constructs a corresponding time series model to predict the indicator's fluctuation range under normal circumstances. Furthermore, anomaly detection models (such as the isolation forest algorithm) are trained based on historical normal data to identify anomalous patterns. The system sets reasonable thresholds for each indicator based on database configuration, business needs, and historical data. During operation, data is continuously monitored in real time, and the current data is compared with the normal patterns predicted by the thresholds, time series models, and anomaly detection models. Once the number of connections exceeds the threshold, the lock wait time exceeds the maximum allowed duration, or the disk space usage reaches the upper limit, the system immediately determines it as an anomaly and triggers a multi-channel alarm mechanism. Through the SMS interface, email sending module, and in-site message system, the alarm message containing the abnormal indicator name, current value, normal threshold range, and the time when the abnormality occurred will be pushed to the relevant operation and maintenance personnel in a timely manner to ensure that the problem is paid attention to at the first time.
[0025] The operation of the automated fault handling module specifically includes the following: Detailed classification of common faults that may occur during database operation, such as deadlocks, connection leaks, and long transaction blocking, and development of detailed repair strategies for each type of fault; Regularly evaluate and update the self-healing strategy library, and promptly adjust and improve the corresponding repair strategies as the database system is upgraded, business needs change, and new fault types emerge.
[0026] The operation of the automated fault handling module specifically includes the following: Works closely with the anomaly detection and alarm module, receives fault information sent by it in real time, and retrieves the corresponding repair strategy from the self-healing strategy library based on the received fault information; Automatically execute the corresponding repair operation according to the retrieved repair strategy. During the execution process, the repair progress and effect are monitored in real time. If the repair operation fails, the failure information is automatically recorded, including the failure reason and the steps executed. After the repair operation is completed, the relevant performance indicators and operating status of the database are checked and evaluated to determine whether the fault has been effectively resolved. The repair results and related data are fed back to the system to provide a basis for subsequent strategy optimization.
[0027] After receiving fault information from the anomaly detection and alerting module, the automated fault handling module accurately matches the fault type based on a pre-defined self-healing strategy library. This library includes detailed repair strategies for common faults such as database deadlocks, master-slave delays, connection leaks, and memory overflows. For example, when handling a deadlock, the system selects the appropriate transaction to roll back based on factors such as the deadlocked transaction's priority and execution time. After determining the repair strategy, the system automatically executes the repair operation and monitors its progress and results in real time. If the repair fails, the system records the cause of failure and the execution steps, and then attempts an alternative strategy or notifies operations and maintenance personnel to intervene. After the repair is complete, the system rechecks the database's performance indicators and operational status to assess whether the fault has been completely resolved. The repair results are then fed back to the system for subsequent strategy optimization, achieving closed-loop management of fault handling.
[0028] The operation of the intelligent indexing and query optimization module specifically includes the following: Use the SQL syntax analyzer to parse the extracted SQL statements. Based on the execution time, number of scanned rows, and number of returned rows in the query log, combined with the performance data of the database system, evaluate the execution efficiency of each SQL statement. Based on the structure of SQL statements and data distribution, it analyzes whether the existing indexes in the current database can meet the query requirements and automatically recommends appropriate index optimization solutions for each inefficient SQL statement. According to the characteristics of inefficient SQL statements and the identified problems, applicable semantically equivalent rewriting rules such as projection pushdown, predicate pushdown, subquery elimination, etc. are matched from the rule base to automatically rewrite the inefficient SQL.
[0029] The intelligent indexing and query optimization module uses a SQL parser to analyze SQL statements using database query logs, capturing their structure, logic, and key elements. It then combines performance data such as execution time, number of rows scanned, and number of rows returned to assess the efficiency of each SQL statement and identify inefficient SQL statements. For inefficient SQL statements, the system analyzes whether existing indexes meet requirements based on statement structure and data distribution. Using a cost-based index recommendation algorithm, the system recommends appropriate index optimization solutions, such as adding new indexes, modifying index structures, or deleting redundant indexes. Furthermore, based on the SQL rewrite rule library, the system automatically rewrites inefficient SQL statements by matching applicable rules, such as converting subqueries into join queries or adjusting the order of table joins. Rewritten SQL statements undergo syntax verification and performance testing to ensure logical correctness and improved performance, thereby optimizing overall database query performance.
[0030] The operation of the backup and recovery management module specifically includes the following: Comprehensively assess backup needs based on database size, business importance, and data update frequency, automatically execute backup tasks according to pre-set backup strategies, and monitor backup progress in real time, recording backup start and end times and backup data volume. Linked with the database monitoring system, it monitors the database operation status in real time and immediately triggers the recovery mechanism when database crash, accidental data deletion, hardware failure, or virus attack is detected; After the recovery process is started, the recovery operation is performed automatically according to the established plan or with the intervention of the operation and maintenance personnel. The database is quickly rebuilt using the backup data, and the data integrity and system consistency are verified after the recovery. The backup data is restored periodically to simulate real failure scenarios and verify the effectiveness of the backup data and recovery process.
[0031] The backup and recovery management module develops differentiated backup strategies based on database size, business importance, and data update frequency, clearly defining the cycles and execution order for full, incremental, and differential backups. The system automatically executes backup tasks according to these strategies, using built-in database tools or specialized software to store and encrypt data on local disk arrays, backup servers, or cloud storage. During the backup process, progress is monitored in real time, key information is recorded, and expired backups are cleared. When the database monitoring system detects anomalies such as a crash, accidental data deletion, hardware failure, or virus attack, the recovery mechanism is immediately triggered. Based on the anomaly scenario, the system selects appropriate backup data and performs recovery operations according to a predefined plan, either automatically or with intervention from operations and maintenance personnel, to rebuild the database. After recovery, integrity and consistency checks are performed, and recovery drills are conducted periodically to verify the availability of backup data and the effectiveness of the recovery process, ensuring data security and business continuity.
[0032] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0033] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. Database automated operation and maintenance system based on intelligent monitoring and self-healing mechanism, characterized by: It includes real-time performance monitoring module, anomaly detection and alarm module, automated fault handling module, intelligent indexing and query optimization module, and backup and recovery management module; The real-time performance monitoring module is used to collect the number of queries per second, number of transactions per second, number of connections, and slow query rate of the database. Through real-time collection and dynamic analysis of the number of queries per second, number of transactions per second, number of connections, and slow query rate, the health status of the database is evaluated. The anomaly detection and alerting module monitors various situations during database operation. For example, it uses the ARIMA time series forecasting model to predict normal fluctuations in key indicators and uses Isolation Forest to detect anomalies. When abnormal situations such as a sudden increase in the number of connections, lock wait timeouts, and insufficient disk space are identified, multi-channel alerts are immediately triggered. The automated fault handling module has a preset self-healing strategy library, which stores repair plans and execution processes by fault type and formulates corresponding repair strategies for common faults. The intelligent indexing and query optimization module is used to analyze the execution of SQL statements in the query log, automatically recommend index optimization solutions, and optimize inefficient SQL using syntax rewriting rules; The backup and recovery management module is used to automatically back up data regularly. It sets a reasonable backup cycle based on the size of the database and business needs, and triggers a rapid recovery mechanism in abnormal situations.
2. The database automated operation and maintenance system based on intelligent monitoring and self-healing mechanism according to claim 1 is characterized by: The real-time performance monitoring module operation specifically includes the following: Deploy a data collection agent on the database server to establish a connection with the database and collect the database's queries per second, transactions per second, number of connections, and slow query rate data at specific time intervals. The collected data is transmitted to the data processing center, and the sliding window algorithm is used to calculate the average value and standard deviation within a certain time window. If the current number of queries per second, number of transactions per second, number of connections, and slow query rate data deviate from the average value μ by more than a certain multiple of the standard deviation σ, it is determined that an abnormal fluctuation has occurred.
3. The database automated operation and maintenance system based on intelligent monitoring and self-healing mechanism according to claim 2 is characterized by: The operation of the anomaly detection and alarm module specifically includes the following: Obtain connection count, lock wait time, and disk space usage data from the real-time performance monitoring module and other relevant data sources of the database system, and perform integration and preprocessing; Analyze its historical data, learn the changing patterns of the number of connections in different time periods, establish a corresponding time series prediction model, and predict the fluctuation range of the number of connections under normal circumstances.
4. The database automated operation and maintenance system based on intelligent monitoring and self-healing mechanism according to claim 3 is characterized by: The operation of the anomaly detection and alarm module specifically includes the following: Based on historical data and normal patterns, we train ARIMA time series forecasting models and Isolation Forest anomaly detection models to identify abnormal patterns. We also set thresholds for connection counts, lock wait times, and disk space usage based on database configuration, business requirements, and historical operational data. Continuously monitor various indicator data in real time, comparing current data with set thresholds and normal fluctuation ranges obtained through time series analysis and anomaly detection models; When the number of connections exceeds the set threshold, or the lock waiting time exceeds the maximum allowed time, or the disk space usage reaches the set upper limit, it is determined to be an abnormal situation; Multiple alarm channels are pre-configured, including SMS interface, email sending module and in-site messaging system. Once an abnormal situation is detected, the alarm logic is immediately triggered to generate an alarm message containing the abnormal indicator name, current value, normal threshold range and the time when the abnormality occurred.
5. The database automated operation and maintenance system based on intelligent monitoring and self-healing mechanism according to claim 4 is characterized by: The operation of the automated fault handling module specifically includes the following: Categorize common faults that occur during database operation and develop detailed repair strategies for each type of fault. Regularly evaluate and update the self-healing strategy library, and promptly adjust and improve the corresponding repair strategies as the database system is upgraded, business needs change, and new fault types emerge.
6. The database automated operation and maintenance system based on intelligent monitoring and self-healing mechanism according to claim 5 is characterized by: The operation of the automated fault handling module specifically includes the following: Works closely with the anomaly detection and alarm module, receives fault information sent by it in real time, and retrieves the corresponding repair strategy from the self-healing strategy library based on the received fault information; Automatically execute the corresponding repair operation according to the retrieved repair strategy. During the execution process, the repair progress and effect are monitored in real time. If the repair operation fails, the failure information is automatically recorded, including the failure reason and the steps executed. After the repair operation is completed, the relevant performance indicators and operating status of the database are checked and evaluated to determine whether the fault has been effectively resolved. The repair results and related data are fed back to the system to provide a basis for subsequent strategy optimization.
7. The database automated operation and maintenance system based on intelligent monitoring and self-healing mechanism according to claim 6 is characterized by: The operation of the intelligent indexing and query optimization module specifically includes the following: Use the SQL syntax analyzer to parse the extracted SQL statements. Based on the execution time, number of scanned rows, and number of returned rows in the query log, combined with the performance data of the database system, evaluate the execution efficiency of each SQL statement. Based on the structure of SQL statements and data distribution, it analyzes whether the existing indexes in the current database can meet the query requirements and automatically recommends appropriate index optimization solutions for each inefficient SQL statement. According to the characteristics of inefficient SQL statements and the identified problems, the applicable semantically equivalent rewriting rules are matched from the rule base to automatically rewrite the inefficient SQL.
8. The database automated operation and maintenance system based on intelligent monitoring and self-healing mechanism according to claim 7 is characterized by: The operation of the backup and recovery management module specifically includes the following: Comprehensively assess backup needs based on database size, business importance, and data update frequency, automatically execute backup tasks according to pre-set backup strategies, and monitor backup progress in real time, recording backup start and end times and backup data volume. Linked with the database monitoring system, it monitors the database operation status in real time and immediately triggers the recovery mechanism when database crash, accidental data deletion, hardware failure, or virus attack is detected; After the recovery process is started, the recovery operation is performed automatically according to the established plan or with the intervention of the operation and maintenance personnel. The database is quickly rebuilt using the backup data, and the data integrity and system consistency are verified after the recovery. The backup data is restored periodically to simulate real failure scenarios and verify the effectiveness of the backup data and recovery process.
Citation Information
Patent Citations
Fault self-recovery system and method based on MySQL database
CN117632651A
Intelligent early warning and disposal method based on historical monitoring data
CN117827608A
Distributed market data acquisition management system
CN117876016A
Data anomaly detection method and system for water supply network, medium and equipment
CN118654234A
Real-time fault monitoring Internet of Things system for chemical production equipment cluster
CN119232773A