Fault detection method and apparatus for database instance

CN115509784BActive Publication Date: 2026-09-18JD DIGITS HAIYI INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211173402.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-26
Publication Date
2026-09-18
Estimated Expiration
2042-09-26

AI Technical Summary

Technical Problem

[0004]现有技术中依靠自动化产品和人工的方式来进行数据库故障检测,对于故障的感知、决策、执行等方面的能力都不能够满足要求,技术迁移能力较弱,不具备直接迁移到其他数据库性能监控的能力;依赖于专家经验,缺乏自动学习能力;且只有故障发生才会报警,不能及时发现系统中存在的隐患,故障检测效果差

Benefits of technology

[0029] One embodiment of the above invention has the following advantages or beneficial effects: It obtains monitoring data and historical fault records corresponding to monitoring indicators of a database instance within a specified time period; determines a time window for indicator processing based on the monitoring data and the historical fault records; processes the monitoring data according to the time window and matches it with the historical fault records to filter a first indicator set from the monitoring indicator set, and determines the weight corresponding to each first indicator in the first indicator set; trains a model based on the monitoring data and weights corresponding to each first indicator to obtain a database instance scoring model; and uses the database instance scoring model to score the monitoring data of the input database instance, and scores the data based on the database... This technical solution for fault detection based on instance scores incorporates core monitoring metrics of database instances into a scoring model. It also combines historical fault records to score the health status of the database instance, improving the ability to perceive, decide on, and execute fault assessments. By simulating various performance monitoring metrics of the database instance and performing feature engineering calculations, it automatically selects important monitoring items and trains a multi-dimensional fault detection scoring model based on historical data. The scoring model has strong transferability, low dependence on expert experience, and can quickly pinpoint the root cause of a score drop, providing a specific deduction for each root cause. Previously undetectable issues, often occurring weekly or monthly, can now be identified and addressed within a single inspection cycle, thus improving fault detection efficiency and providing fault early warning capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115509784B_ABST
    Figure CN115509784B_ABST
Patent Text Reader

Abstract

The application discloses a kind of database instance fault detection method and device, it is related to computer technical field.The specific embodiment of the method includes: obtaining the monitoring data corresponding to the monitoring index of database instance in specified period and historical fault record;According to monitoring data and historical fault record, determine the time window of index processing;According to time window, monitoring data is processed, and is matched with historical fault record to obtain first index set from monitoring index set by screening, and determine the weight corresponding to each first index in first index set;Based on the monitoring data and weight corresponding to each first index, model training is carried out, and database instance scoring model is obtained;The input monitoring data of database instance is scored using database instance scoring model, and fault detection is carried out according to the score of database instance.The embodiment improves the ability of perception, decision, execution and other aspects to the fault of database instance, and improves fault detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method and apparatus for fault detection of a database instance. Background Technology

[0002] In the era of big data, the importance of data storage is self-evident. Once data is stored in a database, the normal operation of the database is crucial to the normal operation of business systems. Currently, database performance and health status testing is mostly conducted manually based on the experience of database experts or through automated products.

[0003] In the process of realizing this invention, the inventors discovered at least the following problems in the prior art:

[0004] Existing technologies rely on automated products and manual methods for database fault detection. However, these methods are insufficient in terms of fault perception, decision-making, and execution capabilities. They also have weak technology transfer capabilities and cannot be directly transferred to other database performance monitoring systems. Furthermore, they depend on expert experience and lack automatic learning capabilities. Moreover, they only trigger alarms when a fault occurs, failing to promptly identify potential problems in the system, resulting in poor fault detection performance. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a fault detection method and apparatus for database instances, which can improve the ability to perceive, make decisions, and execute faults in database instances; the scoring model has strong transferability and low dependence on expert experience and knowledge, and can quickly locate the root cause indicators that cause the score drop when the database instance score drops, and give the specific deduction degree for each root cause indicator; it can promptly discover hidden dangers in the system, improve fault detection efficiency, and has fault early warning capabilities.

[0006] To achieve the above objectives, according to one aspect of the present invention, a fault detection method for a database instance is provided, comprising:

[0007] Retrieve monitoring data and historical fault records corresponding to the monitoring metrics of the database instance within a specified time period;

[0008] The time window for indicator processing is determined based on the monitoring data and the historical fault records.

[0009] The monitoring data is processed according to the time window and matched with the historical fault records to filter out the first indicator set from the monitoring indicator set, and the weight corresponding to each first indicator in the first indicator set is determined.

[0010] The database instance scoring model is obtained by training the model based on the monitoring data and weights corresponding to each of the first indicators.

[0011] The database instance scoring model is used to score the monitoring data of the input database instance, and fault detection is performed based on the database instance's score.

[0012] Optionally, before determining the time window for indicator processing based on the monitoring data and the historical fault records, the method further includes: preprocessing the monitoring data corresponding to the monitoring indicators of the database instance to perform missing value filling and exponential smoothing.

[0013] Optionally, determining the time window for indicator processing based on the monitoring data and the historical fault records includes: processing the monitoring data according to different initial time windows to obtain at least one indicator fluctuation regularity data; calculating the matching degree between each indicator fluctuation regularity data and the fault time point in the historical fault records, and selecting the indicator processing time window from the initial time window based on the matching degree.

[0014] Optionally, processing the monitoring data according to the time window and matching it with the historical fault records to filter out a first set of indicators from the monitoring indicator set, and determining the weight corresponding to each first indicator in the first set of indicators, includes: processing the monitoring data according to the time window to obtain indicator fluctuation regularity data; filtering out a first set of indicators from the monitoring indicator set based on the matching degree between the indicator fluctuation regularity data and the fault time points in the historical fault records; obtaining multi-indicator fluctuation regularity data of the fault time points based on the historical fault records; and determining the weight corresponding to each first indicator in the first set of indicators based on the correlation between the multi-indicator fluctuation regularity data and the fault time points.

[0015] Optionally, before processing the monitoring data according to the time window, the method further includes: pre-screening the monitoring indicators according to the monitoring data and the historical fault records to obtain a second indicator set from the monitoring indicator set, and determining the weight corresponding to each second indicator in the second indicator set; and processing the monitoring data according to the time window and matching it with the historical fault records to obtain a first indicator set from the monitoring indicator set, including: processing the monitoring data corresponding to each second indicator in the second indicator set according to the time window and matching it with the historical fault records to obtain a first indicator set from the second indicator set.

[0016] Optionally, before processing the monitoring data according to the time window, the method further includes: classifying the acquired database instances according to the fault type of the database instances that failed in the historical fault list, resulting in database instance categories including healthy instances, high-load instances, and faulty instances; processing the monitoring data according to the time window and matching it with the historical fault records to filter out a first set of indicators from the monitoring indicator set, and determining the weight corresponding to each first indicator in the first indicator set, including: for each category of database instances, processing the monitoring data of the database instances of that category according to the time window and matching it with the historical fault records to filter out a first set of indicators for the database instances of that category from the monitoring indicator set, and determining the weight corresponding to each first indicator in the first indicator set of the database instances of that category; training a model based on the monitoring data and weight corresponding to each first indicator to obtain a database instance scoring model, including: training a model based on the monitoring data and weight corresponding to each first indicator of each category of database instances to obtain a database instance scoring model.

[0017] Optionally, a database instance scoring model is obtained by training a model based on the monitoring data and weights corresponding to each of the first indicators, including: dividing the first indicator set into multiple indicator groups according to the weight of each of the first indicators, and calculating the average weight of each indicator group; extracting fluctuation features from the monitoring data corresponding to each indicator group, and extracting time-series features from the monitoring data corresponding to the first indicator set; weighting the fluctuation features of each indicator group according to the average weight of each indicator group, and fusing the weighted fluctuation features and the time-series features; and calculating the failure probability of the database instance based on the fused features to train the model and obtain the database instance scoring model.

[0018] Optionally, the first indicator set is divided into multiple indicator groups according to the weight of each first indicator, and the average weight of each indicator group is calculated, including: sorting the first indicators from largest to smallest according to the weight value of each first indicator; determining the number of indicators in each indicator group according to the number of first indicators and the preset number of groups; combining the first indicators in sequence according to the sorting result and the number of indicators in each indicator group to divide the first indicator set into multiple indicator groups; and calculating the average weight of each indicator group according to the weight value of the first indicators included in each indicator group.

[0019] Optionally, fluctuation features are extracted from the monitoring data corresponding to each indicator group, including: inputting the monitoring data corresponding to each indicator group into different first convolutional sub-layers to obtain the initial fluctuation features of each indicator group; and inputting the initial fluctuation features of each indicator group into a second convolutional sub-layer sequentially connected to the first convolutional sub-layer to obtain the fluctuation features of each indicator group.

[0020] Optionally, fault detection is performed based on the database instance's score, including: when the current score of the database instance decreases compared to the previous score, performing anomaly detection on the specified monitoring data corresponding to each of the first indicators of the database instance using the isolated forest algorithm to obtain anomaly indicators, and using the anomaly indicators as the fault detection results. The specified monitoring data includes monitoring data from the most recent half hour, monitoring data from the half hour before and after the same period last year, and monitoring data from the half hour before and after the same period last month.

[0021] According to another aspect of the present invention, a fault detection device for a database instance is provided, comprising:

[0022] The data acquisition module is used to acquire monitoring data and historical fault records corresponding to the monitoring indicators of the database instance within a specified time period;

[0023] The window selection module is used to determine the time window for index processing based on the monitoring data and the historical fault records.

[0024] The indicator filtering module is used to process the monitoring data according to the time window, match it with the historical fault records to filter the first indicator set from the monitoring indicator set, and determine the weight corresponding to each first indicator in the first indicator set.

[0025] The model training module is used to train the model based on the monitoring data and weights corresponding to each of the first indicators to obtain the database instance scoring model.

[0026] The scoring and detection module is used to score the monitoring data of the input database instance using the database instance scoring model, and to perform fault detection based on the database instance's score.

[0027] According to another aspect of the present invention, a fault detection electronic device for a database instance is provided, comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the fault detection method for a database instance provided in the embodiments of the present invention.

[0028] According to another aspect of the present invention, a computer-readable medium is provided having a computer program stored thereon, which, when executed by a processor, implements the fault detection method for a database instance provided in the embodiments of the present invention.

[0029] One embodiment of the above invention has the following advantages or beneficial effects: It obtains monitoring data and historical fault records corresponding to monitoring indicators of a database instance within a specified time period; determines a time window for indicator processing based on the monitoring data and the historical fault records; processes the monitoring data according to the time window and matches it with the historical fault records to filter a first indicator set from the monitoring indicator set, and determines the weight corresponding to each first indicator in the first indicator set; trains a model based on the monitoring data and weights corresponding to each first indicator to obtain a database instance scoring model; and uses the database instance scoring model to score the monitoring data of the input database instance, and scores the data based on the database... This technical solution for fault detection based on instance scores incorporates core monitoring metrics of database instances into a scoring model. It also combines historical fault records to score the health status of the database instance, improving the ability to perceive, decide on, and execute fault assessments. By simulating various performance monitoring metrics of the database instance and performing feature engineering calculations, it automatically selects important monitoring items and trains a multi-dimensional fault detection scoring model based on historical data. The scoring model has strong transferability, low dependence on expert experience, and can quickly pinpoint the root cause of a score drop, providing a specific deduction for each root cause. Previously undetectable issues, often occurring weekly or monthly, can now be identified and addressed within a single inspection cycle, thus improving fault detection efficiency and providing fault early warning capabilities.

[0030] The further effects of the aforementioned unconventional alternative methods will be explained below in conjunction with specific implementation methods. Attached Figure Description

[0031] The accompanying drawings are provided to better understand the invention and are not intended to unduly limit the scope of the invention. Wherein:

[0032] Figure 1 This is a schematic diagram of the main steps of a fault detection method for a database instance according to an embodiment of the present invention;

[0033] Figure 2 This is a schematic diagram of the training process of the scoring model in an embodiment of the present invention;

[0034] Figure 3 This is a schematic diagram of the training process of a scoring model according to an embodiment of the present invention;

[0035] Figure 4 This is a schematic diagram of the training process of a scoring model according to another embodiment of the present invention;

[0036] Figure 5 This is a schematic diagram of the main modules of a fault detection device for a database instance according to an embodiment of the present invention.

[0037] Figure 6 This is an exemplary system architecture diagram in which embodiments of the present invention can be applied;

[0038] Figure 7 This is a schematic diagram of the structure of a computer system suitable for implementing terminal devices or servers of the present invention. Detailed Implementation

[0039] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of the present invention, including various details to aid understanding. These details should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the invention. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0040] To address the technical problems existing in current technologies, this invention establishes an intelligent scoring calculation model. This model incorporates core database monitoring metrics (such as active connections, CPU, memory, and disk usage) into the scoring system. It also combines historical anomaly rates with in-depth analysis of slow logs, deadlocks, and audit logs to identify potential problems and comprehensively score the database's health status. For example, a "database deadlock" issue might cause a service interruption because this fault has never occurred before and is an unknown problem. Therefore, monitoring systems lack relevant root cause analysis rules (database engineers also lack experience with unknown problems, so corresponding monitoring metrics are not configured, failing to detect unknown issues). However, through the intelligent scoring model of this invention, fault replay analysis can identify this fault in advance, thus preventing such a major incident. The database performance monitoring and inspection items involved in the intelligent scoring calculation cover 64 items. This invention simulates data of various database performance monitoring indicators under different states, performs feature engineering calculations, automatically selects important indicator monitoring items, and trains a multi-dimensional anomaly detection network scoring model based on historical data. The intelligent scoring model has strong transferability, low dependence on expert experience and knowledge, and can quickly locate the root cause indicator causing the instance score drop when the instance score drops, and give the specific deduction degree for each item. Potential problems that were previously undetectable at the weekly or monthly level can now be discovered and handled within a single inspection cycle.

[0041] Figure 1This is a schematic diagram illustrating the main steps of a fault detection method for a database instance according to an embodiment of the present invention. Figure 1 As shown, the fault detection method for database instances in this embodiment of the invention mainly includes the following steps S101 to S105.

[0042] Step S101: Obtain monitoring data and historical fault records corresponding to the monitoring metrics of the database instance within a specified time period. In embodiments of the present invention, the specified time period is, for example, a recent period of a specified length, such as monitoring data within the most recent month. The monitoring data corresponds to the monitoring metrics. In embodiments of the present invention, there are, for example, 64 preset monitoring metrics, including, for example, CPU utilization, memory utilization, and write IOPS (Input / Output Operations). PerSecond (read / write operations per second), total disk space usage, network outgoing traffic, network incoming traffic, read IOPS, deletes per second, average row lock acquisition time for InnoDB (one of MySQL's database engines), number of fsync writes completed per second in the InnoDB log, physical writes per second in the InnoDB log, InnoDB write volume per second, InnoDB writes to the buffer pool per second, InnoDB reads to the buffer pool per second, InnoDB read volume per second, number of row lock waits in InnoDB, InnoDB cache pool utilization, InnoDB cache pool dirty block rate, InnoDB cache pool read hit rate, and inserts per second. The system uses various metrics including: Insert_Select (database query insert statement), Replace, Replace_Select, Select, update count per second, number of temporary tables, temporary tablespace usage, shared tablespace usage, instance input traffic per second, instance output traffic per second, total current connections, active current connections, slow queries, log file usage, select count per second (QPS), transactions per second (TPS), user data usage, number of table lock waits, system data usage, number of table lock waits, binlog space usage, total disk space usage, user data usage, undo space usage, system load, lock wait counts, network detection, port detection, database connection detection, database query detection, slave IO thread status, slave SQL thread status, slave latency, queries per second (QPS), transactions per second (TPS), frontend connections, backend connections, QPS, TPS, CPU utilization, memory utilization, disk utilization, network inbound traffic per second, network outbound traffic per second, etc.

[0043] The historical fault log records the time, type, and other information of each database instance failure within the past month.

[0044] Step S102: Determine the time window for indicator processing based on monitoring data and historical fault records. When processing data, a time window needs to be set to determine the scope of each data processing step.

[0045] According to one embodiment of the present invention, before determining the time window for indicator processing based on the monitoring data and the historical fault records, the method further includes: preprocessing the monitoring data corresponding to the monitoring indicators of the database instance to perform missing value imputation and exponential smoothing. The preprocessed data is the data that has already undergone noise reduction and filtering, resulting in better processing performance and facilitating subsequent processing and analysis.

[0046] According to another embodiment of the present invention, determining the time window for indicator processing based on the monitoring data and the historical fault records may specifically include: processing the monitoring data according to different initial time windows to obtain at least one indicator fluctuation regularity data; calculating the matching degree between each indicator fluctuation regularity data and the fault time point in the historical fault record; and selecting the indicator processing time window from the initial time window based on the matching degree. In embodiments of the present invention, the initial time window may include, for example, 10 seconds, 1 minute, and 5 minutes. The optimal data processing time window is selected by comparing the indicator fluctuation regularity data (e.g., indicator fluctuation curve) obtained by processing the original monitoring indicator data according to different initial time windows with the matching degree of the fault time point recorded in the historical fault record.

[0047] Additionally, during step S102, monitoring indicators can be pre-screened based on monitoring data and historical fault records to obtain a second indicator set from the monitoring indicator set, and the weight of each second indicator in the second indicator set can be determined. Since the monitoring indicator set contains a large number of indicators, a pre-screening of the monitoring indicators can be performed in this step to obtain some important monitoring indicators (e.g., 16), forming the second indicator set. Furthermore, based on the fluctuation patterns of multi-indicator monitoring data from each fault time point in the historical fault records, and their correlation with the fault time point (values ​​between 0 and 1), the weight value of each second indicator can be set to further rank the importance of the second indicators.

[0048] Step S103: Process the monitoring data according to the time window and match it with historical fault records to filter the first indicator set from the monitoring indicator set, and determine the weight of each first indicator in the first indicator set.

[0049] According to an embodiment of the present invention, when processing the monitoring data according to the time window and matching it with the historical fault records to filter out a first set of indicators from the monitoring indicator set, and determining the weight corresponding to each first indicator in the first set of indicators, the specific steps can be performed as follows: processing the monitoring data according to the time window to obtain indicator fluctuation regularity data; filtering out a first set of indicators from the monitoring indicator set according to the matching degree between the indicator fluctuation regularity data and the fault time points in the historical fault records; obtaining multi-indicator fluctuation regularity data of the fault time points according to the historical fault records; determining the weight corresponding to each first indicator in the first set of indicators according to the correlation between the multi-indicator fluctuation regularity data and the fault time points. Specifically, this step will filter out the most important first set of indicators from the monitoring indicator set and determine the weight corresponding to each first indicator. The original monitoring indicator data is processed according to the time window selected in step S102, and the matching degree between the processed indicator fluctuation regularity data and the fault time points in the historical fault records is calculated to select nine important indicators from the monitoring indicator set. Assuming a selected time window of 1 minute, data processing is performed once per minute. The data within that minute is processed to obtain data on the regularity of indicator fluctuations, such as the fluctuation curve for each indicator. Then, the fluctuation curve is matched with the fault time points in historical fault records to select the most important primary indicator. The selected primary indicator might be any 9 of the following: load_one (database load), memory_used (memory usage), disk_used (disk usage), sys_cpu_total_time (total CPU time of process sys), sys_cpu_idle_time (CPU idle time of process sys), cpu_time_percent, tps, qps, mysql_network_in (network inbound traffic), mysql_network_out (network outbound traffic), mysql_innodb_rows_read, mysql_slow_logs (database query efficiency (slow SQL)), mysql_active_connections (active connections), mysql_connect_percent, mysql_cpu_active_time, mysql_cpu_percent, etc.

[0050] According to one embodiment of the present invention, before processing the monitoring data according to the time window, the monitoring indicators are pre-screened based on the monitoring data and the historical fault records to obtain a second indicator set from the monitoring indicator set, and the weight corresponding to each second indicator in the second indicator set is determined. Therefore, step S103, when processing the monitoring data according to the time window and matching it with the historical fault records to obtain a first indicator set from the monitoring indicator set, may specifically include: processing the monitoring data corresponding to each second indicator in the second indicator set according to the time window, and matching it with the historical fault records to obtain a first indicator set from the second indicator set. Through these two rounds of monitoring indicator screening, important monitoring indicators can be more accurately identified to constitute the first indicator set.

[0051] Step S104: Train the model based on the monitoring data and weights corresponding to each first indicator to obtain the database instance scoring model.

[0052] According to one embodiment of the present invention, when training a model based on monitoring data and weights corresponding to each first indicator to obtain a database instance scoring model, the specific steps may include: dividing the first indicator set into multiple indicator groups according to the weight of each first indicator, and calculating the average weight of each indicator group; extracting fluctuation features from the monitoring data corresponding to each indicator group, and extracting time-series features from the monitoring data corresponding to the first indicator set; weighting the fluctuation features of each indicator group according to the average weight of each indicator group, and fusing the weighted fluctuation features and the time-series features; calculating the failure probability of the database instance based on the fused features to train the model and obtain the database instance scoring model. In an embodiment of the present invention, the first indicator set can be divided into 3 groups, and fluctuation features can be extracted from the monitoring data corresponding to each first indicator group, while time-series features can be extracted from the monitoring data of all first indicators. Then, the fluctuation features and time-series features of the multiple indicator groups are fused and concatenated by combining the weights of each indicator group; finally, the fused features are used to train the model to obtain the database instance scoring model.

[0053] According to one embodiment of the present invention, when dividing the first indicator set into multiple indicator groups according to the weight of each first indicator and calculating the average weight of each indicator group, the process may include: sorting the first indicators from largest to smallest according to their weight values; determining the number of indicators in each indicator group based on the number of first indicators and a preset number of groups; sequentially combining the first indicators according to the sorting result and the number of indicators in each indicator group to divide the first indicator set into multiple indicator groups; and calculating the average weight of each indicator group based on the weight values ​​of the first indicators included in each indicator group. For the nine first indicators in the selected first indicator set, they are first sorted according to their weights, and then these nine first indicators are divided into three groups: the three first indicators with the highest weights form one group, the three first indicators with the lowest weights form another group, and the remaining three first indicators form a third group. This avoids inaccurate fault identification results due to interference between indicators when many indicators are trained together in the model. Meanwhile, since there is a correlation between the various monitoring indicators, the three primary indicators with strong correlation can be grouped into an indicator group. This takes into account the correlation between the indicators, avoids interference between the indicators, and can also improve the efficiency of model training.

[0054] According to one embodiment of the present invention, fluctuation features are extracted from the monitoring data corresponding to each indicator group, including: inputting the monitoring data corresponding to each indicator group into different first convolutional sub-layers to obtain the initial fluctuation features of each indicator group; and inputting the initial fluctuation features of each indicator group into a second convolutional sub-layer sequentially connected to the first convolutional sub-layer to obtain the fluctuation features of each indicator group.

[0055] According to one embodiment of the present invention, the model training process is as follows:

[0056] Step 1: Group the first set of indicators according to their weights to determine 3 indicator groups;

[0057] The second step is to input each indicator group into a different first convolutional sub-layer, and then input the initial fluctuation feature information output by each of the first convolutional sub-layers into the sequentially connected second convolutional sub-layers. The second convolutional sub-layers can output the fluctuation feature information of the samples. At the same time, the first indicator set is input into the temporal feature extraction sub-module (LSTM) to output temporal feature information.

[0058] Step 3: Integrate fluctuation feature information and time series feature information. After passing through the indicator features, the weighted mean of the indicator group is input into the concat fusion to calculate, where W1 is the weighted mean of the first indicator group, W2 is the weighted mean of the second indicator group, and W3 is the weighted mean of the third indicator group.

[0059] Step 4: Finally, the softmax function outputs the probability of database instance anomalies or failures.

[0060] In model training, the objective function is set as Loss = CrossEntropy(y', y). The Adam optimizer is used to minimize the objective function, and the training parameters of the LSTM and CNN networks are updated through backpropagation. The attention mechanism does not involve parameter updates. Loss = CrossEntropy(y', y) represents the closeness between the predicted label and the true label. The larger the value of this objective function, the better the model training effect.

[0061] Figure 2 This is a schematic diagram illustrating the training process of the scoring model in an embodiment of the present invention. For example... Figure 2 As shown, the model in this embodiment of the invention is, for example, the W-CNN-LSTM model, which consists of three CNN convolutional layers and one LSTM layer. Its training process is as follows: The first indicator set is grouped according to the indicator weights, resulting in three indicator groups. The monitoring data corresponding to each indicator group is input into different first convolutional sub-layers to obtain the initial fluctuation features corresponding to each indicator group. Then, the initial fluctuation features output by each of the first convolutional sub-layers are input into a second convolutional sub-layer sequentially connected to the first convolutional sub-layer, which can output fluctuation features. Simultaneously, the monitoring data corresponding to the first indicator set is input into a temporal feature extraction sub-module (LSTM) to output temporal features. The fluctuation features of each indicator group are weighted according to the average weight of each indicator group, and the weighted fluctuation features and temporal features are fused. Here, W1 is the average weight of the first indicator group, W2 is the average weight of the second indicator group, and W3 is the average weight of the third indicator group. Finally, the probability of database instance anomalies or failures is output through the softmax function.

[0062] Step S105: Use a database instance scoring model to score the monitoring data of the input database instance, and perform fault detection based on the database instance's score. According to an embodiment of the present invention, fault detection based on the database instance's score includes: when the current score of the database instance decreases compared to the previous score, performing anomaly detection on the specified monitoring data corresponding to each of the first indicators of the database instance using an isolated forest algorithm to obtain anomaly indicators, and using the anomaly indicators as the fault detection result. The specified monitoring data includes monitoring data from the most recent half hour, monitoring data from the half hour before and after the same period last year, and monitoring data from the half hour before and after the same period last year. Specifically, when the database instance's score decreases, the monitoring data from the most recent half hour, the monitoring data from the half hour before and after the same period last year, and the monitoring data from the half hour before and after the same period last year corresponding to the nine first indicators of each database instance with a decreased score are transmitted to an isolated forest for anomaly detection, outputting anomaly labels from the most recent half hour. The set of indicators with anomalies in the most recent half hour is diagnosed as root cause indicators. The deduction degree of each root cause indicator is related to the indicator's weight; the greater the weight of the indicator, the greater the deduction when it is abnormal, increasing in multiples of 5.

[0063] According to one embodiment of the present invention, before processing the monitoring data according to the time window, the method further includes: classifying the acquired database instances according to the fault type of the database instances that have failed in the historical fault list, and the resulting database instance categories include healthy instances, high-load instances, and faulty instances.

[0064] The monitoring data is processed according to the time window and matched with the historical fault records to filter the first set of indicators from the monitoring indicator set, and the weight of each first indicator in the first set of indicators is determined. This includes: for each category of database instance, the monitoring data of the database instance of the category is processed according to the time window and matched with the historical fault records to filter the first set of indicators of the database instance of the category from the monitoring indicator set, and the weight of each first indicator in the first set of indicators of the database instance of the category is determined.

[0065] The database instance scoring model is obtained by training the model based on the monitoring data and weights corresponding to each of the first indicators for each category of database instance.

[0066] According to this embodiment, after determining the time window for indicator processing based on monitoring data and historical fault records, a first indicator set can be selected for each category of database instances. The first indicator sets for each category can be different or the same; correspondingly, the weights of each first indicator in the first indicator set for each category can also be different or the same. During subsequent model training, regardless of whether the first indicator sets are the same, model training only needs to be performed based on the monitoring data corresponding to the first indicator set and the weights of each first indicator. This allows for a closer fit to the characteristics of the database instance monitoring data, resulting in a better performance of the trained database instance scoring model.

[0067] Figure 3 This is a schematic diagram illustrating the training process of a scoring model according to an embodiment of the present invention. Figure 3 As shown, in one embodiment of the present invention, the training process of the scoring model is mainly as follows:

[0068] 1. Design of scoring model: Model selection, multi-index anomaly detection model. In the embodiments of this invention, the W-CNN-LSTM model is obtained by improving the TapNet network model.

[0069] 2. Determine the time window for indicator processing and the weight corresponding to each monitoring indicator based on monitoring data and historical fault records. In this step, the monitoring indicators can also be initially screened, and the weight corresponding to each preliminarily screened monitoring indicator can be determined.

[0070] 3. Process the monitoring data using time windows, filter the data from the monitoring indicator set to obtain the first indicator set, and determine the weight of each first indicator in the first indicator set. Specifically, process the monitoring data according to the selected time window, match it with historical fault records to filter the data from the monitoring indicator set to obtain the first indicator set, and determine the weight of each first indicator in the first indicator set.

[0071] 4. Combine the weight of each primary indicator and use the monitoring data corresponding to each primary indicator to train the model and generate a scoring model.

[0072] Figure 4 This is a schematic diagram illustrating the training process of a scoring model according to another embodiment of the present invention. Figure 3 As shown, in another embodiment of the present invention, the training process of the scoring model is mainly as follows:

[0073] 1. Design of scoring model: Model selection, multi-index anomaly detection model. In the embodiments of this invention, the TapNet network model is improved to obtain the W-CNN-LSTM model.

[0074] 2. Determine the time window for indicator processing and the weight corresponding to each monitoring indicator based on monitoring data and historical fault records. In this step, the monitoring indicators can also be initially screened, and the weight corresponding to each preliminarily screened monitoring indicator can be determined.

[0075] 3. Based on the failure types of database instances that occurred in the historical failure list, the database instances are categorized into healthy instances, high-load instances, and failed instances. The categorization can be based on pre-defined criteria for each database instance category.

[0076] 4. Based on the time window, process the monitoring data of healthy instances, high-load instances, and faulty instances separately, filter out the first indicator set from the monitoring indicator set, and determine the weight corresponding to each first indicator in the first indicator set. Specifically, process the monitoring data of healthy instances, high-load instances, and faulty instances according to the selected time window, and match them with historical fault records to filter out the first indicator set from the monitoring indicator set, and determine the weight corresponding to each first indicator in the first indicator set;

[0077] 5. Combine the weight of each primary indicator and use the monitoring data corresponding to each primary indicator to train the model and generate a scoring model.

[0078] Figure 5 This is a schematic diagram of the main modules of a fault detection device for a database instance according to an embodiment of the present invention. Figure 5 As shown, the fault detection device 500 for the database instance of this invention mainly includes a data acquisition module 501, a window selection module 502, an indicator filtering module 503, a model training module 504, and a scoring detection module 505.

[0079] The data acquisition module 501 is used to acquire monitoring data and historical fault records corresponding to the monitoring indicators of the database instance within a specified time period;

[0080] The window selection module 502 is used to determine the time window for index processing based on the monitoring data and the historical fault records.

[0081] The indicator filtering module 503 is used to process the monitoring data according to the time window, match it with the historical fault records to filter the first indicator set from the monitoring indicator set, and determine the weight corresponding to each first indicator in the first indicator set.

[0082] The model training module 504 is used to train the model based on the monitoring data and weights corresponding to each of the first indicators to obtain the database instance scoring model.

[0083] The scoring and detection module 505 is used to score the monitoring data of the input database instance using the database instance scoring model, and to perform fault detection based on the database instance's score.

[0084] According to an embodiment of the present invention, the fault detection device 500 of the database instance further includes a data preprocessing module (not shown in the figure), which is used to: preprocess the monitoring data corresponding to the monitoring indicators of the database instance before determining the time window for indicator processing based on the monitoring data and the historical fault records, so as to perform missing value filling and exponential smoothing processing.

[0085] According to another embodiment of the present invention, the window selection module 502 can also be used to: process the monitoring data according to different initial time windows to obtain at least one indicator fluctuation regularity data; calculate the matching degree between each indicator fluctuation regularity data and the fault time point in the historical fault record, and select the indicator processing time window from the initial time window according to the matching degree.

[0086] According to another embodiment of the present invention, the indicator screening module 503 can also be used to: process the monitoring data according to the time window to obtain indicator fluctuation regularity data; filter from the monitoring indicator set to obtain a first indicator set according to the matching degree between the indicator fluctuation regularity data and the fault time points in the historical fault records; obtain multi-indicator fluctuation regularity data of the fault time points according to the historical fault records; and determine the weight corresponding to each first indicator in the first indicator set according to the correlation between the multi-indicator fluctuation regularity data and the fault time points.

[0087] According to another embodiment of the present invention, the fault detection device 500 of the database instance further includes an indicator pre-screening module (not shown in the figure), which is used to: pre-screen the monitoring indicators according to the monitoring data and the historical fault records before processing the monitoring data according to the time window, so as to screen a second indicator set from the monitoring indicator set and determine the weight corresponding to each second indicator in the second indicator set.

[0088] Furthermore, the indicator filtering module 503 can also be used to: process the monitoring data corresponding to each second indicator in the second indicator set according to the time window, and match it with the historical fault records to filter out the first indicator set from the second indicator set.

[0089] According to another embodiment of the present invention, the database instance fault detection device 500 further includes an instance classification module (not shown in the figure), which is used to classify the acquired database instances according to the fault type of the database instances that have failed in the historical fault list before processing the monitoring data according to the time window, and the obtained database instance categories include healthy instances, high-load instances and faulty instances.

[0090] The indicator filtering module 503 can also be used to: for each category of database instance, process the monitoring data of the database instance of the category according to the time window, and match it with the historical fault records to filter the first indicator set of the database instance of the category from the monitoring indicator set, and determine the weight corresponding to each first indicator in the first indicator set of the database instance of the category.

[0091] The model training module 504 can also be used to: train the model based on the monitoring data and weights corresponding to each of the first indicators for each category of database instance to obtain a database instance scoring model.

[0092] According to another embodiment of the present invention, the model training module 504 can also be used to: divide the first indicator set into multiple indicator groups according to the weight of each first indicator, and calculate the average weight of each indicator group; extract fluctuation features from the monitoring data corresponding to each indicator group, and extract time-series features from the monitoring data corresponding to the first indicator set; weight the fluctuation features of each indicator group according to the average weight of each indicator group, and fuse the weighted fluctuation features and the time-series features; calculate the failure probability of the database instance based on the fused features to train the model and obtain a database instance scoring model.

[0093] According to another embodiment of the present invention, when the model training module 504 divides the first indicator set into multiple indicator groups according to the weight of each first indicator and calculates the weight mean of each indicator group, it may specifically be used to: sort the first indicator set from largest to smallest according to the weight value of each first indicator; determine the number of indicators in each indicator group according to the number of first indicators and the preset number of groups; combine the first indicators in sequence according to the sorting result and the number of indicators in each indicator group to divide the first indicator set into multiple indicator groups; and calculate the weight mean of each indicator group according to the weight value of the first indicator included in each indicator group.

[0094] According to another embodiment of the present invention, when the model training module 504 extracts fluctuation features from the monitoring data corresponding to each indicator group, it can also be used to: input the monitoring data corresponding to each indicator group into different first convolutional sub-layers to obtain the initial fluctuation features of each indicator group; input the initial fluctuation features of each indicator group into a second convolutional sub-layer sequentially connected to the first convolutional sub-layer to obtain the fluctuation features of each indicator group.

[0095] According to another embodiment of the present invention, the scoring detection module 505 can also be used to: when the current score of the database instance decreases compared with the previous score, perform anomaly detection on the specified monitoring data corresponding to each of the first indicators of the database instance through the isolated forest algorithm to obtain anomaly indicators, and use the anomaly indicators as fault detection results, wherein the specified monitoring data includes monitoring data of the most recent half hour, monitoring data of the half hour before and after the same period last year, and monitoring data of the half hour before and after the same period last month.

[0096] According to the technical solution of this embodiment of the invention, monitoring data and historical fault records corresponding to monitoring indicators of a database instance within a specified time period are obtained; a time window for indicator processing is determined based on the monitoring data and the historical fault records; the monitoring data is processed according to the time window and matched with the historical fault records to filter a first indicator set from the monitoring indicator set, and the weight corresponding to each first indicator in the first indicator set is determined; a model is trained based on the monitoring data and weight corresponding to each first indicator to obtain a database instance scoring model; the database instance scoring model is used to score the monitoring data of the input database instance, and the score is determined based on the database instance's score. This technical solution for fault detection incorporates core monitoring metrics of database instances into a scoring model. It also combines historical fault records to score the health status of the database instance, improving the ability to perceive, decide on, and execute fault detection measures. By simulating various performance monitoring metrics of the database instance and performing feature engineering calculations, it automatically selects important monitoring items and trains a multi-dimensional fault detection scoring model based on historical data. The scoring model has strong transferability, low reliance on expert experience, and can quickly pinpoint the root cause of a decline in the database instance score, providing a specific deduction for each root cause. Previously undetectable issues, often occurring weekly or monthly, can now be identified and addressed within a single inspection cycle, thus improving fault detection efficiency and providing fault early warning capabilities.

[0097] Figure 6 An exemplary system architecture 600 is shown, in which the fault detection method or the fault detection apparatus for a database instance can be applied according to embodiments of the present invention.

[0098] like Figure 6 As shown, system architecture 600 may include terminal devices 601, 602, and 603, a network 604, and a server 605. Network 604 serves as the medium for providing communication links between terminal devices 601, 602, and 603 and server 605. Network 604 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0099] Users can use terminal devices 601, 602, and 603 to interact with server 605 via network 604 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 601, 602, and 603, such as database applications, data processing applications, rating applications, data monitoring tools, etc. (for example only).

[0100] Terminal devices 601, 602, and 603 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.

[0101] Server 605 can be a server providing various services, such as a backend management server supporting shopping websites browsed by users using terminal devices 601, 602, and 603 (for example only). The backend management server can, upon receiving a fault detection request for a database instance, obtain monitoring data and historical fault records corresponding to monitoring indicators of the database instance within a specified time period; determine a time window for indicator processing based on the monitoring data and historical fault records; process the monitoring data according to the time window and match it with the historical fault records to filter a first indicator set from the monitoring indicator set, and determine the weight corresponding to each first indicator in the first indicator set; train a model based on the monitoring data and weight corresponding to each first indicator to obtain a database instance scoring model; use the database instance scoring model to score the input monitoring data of the database instance, perform fault detection and other processing based on the database instance's score, and feed back the processing results (e.g., the database instance's score, fault detection results—for example only) to the terminal device.

[0102] It should be noted that the database instance fault detection method provided in this embodiment of the invention is generally executed by server 605, and correspondingly, the database instance fault detection device is generally set in server 605.

[0103] It should be understood that Figure 6 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0104] The following is for reference. Figure 7 It shows a schematic diagram of the structure of a computer system 700 suitable for implementing terminal devices or servers of the present invention. Figure 7 The terminal device or server shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.

[0105] like Figure 7 As shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 702 or programs loaded from storage section 708 into random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the system 700. The CPU 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0106] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.

[0107] In particular, according to the embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 709, and / or installed from removable medium 711. When the computer program is executed by central processing unit (CPU) 701, it performs the functions defined above in the system of this invention.

[0108] It should be noted that the computer-readable medium shown in this invention can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable medium other than computer-readable storage media, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0109] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0110] The units or modules described in the embodiments of the present invention can be implemented in software or hardware. The described units or modules can also be housed in a processor; for example, a processor can be described as including a data acquisition module, a window selection module, an indicator filtering module, a model training module, and a scoring and detection module. The names of these units or modules do not necessarily limit the specific unit or module itself; for example, a data acquisition module can also be described as "a module for acquiring monitoring data and historical fault records corresponding to monitoring indicators of a database instance within a specified time period."

[0111] In another aspect, the present invention also provides a computer-readable medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable medium carries one or more programs, which, when executed by the device, cause the device to include: acquiring monitoring data and historical fault records corresponding to monitoring indicators of a database instance within a specified time period; determining a time window for indicator processing based on the monitoring data and the historical fault records; processing the monitoring data according to the time window and matching it with the historical fault records to filter a first indicator set from the monitoring indicator set, and determining the weight corresponding to each first indicator in the first indicator set; training a model based on the monitoring data and weight corresponding to each first indicator to obtain a database instance scoring model; using the database instance scoring model to score the input monitoring data of the database instance, and performing fault detection based on the database instance's score.

[0112] According to the technical solution of this embodiment of the invention, monitoring data and historical fault records corresponding to monitoring indicators of a database instance within a specified time period are obtained; a time window for indicator processing is determined based on the monitoring data and the historical fault records; the monitoring data is processed according to the time window and matched with the historical fault records to filter a first indicator set from the monitoring indicator set, and the weight corresponding to each first indicator in the first indicator set is determined; a model is trained based on the monitoring data and weight corresponding to each first indicator to obtain a database instance scoring model; the database instance scoring model is used to score the monitoring data of the input database instance, and the score is determined based on the database instance's score. This technical solution for fault detection incorporates core monitoring metrics of database instances into a scoring model. It also combines historical fault records to score the health status of the database instance, improving the ability to perceive, decide on, and execute fault detection measures. By simulating various performance monitoring metrics of the database instance and performing feature engineering calculations, it automatically selects important monitoring items and trains a multi-dimensional fault detection scoring model based on historical data. The scoring model has strong transferability, low reliance on expert experience, and can quickly pinpoint the root cause of a decline in the database instance score, providing a specific deduction for each root cause. Previously undetectable issues, often occurring weekly or monthly, can now be identified and addressed within a single inspection cycle, thus improving fault detection efficiency and providing fault early warning capabilities.

[0113] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for fault detection of a database instance, characterized in that, include: Retrieve monitoring data and historical fault records corresponding to the monitoring metrics of the database instance within a specified time period; The time window for indicator processing is determined based on the monitoring data and the historical fault records. The monitoring data is processed according to the time window and matched with the historical fault records to filter out the first indicator set from the monitoring indicator set, and the weight corresponding to each first indicator in the first indicator set is determined. The database instance scoring model is obtained by training the model based on the monitoring data and weights corresponding to each of the first indicators. The database instance scoring model is used to score the monitoring data of the input database instance, and fault detection is performed based on the database instance score. The monitoring data is processed according to the time window and matched with the historical fault records to filter out a first set of indicators from the monitoring indicator set, and the weight of each first indicator in the first set of indicators is determined. This includes: processing the monitoring data according to the time window to obtain indicator fluctuation regularity data; filtering out a first set of indicators from the monitoring indicator set based on the matching degree between the indicator fluctuation regularity data and the fault time points in the historical fault records; obtaining multi-indicator fluctuation regularity data of fault time points based on the historical fault records; and determining the weight of each first indicator in the first set of indicators based on the correlation between the multi-indicator fluctuation regularity data and the fault time points. Fault detection based on the database instance's score includes: when the current score of the database instance decreases compared to the previous score, performing anomaly detection on the specified monitoring data corresponding to each of the first indicators of the database instance using the isolated forest algorithm to obtain anomaly indicators, and using the anomaly indicators as the fault detection results.

2. The method according to claim 1, characterized in that, Before determining the time window for indicator processing based on the monitoring data and the historical fault records, the process also includes: Preprocess the monitoring data corresponding to the monitoring metrics of the database instance to perform missing value filling and exponential smoothing.

3. The method according to claim 1, characterized in that, The time window for indicator processing is determined based on the monitoring data and the historical fault records, including: The monitoring data is processed according to different initial time windows to obtain at least one indicator fluctuation regularity data; Calculate the matching degree between the fluctuation regularity data of each indicator and the fault time point in the historical fault record, and select the time window for indicator processing from the initial time window based on the matching degree.

4. The method according to claim 1, characterized in that, Before processing the monitoring data according to the time window, the process also includes: The monitoring indicators are pre-screened based on the monitoring data and the historical fault records to obtain a second indicator set from the monitoring indicator set, and the weight of each second indicator in the second indicator set is determined. Furthermore, the monitoring data is processed according to the time window and matched with the historical fault records to filter out a first set of indicators from the monitoring indicator set, including: The monitoring data corresponding to each second indicator in the second indicator set is processed according to the time window, and matched with the historical fault records to filter out the first indicator set from the second indicator set.

5. The method according to claim 1, characterized in that, Before processing the monitoring data according to the time window, the process also includes: Based on the failure types of the database instances that failed in the historical failure list, the obtained database instances are classified into healthy instances, high-load instances, and failed instances. The monitoring data is processed according to the time window and matched with the historical fault records to filter out a first set of indicators from the monitoring indicator set, and the weight corresponding to each first indicator in the first indicator set is determined, including: For each category of database instance, the monitoring data of the database instance of the category is processed according to the time window and matched with the historical fault records to filter out the first indicator set of the database instance of the category from the monitoring indicator set, and the weight corresponding to each first indicator in the first indicator set of the database instance of the category is determined. Based on the monitoring data and weights corresponding to each of the first indicators, a database instance scoring model is trained to obtain the following: The database instance scoring model is obtained by training the model based on the monitoring data and weights corresponding to each of the first indicators for each category of database instance.

6. The method according to claim 1, characterized in that, Based on the monitoring data and weights corresponding to each of the first indicators, a database instance scoring model is trained to obtain the following: The first indicator set is divided into multiple indicator groups according to the weight of each first indicator, and the average weight of each indicator group is calculated. Fluctuation features are extracted from the monitoring data corresponding to each indicator group, and time-series features are extracted from the monitoring data corresponding to the first indicator set. The fluctuation characteristics of each indicator group are weighted according to the weighted mean of each indicator group, and the weighted fluctuation characteristics are fused with the time series characteristics. The failure probability of the database instance is calculated based on the fused features to train the model and obtain the database instance scoring model.

7. The method according to claim 6, characterized in that, The first indicator set is divided into multiple indicator groups according to the weight of each first indicator, and the average weight of each indicator group is calculated, including: Sort the results from largest to smallest according to the weight value of each of the first indicators; The number of indicators in each indicator group is determined based on the number of the first indicators and the preset number of groups; Based on the sorting results and the number of indicators in each indicator group, the first indicators are combined in sequence to divide the first indicator set into multiple indicator groups; The average weight of each indicator group is calculated based on the weight value of the first indicator included in each indicator group.

8. The method according to claim 6, characterized in that, Fluctuation characteristics were extracted from the monitoring data corresponding to each indicator group, including: The monitoring data corresponding to each indicator group is input into different first convolutional sub-layers to obtain the initial fluctuation characteristics of each indicator group; The initial fluctuation characteristics of each indicator group are input into the second convolutional sub-layer, which is sequentially connected to the first convolutional sub-layer, to obtain the fluctuation characteristics of each indicator group.

9. The method according to claim 1, characterized in that, The specified monitoring data includes monitoring data from the most recent half hour, monitoring data from the half hour before and after the same period last year, and monitoring data from the half hour before and after the same period last month.

10. A fault detection device for a database instance, characterized in that, include: The data acquisition module is used to acquire monitoring data and historical fault records corresponding to the monitoring indicators of the database instance within a specified time period; The window selection module is used to determine the time window for index processing based on the monitoring data and the historical fault records. The indicator filtering module is used to process the monitoring data according to the time window, match it with the historical fault records to filter the first indicator set from the monitoring indicator set, and determine the weight corresponding to each first indicator in the first indicator set. The indicator filtering module is further configured to process the monitoring data according to the time window to obtain indicator fluctuation regularity data; filter the first indicator set from the monitoring indicator set according to the matching degree between the indicator fluctuation regularity data and the fault time points in the historical fault records; and obtain multi-indicator fluctuation regularity data of fault time points according to the historical fault records. Based on the correlation between the fluctuation regularity data of the multi-indicator and the time point of the failure, the weight corresponding to each first indicator in the first indicator set is determined; The model training module is used to train the model based on the monitoring data and weights corresponding to each of the first indicators to obtain the database instance scoring model. The scoring and detection module is used to score the monitoring data of the input database instance using the database instance scoring model, and to perform fault detection based on the database instance's score. The scoring and detection module is further configured to, when the current score of the database instance decreases compared to the previous score, perform anomaly detection on the specified monitoring data corresponding to each of the first indicators of the database instance using the isolated forest algorithm to obtain anomaly indicators, and use the anomaly indicators as fault detection results.

11. An electronic device for fault detection of a database instance, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-9.

12. A computer-readable medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Electronic equipment fault diagnosis system

    CN112101431A

  • Circuit breaker fault detection system and method based on CNN combined with LSTM algorithm

    CN114839523A