Database cluster anomaly detection method and device and electronic equipment
By performing unsupervised classification on the operating status data of the database cluster and supervised classification on the business attribute data, combined with a decision tree model, the problem of inaccurate detection in existing technologies is solved, automated anomaly detection is achieved, accuracy is improved, and the workload of operation and maintenance is reduced.
Patent Information
- Application Number
- CN202511203535.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-27
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-27
AI Technical Summary
Existing database cluster anomaly detection methods cannot effectively understand the impact of business characteristics on operating load, resulting in inaccurate detection when the load varies greatly in different time windows, and requiring a large amount of manual maintenance work.
By performing unsupervised classification on the operating status data of the database cluster, combining it with supervised classification of business attribute data, and using a decision tree model to analyze load levels, automated anomaly detection can be achieved.
It improves the accuracy of anomaly detection, reduces the workload of operation and maintenance, and can maintain detection accuracy when business characteristics cause large load changes. Automated processing reduces manual intervention.
Smart Images

Figure CN120724263A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic digital data processing, and in particular to a database cluster anomaly detection method, device and electronic equipment. Background Art
[0002] As the operating load of databases and database clusters becomes increasingly complex, business changes at different times, resulting in greater and greater variations in operating load at different times, making false alarms more likely to occur and making it more difficult to detect abnormal operating load status of databases / database clusters.
[0003] Existing classification and anomaly detection methods usually determine whether an anomaly is based on comparing single indicator data with a threshold, or perform graded judgments based on resource usage to determine whether an anomaly is present. They do not analyze the operating load classification from the perspective of business characteristics. Therefore, they cannot gain insight into the impact of business characteristics changing over time on the operating load, reducing the accuracy of load anomaly detection. Summary of the Invention
[0004] In view of the above problems, the present invention provides a method, device and electronic device for detecting anomalies in a database cluster.
[0005] According to a first aspect of the present invention, a database cluster anomaly detection method is provided, comprising: in response to an anomaly detection request for a database cluster in a target time window, performing unsupervised classification on operating status data within the target time window to obtain a first load level classification result, wherein the operating status data represents status data of the database cluster during operation; performing supervised classification on business attribute data of the target time window to obtain a second load level classification result, wherein the business attribute data represents the degree to which the target time window has a periodic impact on the category and frequency of business transactions; and determining a detection result based on the first load level classification result and the second load level classification result.
[0006] Optionally, the database cluster includes multiple database nodes; wherein, in response to an anomaly detection request for the database cluster in a target time window, unsupervised classification is performed on the operating status data within the target time window to obtain a first load level classification result, including: scene recognition is performed on the anomaly detection request to obtain a target detection scene; based on the scene mapping relationship, at least one operating status data corresponding to the target detection scene in the target time window is obtained from multiple database nodes respectively, wherein the scene mapping relationship represents the correspondence between the detection scene and the operating status data; at least one operating status data is merged to obtain merged data; and unsupervised classification is performed on the merged data to obtain a first load level classification result.
[0007] Optionally, obtaining at least one operating status data corresponding to the target detection scenario within the target time window includes: when the target detection scenario is an operating load detection scenario, obtaining the operating status data includes at least one of the following: the average number of database connections, the number of insert operations, the number of update operations, the number of delete operations, the number of precise queries, the number of fuzzy queries, and the number of controlled user access transactions; when the target detection scenario is a permission security detection scenario, obtaining the operating status data includes at least one of the following: the number of accessed network addresses, the number of successful logins, the number of unlocked users, the number of failed logins, and the number of locked users.
[0008] Optionally, the merged data includes the merged number of insert operations; wherein, merging processing is performed on at least one operation status data to obtain the merged data, including: summing the numbers of insert operations respectively obtained from multiple database nodes to obtain the merged number of insert operations.
[0009] Optionally, the business attribute data includes at least one of the following: maintenance period attribute data, batch processing period attribute data, holiday period attribute data, promotion period attribute data, attribute data of the first period, attribute data of the second period, and attribute data of the third period. The business transaction frequency of the first period is greater than the business transaction frequency of the second period, and the business transaction frequency of the second period is greater than the business transaction frequency of the third period.
[0010] Optionally, the attribute data of the second period is determined based on the following operation: statistically processing the login business sub-attribute data, the initial business sub-attribute data of the day, the normal business sub-attribute data and the daily settlement business sub-attribute data to obtain the attribute data of the second period.
[0011] Optionally, the second load level classification result is obtained by supervised classification based on a decision tree model, and the training of the decision tree model is determined based on the following operations: obtaining training samples, the training samples including sample business attribute data and sample operation status data corresponding to each of the multiple sample time windows; for each of the multiple sample time windows, performing unsupervised classification on the sample operation status data to obtain the sample first load level classification results corresponding to each of the multiple sample time windows, wherein the sample first load level classification results represent the load level pseudo-label of the sample business attribute data in the sample time window; training the initial decision tree model based on the sample business attribute data and the sample first load level classification results corresponding to each of the multiple sample time windows to obtain a decision tree model.
[0012] Optionally, determining the detection result based on the first load level classification result and the second load level classification result includes: when the first load level classification result is the same as the second load level classification result, obtaining a normal detection result; when the first load level classification result is different from the second load level classification result, obtaining an abnormal detection result, triggering a system warning.
[0013] The second aspect of the present invention provides a database cluster anomaly detection device, including: a first classification module, used to respond to an anomaly detection request for the database cluster in a target time window, perform unsupervised classification on the operating status data within the target time window, and obtain a first load level classification result, wherein the operating status data represents the status data of the database cluster during operation; a second classification module, used to perform supervised classification on the business attribute data of the target time window, and obtain a second load level classification result, wherein the business attribute data represents the degree to which the target time window has a periodic impact on the business transaction category and frequency; a detection module, used to determine the detection result based on the first load level classification result and the second load level classification result.
[0014] A third aspect of the present invention provides an electronic device comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned database cluster anomaly detection method.
[0015] According to the database cluster anomaly detection method, device and electronic device provided by the present invention, by introducing business attribute data, the business attribute data is classified to obtain a second load level classification result, and the load anomaly is detected by comparing the first load level classification result and the second load level classification result of the target time window, thereby ensuring that the accuracy of anomaly detection is maintained even when the business attribute data causes a large load change, and solving the technical problem that the business characteristics cause the database operation load to vary greatly in different time windows, resulting in inaccurate anomaly detection; in addition, by performing unsupervised learning on the operation status data, automatic classification is achieved, reducing the operation and maintenance workload. By performing supervised classification on the business attribute data, the impact of various types of business attribute data on the operation load can be understood through the decision path. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above and other objects, features and advantages of the present invention will become more apparent from the following description of the embodiments of the present invention with reference to the accompanying drawings.
[0017] Figure 1 The application scenario of the database cluster anomaly detection method and device according to the embodiment of the present invention is shown.
[0018] Figure 2A flowchart of a database cluster anomaly detection method according to an embodiment of the present invention is shown.
[0019] Figure 3 A schematic diagram of an anomaly detection process according to an embodiment of the present invention is shown.
[0020] Figure 4 A structural block diagram of a database cluster anomaly detection device according to an embodiment of the present invention is shown.
[0021] Figure 5 A block diagram of an electronic device suitable for implementing a database cluster anomaly detection method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0022] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present invention. In the following detailed description, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of embodiments of the present invention. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessary confusion of the concept of the present invention.
[0023] The terms used herein are only for describing specific embodiments and are not intended to limit the present invention. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0024] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0025] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0026] In the process of implementing the present invention, it was found that there are two main types of existing classification and anomaly detection methods. One is rule-based classification and anomaly detection, such as: when a certain value is within a certain range, it is a state, when a certain value is greater than a specified threshold, an alarm is issued, etc.; the other is machine learning-based classification and anomaly detection, which is usually based on time series anomaly detection, and is generally aimed at resource usage classification, rather than operating load classification.
[0027] In this scenario, the main problems of existing anomaly detection methods are: (1) Rule-based classification and anomaly detection cannot be dynamically adjusted according to the operating environment. If the threshold is set too low, it is easy to have false positives, and if it is set too high, it is easy to have missed negatives. It is also difficult to make comprehensive judgments on more items, and manual operation and maintenance are required, which is also a lot of work. (2) Classification and anomaly detection based on machine learning: Usually, isolated points are detected based on time series, or supervised learning is performed based on operating data items. Neither of them considers the correlation with business characteristics. In fact, the database operating load is greatly related to business characteristics. For example, during the maintenance window, the operating load may drop by 50% (load balancing between two databases, maintaining one of the databases) or drop to 100% (service termination during maintenance). Usually, the load at night is much smaller than that during the day, but there may be batch processing business in a certain time window at night, and the load may be larger, but the load type is different. Therefore, traditional anomaly detection based on machine learning is quite inaccurate for scenarios where the business changes greatly in each time window. (3) The above anomaly detection methods all require manual participation, and the maintenance workload is relatively large. (4) The above anomaly detection methods are all performed on a single database. Even for a database cluster, if there are multiple computing service nodes that do not provide a unified interface, it is impossible to monitor all nodes. For example, when the system uses multiple databases with separate libraries and tables to implement data storage and access, it is impossible to logically unify these databases into a single database for processing. Therefore, the above methods do not analyze the operating load classification and perform anomaly detection from the perspective of business characteristics, and thus cannot provide insight into the impact of business characteristics on the operating load.
[0028] In view of this, the present invention provides a database cluster anomaly detection method, apparatus, and electronic device. The method comprises: in response to an anomaly detection request for a database cluster in a target time window, performing unsupervised classification on operating status data within the target time window to obtain a first load level classification result, wherein the operating status data represents the state data of the database cluster during operation; performing supervised classification on business attribute data in the target time window to obtain a second load level classification result, wherein the business attribute data represents the degree to which the target time window periodically affects the type and frequency of business transactions; and determining a detection result based on the first load level classification result and the second load level classification result.
[0029] In the technical solution of the present invention, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0030] Figure 1 The application scenario of the database cluster anomaly detection method and device according to the embodiment of the present invention is shown.
[0031] like Figure 1 As shown, the business system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0032] The user may use at least one of the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0033] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.
[0034] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using first terminal device 101, second terminal device 102, and third terminal device 103 (this is just an example). A database cluster is deployed on server 105. Server 105 includes multiple server nodes that can analyze and process received data such as user requests and feedback the processing results (such as web pages, information, or data obtained or generated based on user requests) to the terminal devices.
[0035] It should be noted that the database cluster anomaly detection method provided in the embodiments of the present invention can generally be executed by the server 105. Accordingly, the database cluster anomaly detection apparatus provided in the embodiments of the present invention can generally be set in the server 105. The database cluster anomaly detection method provided in the embodiments of the present invention can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the database cluster anomaly detection apparatus provided in the embodiments of the present invention can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105.
[0036] Alternatively, the database cluster anomaly detection method provided in the embodiment of the present invention may also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or may also be executed by another terminal device different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the database cluster anomaly detection apparatus provided in the embodiment of the present invention may also be provided in the first terminal device 101, the second terminal device 102, or the third terminal device 103, or may be provided in another terminal device different from the first terminal device 101, the second terminal device 102, or the third terminal device 103.
[0037] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0038] It should be noted that the sequence numbers of the operations in the following method are only used to indicate the operation for the purpose of description, and should not be regarded as indicating the order in which the operations should be performed. Unless explicitly stated, the method does not need to be performed in the order shown.
[0039] Figure 2 A flowchart of a database cluster anomaly detection method according to an embodiment of the present invention is shown.
[0040] like Figure 2 As shown, the database cluster anomaly detection method includes operations S210 to S230.
[0041] In operation S210 , in response to an anomaly detection request for a database cluster in a target time window, unsupervised classification is performed on the operating status data in the target time window to obtain a first load level classification result.
[0042] Optionally, the database cluster includes multiple database deployments, which can be 1 to N centralized databases and 0 to N database clusters, where N is a positive integer.
[0043] Optionally, the database cluster includes multiple database nodes, each of which stores real-time data generated based on business transactions.
[0044] Optionally, the running status data represents the status data of the database cluster during operation. The server periodically reads the running data of each data item from the database cluster and performs statistical analysis according to a time window to obtain the running status data.
[0045] Optionally, the database cluster sets commonly used data items and can also support the configuration or addition of custom data items based on actual required scenarios.
[0046] For example, the data items include application login items, login failure items, and query items, and the running status data includes the number of application logins, the number of login failures, the number of queries, and the like.
[0047] Optionally, the target time window is the time window to be tested. For example, if the state of the database cluster is tested from 8:00 to 9:00, the target time window is from 8:00 to 9:00.
[0048] Optionally, the number of time windows in a day refers to the number of time windows set in a day. The default setting is 24, which means the time window is 1 hour.
[0049] Optionally, in response to the anomaly detection request for the database cluster in the target time window, the server may perform unsupervised classification on the operating status data in the target time window based on an unsupervised classification algorithm to obtain a first load level classification result.
[0050] For example, an unsupervised classification algorithm can be a clustering algorithm, such as the SimpleKMeans or SKMeans clustering algorithm. Algorithm parameters include the number of categories, which is set based on the load categories of different business time windows. The training parameters are adjusted based on the accuracy of actual training results. The number of categories is set to eight to account for the following eight possible load categories: normal business period load category, off-peak special period load category, peak special period load category, business processing period load category, maintenance period load category, promotion day normal load category, promotion day peak load category, and holiday load category. The business processing time load category represents the load category corresponding to the batch processing time window.
[0051] Optionally, the number of load categories can be set based on actual circumstances. Furthermore, after machine learning based on a large number of training samples, the number of load categories obtained is not limited to these eight. Furthermore, the load categories obtained based on machine learning algorithms and training data may not necessarily correspond exactly to the eight load categories described above, which is normal.
[0052] Optionally, the first load level classification result includes a service load category corresponding to the target time window. For example, unsupervised classification is performed on the operating status data between 8:00 and 9:00, and the first load level classification result is a low-peak special period load category.
[0053] In operation S220 , supervised classification is performed on the service attribute data of the target time window to obtain a second load level classification result.
[0054] Optionally, the business attribute data represents the degree to which the target time window has a periodic impact on the business transaction category and frequency.
[0055] Optionally, there may be special services within a certain period, such as batch processing services during the time window of 24:00-1:00, but not during normal daytime trading hours; during maintenance periods, such as Saturday at 24:00, there may be maintenance services or flow control services.
[0056] For example, business attribute data can be date attributes, time window sorting attributes, week attributes, holiday attributes, promotion day attributes, time period transaction frequency attributes, maintenance period attributes, etc. Among them, the time window sorting attribute is the time window in which the target time window is located in a day, and the time period transaction frequency attribute indicates whether the time period corresponding to the target time window is a peak trading period, a low-peak trading period, or a normal trading period.
[0057] Optionally, a supervised classification algorithm may be used to perform supervised classification on the service attribute data of the target time window to obtain a second load level classification result.
[0058] For example, the supervised classification algorithm may be a decision tree algorithm.
[0059] Optionally, the second load level classification result represents the service load category corresponding to the target time window. For example, supervised classification is performed on service attribute data between 9:00 and 10:00, and the second load level classification result is the normal service period load category.
[0060] In operation S230 , a detection result is determined according to the first load level classification result and the second load level classification result.
[0061] Optionally, when the first load level classification result and the second load level classification result are the same, the detection result is a normal load state.
[0062] Optionally, when the first load level classification result and the second load level classification result are different, the detection result is an abnormal load state, and an early warning is triggered.
[0063] Optionally, business attribute data is highly correlated with the database cluster's operational load changes over different time windows. For example, during a promotional day window, the operational load might increase by 50% due to frequent transaction activity; during a maintenance window, the operational load might decrease by 50% or even reach 100% (with service interruption during maintenance). However, these conditions are considered normal during promotional day and maintenance window periods; however, they are considered abnormal during normal business hours. Therefore, by incorporating business attribute data, anomaly detection and verification of database load can be performed using both operational status data and business attribute data, improving detection accuracy.
[0064] For example, if the amount of operating status data is large, the first load level classification result is the promotion day peak load category; if the second load level classification result is the normal business period load category, the database cluster is in an abnormal load state; if the second load level classification result is the promotion day peak load category, the database cluster is in a normal load state.
[0065] Optionally, due to the introduction of business attribute data, by classifying the business attribute data, a second load level classification result is obtained, and load anomalies are detected by comparing the first load level classification result and the second load level classification result of the target time window, thereby ensuring that the accuracy of anomaly detection is maintained even when the load changes greatly due to business attribute data. This solves the technical problem of inaccurate anomaly detection caused by large changes in the database operating load in different time windows due to business characteristics. In addition, by performing unsupervised learning on the operating status data for classification, automatic classification is achieved, reducing the operation and maintenance workload. By performing supervised classification on business attribute data, the impact of various types of business attribute data on the operating load can be understood through the decision path.
[0066] Optionally, the database cluster includes multiple database nodes; wherein, in response to an anomaly detection request for the database cluster in a target time window, unsupervised classification is performed on the operating status data within the target time window to obtain a first load level classification result, including: scene recognition is performed on the anomaly detection request to obtain a target detection scene; based on the scene mapping relationship, at least one operating status data corresponding to the target detection scene in the target time window is obtained from multiple database nodes respectively, wherein the scene mapping relationship represents the correspondence between the detection scene and the operating status data; at least one operating status data is merged to obtain merged data; and unsupervised classification is performed on the merged data to obtain a first load level classification result.
[0067] Optionally, scenario identification includes identification of business category scenarios, data security detection scenarios, etc.
[0068] Optionally, scenario recognition is performed on the anomaly detection request to obtain a target detection scenario, which may be a high-frequency trading business detection scenario.
[0069] Optionally, the server responds to the anomaly detection request and determines that the target detection scenario is a high-frequency trading business detection scenario, and then obtains the operating status data associated with the high-frequency trading business detection scenario in the target time window from multiple database nodes based on the scenario mapping relationship.
[0070] Optionally, the database cluster stores multiple data items, and a scene mapping relationship between the detection scene and the data items is constructed. The running state data is obtained by statistically analyzing the data items within the time window, thereby obtaining the scene mapping relationship between the detection scene and the running state data.
[0071] For example, operating status data A1 of data item A and operating status data B1 of data item B, which are associated with the high-frequency trading business detection scenario within the target time window, are obtained from the first database node. Operating status data A2 of data item A and operating status data B2 of data item B, which are associated with the high-frequency trading business detection scenario within the target time window, are obtained from the second database node. The operating status data A1 and A2 of data item A are merged to obtain the merged data of data item A. The operating status data B1 and B2 of data item B are merged to obtain the merged data of data item B.
[0072] Optionally, an unsupervised classification algorithm is used to perform unsupervised classification on the merged data to obtain a first load level classification result of the database cluster in the target time window.
[0073] Optionally, the operation status data is collected and time window statistics are performed to form an operation status data set, and an unsupervised classification algorithm is used for training to obtain the classification results of each operation status data in the data set; in addition, by collecting statistics according to the time window, the operation data characteristics of the specific window can be effectively divided, and the division rules of the time window can be customized to effectively cooperate with the business attribute data. For example, 24 time windows are set in a day, one time window per hour, so as to effectively obtain the operation status data every hour, so as to determine whether the operation status every hour is abnormal, and the system can run automatically without human intervention, thereby improving the classification efficiency.
[0074] Optionally, obtaining at least one operating status data corresponding to the target detection scenario within the target time window includes: when the target detection scenario is an operating load detection scenario, obtaining at least one of the following: the average number of database connections, the number of insert operations, the number of update operations, the number of delete operations, the number of precise queries, the number of fuzzy queries, and the number of controlled user access transactions; when the target detection scenario is a permission security detection scenario, obtaining operating status data includes at least one of the following: the number of accessed network addresses, the number of successful logins, the number of unlocked users, the number of failed logins, and the number of locked users.
[0075] Optionally, you can configure data items in the database cluster based on the actual scenario, or add data items to obtain statistics on data items corresponding to the scenario and obtain operational status data. When adding data items, you need to specify the collection method.
[0076] Optionally, the load detection scenario is to detect a scenario in which the database cluster has an abnormal load.
[0077] Optionally, the average number of database connections is the average number of connections to the database cluster within the target time window.
[0078] Optionally, the number of insert operations is the number of data insert operations performed in the database cluster within the target time window.
[0079] Optionally, the number of update operations is the number of data update operations performed in the database cluster within the target time window.
[0080] Optionally, the number of delete operations is the number of data delete operations performed in the database cluster within the target time window.
[0081] Optionally, the number of controlled user access transactions is the number of data definition language (DDL) operations or data control language (DCL) operations performed in the database cluster within the target time window.
[0082] For example, when at least one of the operating status data such as the average number of database connections, the number of insert operations, the number of update operations, the number of delete operations, and the number of controlled user access transactions is greater than the first preset threshold, the first load level classification result corresponding to this target time window is more likely to belong to the peak special period load category or the promotion day peak load category.
[0083] For example, if the sum of the average number of database connections, insert operations, update operations, delete operations, and user access control transactions is less than a second preset threshold, the first load level classification result corresponding to the target time window is likely to belong to the off-peak special period load category. The first preset threshold is much greater than the second preset threshold.
[0084] Optionally, the permission security detection scenario is a scenario of detecting security anomalies of access permission of a database cluster.
[0085] Optionally, the number of successful logins is the number of successful logins to the database cluster within the target time window.
[0086] Optionally, when the target detection scenario is a permission security detection scenario, the acquired operation status data may further include the number of applications associated with the database cluster, the number of user accesses, and the like.
[0087] Optionally, operational status data can be flexibly configured and adjusted based on actual scenarios, enabling more accurate characterization of the desired operational status for anomaly detection, rather than focusing on a single scenario or overly broad detection. Flexible extensions are also provided to detect user-defined operational scenarios.
[0088] Optionally, the merged data includes the merged number of insert operations; wherein, merging processing is performed on at least one operation status data to obtain the merged data, including: summing the numbers of insert operations respectively obtained from multiple database nodes to obtain the merged number of insert operations.
[0089] Optionally, for each data item in the database cluster, the storage data corresponding to the data item is obtained from multiple database nodes based on the target time window for statistical analysis to obtain the operating status data of the data items corresponding to each of the multiple database nodes, thereby merging the multiple operating status data to obtain merged data.
[0090] For example, for the data item: insert operation, the number of insertions corresponding to the insert operation between 9:00 and 10:00 is obtained from the three database nodes for statistical analysis, and the number of insert operations corresponding to each of the three database nodes is obtained. Then, the number of insert operations of each of the three database nodes is summed up to obtain the merged number of insert operations.
[0091] For example, for the data item: successful login operation, the number of successful login operations corresponding to the successful login operations between 9:00 and 10:00 is obtained from the three database nodes for statistical analysis, and the number of successful login operations corresponding to each of the three database nodes is obtained. The number of successful login operations of each of the three database nodes is then summed up to obtain the merged number of successful login operations.
[0092] For example, for the data item: database connection operation, the number of times corresponding to the database connection operation between 9:00 and 10:00 is obtained from the three database nodes for statistical analysis, and the number of database connection operations corresponding to each of the three database nodes is obtained, and then the number of database connection operations of each of the three database nodes is averaged to obtain the average number of database connection operations.
[0093] Optionally, the database cluster is a cluster of multiple database nodes. The server periodically collects operational status data corresponding to the same data item from multiple database nodes according to a time window and merges the data to obtain merged data. When training an unsupervised machine learning model, the system periodically uses the historical merged data from the historical time window as a training dataset for periodic classification training of the unsupervised machine learning model.
[0094] Optionally, data collection supports configuring multiple database nodes. The operating status data corresponding to the same data item collected by multiple database nodes can be merged into a logical database node to ensure unified detection of database clusters or databases, which is beneficial for evaluating scenarios where the system uses database clusters / databases.
[0095] Optionally, the business attribute data includes at least one of the following: maintenance period attribute data, batch processing period attribute data, holiday period attribute data, promotion period attribute data, attribute data of the first period, attribute data of the second period, and attribute data of the third period. The business transaction frequency of the first period is greater than the business transaction frequency of the second period, and the business transaction frequency of the second period is greater than the business transaction frequency of the third period.
[0096] Optionally, the maintenance window attribute data indicates whether the database cluster is in a maintenance window during the target time window. For example, every Saturday from 23:00 to 24:00 is an online maintenance window, and the last weekend of each month is an offline maintenance window. Alternatively, a specific date, time period, and maintenance type can be specified.
[0097] Optionally, the batch processing period attribute data indicates whether the target time window is in a batch transaction processing period.
[0098] Optionally, the holiday period attribute data indicates whether the date attribute of the target time window is a statutory holiday.
[0099] Optionally, the promotion period attribute data indicates whether the date attribute of the target time window is a promotion day or whether the target time window is a promotion period.
[0100] Optionally, the first time period may be a target time window that is a peak trading period of the day.
[0101] Optionally, the second time period may be a target time window that is a normal trading period of the day.
[0102] Optionally, the third time period may be a target time window that is a low-peak trading period of the day.
[0103] Optionally, the business attribute data may also include date attribute data, day of the week attribute data, and custom time period attribute data. For example, custom time period attribute data may represent the night batch processing business period, or any custom attribute type. For example, the custom time period attribute data for the 22:00-23:00 period represents the daily report generation business period; the custom time period attribute data for the 23:00-24:00 period represents the data synchronization business period.
[0104] Optionally, each attribute data can be represented by a mask. For example, the holiday period attribute data of 1 represents Mid-Autumn Festival, and the holiday period attribute data of 1 represents National Day. 9:00-10:00 on xx / xx / xx is both Mid-Autumn Festival and National Day, so the holiday period attribute data is represented by 1+2=3.
[0105] Optionally, the server periodically collects business attribute data according to a time window. When training a supervised machine learning model, the system periodically uses the business attribute data corresponding to the historical time window as a training data set to periodically classify and train the supervised machine learning model.
[0106] Optionally, business attribute data such as maintenance period attribute data, batch period attribute data, holiday period attribute data, and promotion period attribute data can be configured for each time window, which reduces storage space, reduces calculation dimensions and amount, and can also handle multi-type overlapping problems. In addition, classification is performed based on rich business attribute data to determine whether the load is abnormal, avoiding the problem of false alarm detection when the load of each time window changes greatly due to business characteristics.
[0107] Optionally, the attribute data of the second period is determined based on the following operation: statistically processing the login business sub-attribute data, the initial business sub-attribute data of the day, the normal business sub-attribute data and the daily settlement business sub-attribute data to obtain the attribute data of the second period.
[0108] Optionally, the second time period can be divided into a login business period, an initial business period of the day, a normal business period, and a daily settlement business period.
[0109] Optionally, a time window can represent multiple time period types. A time period type represents a business type with a high transaction frequency or a custom-configured business type within that time period.
[0110] For example, the 8:00-9:00 time period can represent the login business period, the daily business period and the normal business period at the same time; the 10:00-11:00, 13:00-14:00, 14:00-15:00 and 15:00-16:00 time periods represent the normal business period; and the 17:00-18:00 time period can represent the daily settlement business period and the normal business period at the same time.
[0111] Optionally, the login service sub-attribute data is 1 to represent a login service period, and the login service sub-attribute data is 0 to represent a non-login service period.
[0112] Optionally, the initial business sub-attribute data of the day is 1 to represent the initial business period of the day, and the initial business sub-attribute data of the day is 0 to represent a non-initial business period of the day.
[0113] Optionally, the normal business sub-attribute data is 1 to represent a normal business period, and the normal business sub-attribute data is 0 to represent an abnormal business period.
[0114] Optionally, the daily settlement business sub-attribute data is 1 to represent the daily settlement business period, and the daily settlement business sub-attribute data is 0 to represent the non-daily settlement business period.
[0115] For example, the time period of 8:00-9:00 can represent the login business period, the daily business period, and the normal business period, and the attribute data of the time period of 8:00-9:00 is 1110. Based on the attribute data 1110 of the second time period of 8:00-9:00, supervised classification is performed, and the second load level classification result corresponding to the second time period of 8:00-9:00 is the normal business period load category.
[0116] Optionally, the second load level classification result is obtained by supervised classification based on a decision tree model, and the training of the decision tree model is determined based on the following operations: obtaining training samples, the training samples including sample business attribute data and sample operation status data corresponding to each of the multiple sample time windows; for each of the multiple sample time windows, performing unsupervised classification on the sample operation status data to obtain the sample first load level classification results corresponding to each of the multiple sample time windows, wherein the sample first load level classification results represent the load level pseudo-label of the sample business attribute data in the sample time window; training the initial decision tree model based on the sample business attribute data and the sample first load level classification results corresponding to each of the multiple sample time windows to obtain a decision tree model.
[0117] Optionally, the sample time window is a historical time window. The system periodically obtains sample business attribute data and sample operational status data corresponding to the historical time window. Sample business attribute data indicates the degree to which the sample time window periodically affects the transaction type and frequency. Sample operational status data indicates the status data of the database cluster during each sample time window during operation. Each sample time window is configured to be 1 hour.
[0118] Optionally, for each sample time window, unsupervised classification is performed on the sample operation status data to obtain a sample first load level classification result corresponding to each sample time window.
[0119] Optionally, the first load level classification result of the sample predicted by unsupervised classification is used as the load level pseudo-label corresponding to the sample time window, and the initial decision tree model is trained in combination with the sample business attribute data, and the parameters of the initial decision tree model are adjusted to obtain the decision tree model.
[0120] For example, if the sample time window is 10:00-11:00, and the first load level classification result of the sample is the normal business period load category, the normal business period load category is used as the load level pseudo label for the sample time window 10:00-11:00.
[0121] Optionally, the initial decision tree model may be constructed based on an improved decision tree algorithm (J48 Decision Tree Algorithm).
[0122] Optionally, an initial decision tree model is used to process sample business attribute data corresponding to each of multiple sample time windows to obtain a second load level classification result of the sample. The verification success rate is calculated based on the second load level classification result and the first load level classification result of the sample as a load level pseudo-label. The model parameters are adjusted and optimized based on the verification success rate to obtain a trained initial decision tree model.
[0123] Optionally, the business attribute data is classified using a supervised machine learning model, which may include but is not limited to a decision tree model.
[0124] Optionally, a supervised machine learning model is used to classify business attribute data. During the training process of the supervised machine learning model, the first load level classification results of the samples in the sample time window are used as pseudo-labels for the load levels. The first load level classification results are applied to the business attribute data as classification attributes, avoiding manual maintenance of the classification, thereby achieving automatic operation and reducing the workload of operation and maintenance. The second load level classification results are inferred from the business attribute data in the supervised machine learning time window. The supervised machine learning model uses a decision tree model to clearly see the decision path, facilitating analysis of the operation status. The load level classification results reveal the current operating status of the database cluster at a macro level. By observing the decision-making process, one can gain insight into the impact of various types of business attribute data on the database load.
[0125] Optionally, determining the detection result based on the first load level classification result and the second load level classification result includes: when the first load level classification result is the same as the second load level classification result, obtaining a normal detection result; when the first load level classification result is different from the second load level classification result, obtaining an abnormal detection result, triggering a system warning.
[0126] Optionally, by comparing the first load level classification result with the second load level classification result, if the first load level classification result is the same as the second load level classification result, a normal detection result is obtained. The normal detection result means that the database cluster is in a normal load state after comparison and verification from both business attributes and operating status.
[0127] Optionally, when the first load level classification result is different from the second load level classification result, an abnormal detection result is obtained, triggering a system warning.
[0128] Optionally, the anomaly detection result represents that the database cluster is in an abnormal load state, as determined by comparing and verifying the business attributes and the operating status.
[0129] For example, the first load level classification result is the peak special period load category, and the second load level classification result is the normal business period load category, that is, a high load abnormality occurs during the normal business period.
[0130] Optionally, the first load level classification result, the second load level classification result, and the detection result are stored in a database cluster for subsequent analysis. For abnormal detection results, manual vulnerability repair can also be performed.
[0131] Optionally, by using both unsupervised machine learning models and supervised machine learning models to predict load level classification results and then performing matching and comparison, time windows of similar business attribute data and abnormal operating status data can be effectively discovered, and then the model can be corrected and periodically trained to ensure a mechanism for continuous improvement in anomaly detection accuracy.
[0132] Figure 3 A schematic diagram of an anomaly detection process according to an embodiment of the present invention is shown.
[0133] like Figure 3 As shown, the database cluster includes database node Node1, database node Node2, and database node Node3. The operating status data is collected from database node Node1, database node Node2, and database node Node3 according to the time window, and the operating status data is merged. The operating status data is classified based on the unsupervised learning model to obtain the first load level classification result. The business attribute data can be customized for the business feature library, and the business attribute data is collected from the business feature library according to the time window and saved. The business attribute data is input into the supervised learning model for classification to obtain the second load level classification result. The first load level classification result and the second load level classification result are compared to obtain the detection result. The operating status data, business attribute data, the first load level classification result, the second load level classification result, and the detection result are saved to the database.
[0134] Based on the above database cluster anomaly detection method, the present invention also provides a database cluster anomaly detection device. Figure 4 The device is described in detail.
[0135] Figure 4 A structural block diagram of a database cluster anomaly detection device according to an embodiment of the present invention is shown.
[0136] like Figure 4 As shown, the database cluster anomaly detection device 400 of this embodiment includes a first classification module 410 , a second classification module 420 and a detection module 430 .
[0137] First classification module 410 is configured to, in response to an anomaly detection request for the database cluster within a target time window, perform unsupervised classification on the operating status data within the target time window to obtain a first load level classification result, where the operating status data represents the state data of the database cluster during operation. In one embodiment, first classification module 410 can be configured to perform operation S210 described above and will not be further described here.
[0138] Second classification module 420 is configured to perform supervised classification on the service attribute data of the target time window to obtain a second load level classification result, where the service attribute data indicates the degree to which the target time window periodically impacts the type and frequency of service transactions. In one embodiment, second classification module 420 may be configured to perform operation S220 described above, and will not be further described here.
[0139] The detection module 430 is configured to determine a detection result based on the first load level classification result and the second load level classification result. In one embodiment, the detection module 430 may be configured to perform the operation S230 described above, which will not be described in detail herein.
[0140] Optionally, the first classification module 410 includes a first classification submodule, a second classification submodule, a third classification submodule and a fourth classification submodule.
[0141] The first classification submodule is used to perform scene recognition on the anomaly detection request to obtain the target detection scene.
[0142] The second classification submodule is used to obtain at least one operating status data corresponding to the target detection scene in the target time window from multiple database nodes based on the scene mapping relationship, wherein the scene mapping relationship represents the correspondence between the detection scene and the operating status data.
[0143] The third classification submodule is used to merge the at least one running status data to obtain merged data.
[0144] The fourth classification submodule is used to perform unsupervised classification on the merged data to obtain a first load level classification result.
[0145] Optionally, the second classification submodule includes a first classification unit and a second classification unit.
[0146] The first classification unit is used to obtain operation status data including at least one of the following when the target detection scenario is an operation load detection scenario: the average number of database connections, the number of insert operations, the number of update operations, the number of delete operations, the number of precise queries, the number of fuzzy queries, and the number of controlled user access transactions.
[0147] The second classification unit is used to obtain operation status data including at least one of the following when the target detection scenario is a permission security detection scenario: the number of accessed network addresses, the number of successful logins, the number of unlocked users, the number of failed logins, and the number of locked users.
[0148] Optionally, the third classification submodule includes a third classification unit.
[0149] The third classification unit is used to sum the numbers of insert operations respectively obtained from multiple database nodes to obtain a merged number of insert operations.
[0150] Optionally, the second classification module 420 includes an acquisition submodule, a obtaining submodule and a training submodule.
[0151] The acquisition submodule is used to acquire training samples, where the training samples include sample business attribute data and sample operation status data corresponding to each of a plurality of sample time windows.
[0152] A submodule is obtained, which is used to perform unsupervised classification on the sample operation status data for each sample time window of multiple sample time windows, and obtain the sample first load level classification results corresponding to each of the multiple sample time windows, wherein the sample first load level classification results represent the load level pseudo-label of the sample business attribute data in the sample time window.
[0153] The training submodule is used to train an initial decision tree model according to the sample business attribute data corresponding to each of the multiple sample time windows and the sample first load level classification results to obtain a decision tree model.
[0154] Optionally, the detection module 430 includes a first detection submodule and a second detection submodule.
[0155] The first detection submodule is configured to obtain a normal detection result when the first load level classification result is the same as the second load level classification result.
[0156] The second detection submodule is used to obtain an abnormal detection result and trigger a system warning when the first load level classification result is different from the second load level classification result.
[0157] Optionally, any multiple modules among the first classification module 410, the second classification module 420, and the detection module 430 may be combined into a single module, or any one of these modules may be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules may be combined with at least part of the functionality of other modules and implemented in a single module. Optionally, at least one of the first classification module 410, the second classification module 420, and the detection module 430 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or may be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or may be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination thereof. Alternatively, at least one of the first classification module 410, the second classification module 420, and the detection module 430 may be at least partially implemented as a computer program module that, when executed, performs the corresponding functionality.
[0158] Figure 5 A block diagram of an electronic device suitable for implementing a database cluster anomaly detection method according to an embodiment of the present invention is shown.
[0159] Figure 5 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.
[0160] like Figure 5 As shown, a computer electronic device 500 according to an embodiment of the present invention includes a processor 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage portion 508 into a random access memory (RAM) 503. The processor 501 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 501 may also include onboard memory for caching purposes. The processor 501 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0161] Various programs and data required for the operation of the electronic device 500 are stored in the RAM 503. The processor 501, ROM 502, and RAM 503 are connected to each other via a bus 504. The processor 501 executes the programs in the ROM 502 and / or RAM 503 to perform various operations according to the method flow of the embodiment of the present invention. It should be noted that the programs may also be stored in one or more memories other than the ROM 502 and RAM 503. The processor 501 may also execute the programs stored in one or more memories to perform various operations according to the method flow of the embodiment of the present invention.
[0162] Optionally, electronic device 500 may further include an input / output (I / O) interface 505, which is also connected to bus 504. Electronic device 500 may also include one or more of the following components connected to I / O interface 505: an input section 506 including a keyboard, mouse, etc.; an output section 507 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 508 including a hard disk; and a communication section 509 including a network interface card such as a LAN card or modem. Communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to I / O interface 505 as needed. Removable media 511, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 510 as needed, so that computer programs read from the removable media can be installed into storage section 508 as needed.
[0163] Optionally, the method flow according to an embodiment of the present invention can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the processor 501, the above-mentioned functions defined in the system of the embodiment of the present invention are performed. Optionally, the system, device, apparatus, module, unit, etc. described above can be implemented by a computer program module.
[0164] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, which, when executed, implement the database cluster anomaly detection method according to the embodiments of the present invention.
[0165] Optionally, the computer-readable storage medium may be a non-volatile computer-readable storage medium. Examples include, but are not limited to, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0166] For example, optionally, the computer-readable storage medium may include the ROM 502 and / or RAM 503 described above and / or one or more memories other than the ROM 502 and RAM 503 .
[0167] An embodiment of the present invention also includes a computer program product, which includes a computer program, which contains program code for executing the method provided by the embodiment of the present invention. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the database cluster anomaly detection method provided by the embodiment of the present invention.
[0168] When the computer program is executed by the processor 501, the above functions defined in the system / device of the embodiment of the present invention are performed. Optionally, the above-described systems, devices, modules, units, etc. can be implemented by computer program modules.
[0169] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 509, and / or installed from a removable medium 511. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0170] Optionally, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, Java, C++, Python, "C" language, or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0171] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram may represent a module, program segment, or portion of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the boxes may occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession may actually be executed substantially in parallel, or they may sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, as well as the combination of boxes in the block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or may be implemented using a combination of dedicated hardware and computer instructions. It will be understood by those skilled in the art that the features described in the various embodiments of the present invention may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in the present invention. In particular, the features described in the various embodiments of the present invention may be combined and / or coupled in various ways without departing from the spirit and teachings of the present invention. All such combinations and / or combinations fall within the scope of the present invention.
[0172] The above describes embodiments of the present invention. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.
Claims
1. A database cluster anomaly detection method, characterized in that: The method comprises: In response to an anomaly detection request for a database cluster in a target time window, performing unsupervised classification on operating status data in the target time window to obtain a first load level classification result, wherein the operating status data represents status data of the database cluster during operation; Performing supervised classification on the business attribute data of the target time window to obtain a second load level classification result, wherein the business attribute data represents the degree to which the target time window has a periodic impact on the business transaction category and frequency; A detection result is determined according to the first load level classification result and the second load level classification result.
2. The method according to claim 1, characterized in that The database cluster includes a plurality of database nodes; The step of performing unsupervised classification on the operating status data within the target time window in response to the anomaly detection request for the database cluster in the target time window to obtain a first load level classification result includes: Performing scene recognition on the anomaly detection request to obtain a target detection scene; Based on the scene mapping relationship, respectively obtain at least one operating status data corresponding to the target detection scene within the target time window from the plurality of database nodes, wherein the scene mapping relationship represents the correspondence between the detection scene and the operating status data; Merging at least one of the operating status data to obtain merged data; Unsupervised classification is performed on the merged data to obtain a first load level classification result.
3. The method according to claim 2, characterized in that Acquiring at least one operating status data corresponding to the target detection scenario within the target time window includes: In the case where the target detection scenario is an operation load detection scenario, obtaining the operation status data includes at least one of the following: an average number of database connections, a number of insert operations, a number of update operations, a number of delete operations, a number of precise queries, a number of fuzzy queries, and a number of controlled user access transactions; In the case where the target detection scenario is a permission security detection scenario, obtaining the operating status data includes at least one of the following: the number of accessed network addresses, the number of successful logins, the number of unlocked users, the number of failed logins, and the number of locked users.
4. The method according to claim 3, characterized in that The merged data includes the number of insert operations merged; The merging of at least one of the operating status data to obtain merged data includes: The numbers of insert operations respectively obtained from the plurality of database nodes are summed to obtain the merged number of insert operations.
5. The method according to claim 1, wherein The business attribute data includes at least one of the following: maintenance period attribute data, batch processing period attribute data, holiday period attribute data, promotion period attribute data, attribute data of the first period, attribute data of the second period, and attribute data of the third period. The business transaction frequency of the first period is greater than the business transaction frequency of the second period, and the business transaction frequency of the second period is greater than the business transaction frequency of the third period.
6. The method according to claim 5, characterized in that The attribute data of the second period is determined based on the following operations: Statistical processing is performed on the login service sub-attribute data, the initial service sub-attribute data of the day, the normal service sub-attribute data, and the daily settlement service sub-attribute data to obtain attribute data of the second time period.
7. The method according to claim 1, characterized in that The second load level classification result is obtained by performing supervised classification based on a decision tree model, and the training of the decision tree model is determined based on the following operations: Acquire training samples, where the training samples include sample business attribute data and sample operation status data corresponding to each of a plurality of sample time windows; For each of the plurality of sample time windows, performing unsupervised classification on the sample operation status data to obtain a sample first load level classification result corresponding to each of the plurality of sample time windows, wherein the sample first load level classification result represents a load level pseudo-label of the sample service attribute data within the sample time window; An initial decision tree model is trained according to the sample business attribute data and the sample first load level classification results corresponding to each of the plurality of sample time windows to obtain the decision tree model.
8. The method according to claim 1, characterized in that The determining of a detection result according to the first load level classification result and the second load level classification result includes: When the first load level classification result is the same as the second load level classification result, a normal detection result is obtained; In the case that the first load level classification result is different from the second load level classification result, an abnormal detection result is obtained, triggering a system warning.
9. A database cluster anomaly detection device, characterized in that: The device comprises: a first classification module, configured to, in response to an anomaly detection request for a database cluster in a target time window, perform unsupervised classification on operating status data within the target time window to obtain a first load level classification result, wherein the operating status data represents status data of the database cluster during operation; a second classification module, configured to perform supervised classification on the business attribute data of the target time window to obtain a second load level classification result, wherein the business attribute data represents the degree to which the target time window has a periodic impact on the business transaction category and frequency; The detection module is used to determine a detection result according to the first load level classification result and the second load level classification result.
10. An electronic device comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors call the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Database cluster exception handling method and device, equipment and medium
CN114385453A
Abnormality detection method and device, electronic equipment and readable storage medium
CN117610933A
Abnormal event detection method and device, equipment, medium and product
CN119203121A
Fault alarm self-healing system and method
CN119292810A
Flow limiting management method, system and device applied to real-time database and medium
CN119739698A