Database alarm detection method and device, equipment and medium
By configuring alarm rules in the database system and combining them with alarm collection systems and data nodes for detection, the problems of insufficient real-time performance and accuracy in the database alarm system are solved, enabling fast and accurate anomaly detection and improving the stability of database operation.
Patent Information
- Application Number
- CN202511180638.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-18
AI Technical Summary
Existing database alarm systems rely on manual operation, making it difficult to achieve real-time fault detection and rapid location, which leads to the escalation of problems and affects the availability of database services.
By configuring alarm rules in both the alarm collection system and data nodes, and combining the alarm collection system and data nodes for alarm detection, the detection dimensions and scope are increased. By fusing coarse-grained and fine-grained detection results, rapid and accurate anomaly detection can be achieved.
It improves the accuracy and real-time performance of database anomaly detection, prevents problems from escalating, and enhances the stability of database operation.
Smart Images

Figure CN120973584A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of database technology, and in particular to a database alarm detection method, apparatus, device, and medium. Background Technology
[0002] As the carrier of core data, the stability of databases directly impacts daily business operations. To ensure database availability, operational and maintenance methods are crucial. Troubleshooting and repair, as well as performance monitoring and optimization, are routine tasks in database operations and maintenance. These involve addressing various issues that may arise during database system operation, such as component offline, data inconsistency, and performance degradation, and locating problems through log analysis and performance monitoring tools.
[0003] Database alerts facilitate real-time fault detection and rapid response, enabling risk prevention and performance optimization. For database system operators, alerts can help locate problems and greatly reduce the likelihood of problems escalating and affecting the availability of database services. Therefore, databases need to evolve towards automated and refined problem exposure, and optimizing the alerting system is a crucial part of this process.
[0004] Currently, database alerts rely on operations and maintenance personnel. This method makes it difficult to guarantee the real-time detection and location of problems, which may lead to further escalation of the problem and affect the normal operation of the business. Summary of the Invention
[0005] This invention provides a database alarm detection method, apparatus, device, and medium, which can increase detection dimensions, improve the accuracy and real-time performance of anomaly detection, and enhance the stability of database operation.
[0006] According to one aspect of the present invention, an embodiment of the present invention provides a database alarm detection method, the method comprising:
[0007] Retrieve alarm rule information from the database;
[0008] The alarm rule information is sent to the alarm processing object; the alarm processing object includes the alarm acquisition system or data node.
[0009] Based on the alarm rule information, the alarm data reported by the alarm collection system is used to make alarm judgments, and a first detection result is obtained;
[0010] Obtain the second detection result reported by the data node;
[0011] The alarm detection results of the database are determined based on the first detection result and the second detection result.
[0012] According to another aspect of the present invention, embodiments of the present invention also provide a database alarm detection device, the device comprising:
[0013] The alarm rule information acquisition module is used to acquire alarm rule information from the database.
[0014] An alarm rule information sending module is used to send the alarm rule information to the alarm processing object; the alarm processing object includes an alarm acquisition system or a data node;
[0015] The first detection module is used to perform alarm judgment on the collected data reported by the alarm collection system according to the alarm rule information, and obtain the first detection result;
[0016] The second detection module is used to obtain the second detection result reported by the data node;
[0017] An alarm detection module is used to determine the alarm detection results of the database based on the first detection result and the second detection result.
[0018] According to another aspect of the present invention, embodiments of the present invention also provide an electronic device, the electronic device comprising:
[0019] At least one processor; and
[0020] A memory that is communicatively connected to at least one processor; wherein,
[0021] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to execute the database alarm detection method of any embodiment of the present invention.
[0022] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the database alarm detection method of any embodiment of the present invention.
[0023] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the database alarm detection method according to any embodiment of the present invention.
[0024] The technical solution of this invention configures alarm rule information for the alarm collection system and data nodes respectively, and performs alarm detection based on the corresponding alarm rule information by the alarm collection system and data nodes, thereby increasing the detection dimension and scope. This solves the problem of difficulty in discovering and quickly locating problems in the prior art, and can quickly and accurately detect database anomalies and issue real-time alarms to prevent the problem from escalating further, thereby improving the stability of database operation.
[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a flowchart of a database alarm detection method provided according to an embodiment of the present invention;
[0028] Figure 2 This is a flowchart of a database alarm detection method provided according to an embodiment of the present invention;
[0029] Figure 3 This is a scene diagram of a database alarm detection method provided according to an embodiment of the present invention;
[0030] Figure 4 This is a structural diagram of a database alarm detection device provided according to an embodiment of the present invention;
[0031] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0032] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0034] The acquisition, storage, and application of driving trajectory points and other related technologies in the technical solutions of this invention comply with relevant laws and regulations and do not violate public order and good morals.
[0035] Figure 1 This is a flowchart illustrating a database alarm detection method provided in an embodiment of the present invention. This embodiment is applicable to database alarm detection, and the method can be executed by a database alarm detection device, which can be implemented in hardware and / or software.
[0036] The database alarm detection method of this invention is applied to a database system, which includes a database management module (Insight), a metadata module, data nodes, and an alarm collection system. The database management module is an intelligent operation and maintenance system for the distributed database. It configures, displays, and distributes alarm rules, and shows alarm rules and alarm detection results. The database management module is a unified operation and maintenance management service for the distributed database, capable of creating and maintaining database instances, and providing monitoring and alarm functions. The distributed database can be GoldenDB (Database), a financial-grade transactional distributed database. The metadata module (RelationDataBase, RDB) can refer to the database cluster metadata module, which stores distributed database metadata, operation and maintenance alarm data, and backup management of the aforementioned data. The metadata module receives configuration parameters from the database management module and stores them in a table. Data nodes store data and perform database operations. Database agents (DataBase Agent, DBAgent) and operation, maintenance, and monitoring management agents (Operations, Maintenance & Monitoring Manager Agent, OMMAgent) are deployed on the data nodes. Data nodes record anomalies from their own logs in a system table, allowing the database agent to directly read from this table and perform alarm detection according to rules. The database agent monitors and starts / stops the daemons of database instances. The database agent connects to data nodes to automatically detect alarm rules corresponding to the data node logs issued by the database management module, generating and reporting corresponding alarm messages. The operations and maintenance monitoring and management agent provides extended operations and maintenance functions for the upper-layer platform, providing high availability for management nodes. As an information transmission channel between the database agent and the database management module, the operations and maintenance monitoring and management agent facilitates the database agent's acquisition of alarm rule information and displays the alarm messages reported by the database agent on the database management module interface.
[0037] The alarm process mainly involves the database management module, the operation and maintenance monitoring management agent, and the database agent. The database management module can generate an alarm display page, showing detailed alarm information. The operation and maintenance monitoring management agent acts as a conduit for information transmission between the database agent and the database management module. The database agent integrates and filters anomalies matched with alarm rules from data nodes, packages them into key alarm information, and transmits it upwards.
[0038] See Figure 1 The database alarm detection method shown includes:
[0039] S101. Obtain alarm rule information from the database.
[0040] The alarm rule information is used to detect database anomalies. In some embodiments, the alarm rule information may include CPU (Central Processing Unit) utilization exceeding 60%; for example, the alarm rule information may include CPU utilization exceeding 80%. The alarm rule information can be manually configured. In this embodiment of the invention, it can be executed through the database system, specifically through the database management module.
[0041] In some embodiments, obtaining alarm rule information from the database includes: obtaining rule information input by the user through a rules page and generating alarm rules; storing the generated alarm rules in the metadata module of the database; extracting the stored alarm rules from the metadata module and using them as alarm rules for the database.
[0042] The database system provides an interactive page where users can trigger rule configuration controls to navigate to the rules page. Users can configure alarm rules on the rules page. Users can input rule information or select and configure options based on existing rule parameters. The metadata module stores alarm rules. Users can continuously update the alarm rules in the metadata module. When an alarm is triggered, the database system retrieves the latest alarm rule from the metadata module and uses it as the database's alarm rule.
[0043] In some embodiments, the rules page can display the alarm rules configured in the database system. An alarm rule includes multiple configuration items, such as: alarm description (alarm level, alarm description, alarm cause description, remediation suggestions, alarm code, and alarm cause code, etc.), alarm type, alarm detection cycle, recovery method (automatic clearing switch; if automatic clearing is enabled, the clearing cycle also needs to be set) or duration, etc. For configured alarm rules, metadata can also be added. Metadata may include at least one of the following: rule name, alarm code, alarm cause code, alarm level, component type, object identifier, metric type, status, detection cycle, and operating user, etc. Metadata can be directly extracted from configuration items or automatically generated based on preset rules. Alarm rules are grouped according to their metadata; the same rule can belong to multiple groups. The rules page supports various operations on rules, including displaying details, editing, enabling, disabling, deleting, and querying.
[0044] Optionally, alarm rules can be further divided into sub-rules, with different thresholds, alarm codes, alarm reason codes, and alarm levels configured for different sub-rules. A single alarm rule cannot contain two sub-rules with the same level. For example, if CPU alarm is rule 1, sub-rule 1.1: CPU utilization exceeds 60%, alarm level is critical; sub-rule 1.2: CPU utilization exceeds 80%, alarm level is urgent.
[0045] Alarm rules can generally be divided into two categories: one corresponding to the alarm acquisition system and the other corresponding to the data nodes. Based on the type of alarm rule, the object to which the alarm is issued can be determined.
[0046] The current full set of alarm rules is retrieved from the metadata module and displayed on the rules page of the database management module. The database management module edits the alarm rules and stores them in the metadata module.
[0047] In some embodiments, after the database management module edits an alarm rule, it is synchronized to the metadata module for storage. The metadata module stores all alarm rules, including built-in alarm rules and user-defined alarm rules. Built-in alarm rules can refer to general database alarm rules, which are usually configured by the development user. Built-in alarm rules are typically configured directly during development, rather than through the rules page. User-defined alarm rules can refer to alarm rules for additional databases, which are usually configured by the user.
[0048] When the database management module or database agent restarts, it will obtain alarm rules from the metadata module. The database agent only obtains alarm rules related to data nodes, while the database management module obtains all alarm rules, including alarm rules of data nodes and alarm rules of the alarm collection system.
[0049] As can be seen, generating alarm rules by obtaining user-input rule information through the rules page and storing them in the metadata module allows users to customize alarm rules, which can enrich the content of alarm rules, increase the scope of detectable anomalies, and improve the accuracy of anomaly detection.
[0050] S102. Send the alarm rule information to the alarm processing object; the alarm processing object includes the alarm acquisition system or data node.
[0051] The alarm collection system can refer to an application configured within the database management module. It collects multi-dimensional operational information from the database and performs alarm detection based on this information and alarm rules. A data node can refer to the core storage and execution unit in a distributed database. Data nodes store data and perform database operations (CRUD operations). The alarm collection system also collects operational information from other servers and modules within the database system, excluding data nodes. Deployed within the database management module, the alarm collection system is specifically designed to collect operational information from the database system. Data nodes themselves have monitoring capabilities and can collect their own operational information. Therefore, the functionality of the data nodes themselves can be directly reused to perform alarm detection based on alarm rules.
[0052] S103. Based on the alarm rule information, perform alarm judgment on the collected data reported by the alarm collection system to obtain the first detection result.
[0053] The first detection result can refer to the result obtained by the alarm acquisition system from alarm detection in the database. The alarm acquisition system is only used to acquire and report data. The database management module is used to perform alarm detection based on the acquired data.
[0054] S104. Obtain the second detection result reported by the data node.
[0055] The second detection result can refer to the result obtained by the data node in detecting alarms in the database.
[0056] S105. Determine the alarm detection results of the database based on the first detection result and the second detection result.
[0057] The first detection result can be understood as a global, coarse-grained alarm detection result, while the second detection result can be understood as a local, fine-grained alarm detection result. The first and second detection results are then combined to obtain the alarm detection results for the database. Alternatively, the first and second detection results can be directly concatenated and combined to obtain the database's alarm detection results.
[0058] The technical solution of this invention configures alarm rule information for the alarm collection system and data nodes respectively, and performs alarm detection based on the corresponding alarm rule information by the alarm collection system and data nodes, thereby increasing the detection dimension and scope. This solves the problem of difficulty in discovering and quickly locating problems in the prior art, and can quickly and accurately detect database anomalies and issue real-time alarms to prevent the problem from escalating further, thereby improving the stability of database operation.
[0059] Figure 2This is a flowchart illustrating a database alarm detection method provided in an embodiment of the present invention. Based on the above embodiments, this embodiment sends the alarm rule information to the alarm processing object, specifically as follows: A first alarm rule and a second alarm rule are determined according to the metadata of each alarm rule in the alarm rule information; the first alarm rule is configured with collection information, calculation information, alarm threshold, and script information; the second alarm rule is configured with a detection period and alarm content range; the first alarm rule is sent to the alarm collection system; and the second alarm rule is sent to the data node.
[0060] It should be noted that for parts not described in detail in the embodiments of the present invention, please refer to the descriptions in other embodiments.
[0061] See Figure 2 The database alarm detection method shown includes:
[0062] S201. Obtain alarm rule information from the database.
[0063] S202. Based on the metadata of each alarm rule in the alarm rule information, determine the first alarm rule and the second alarm rule; the first alarm rule is configured with collection information, calculation information, alarm threshold and script information; the second alarm rule is configured with detection period and alarm content range; the alarm processing object includes alarm collection system or data node.
[0064] The metadata of the alarm rules describes the alarm rules. The recipients are determined based on the metadata. The first alarm rule is used by the alarm collection system for data acquisition. The second alarm rule is used by data nodes for alarm detection. The first and second alarm rules are different.
[0065] For the first alarm rule, the collected information describes the collected content and method. The calculation information describes the calculation method applied to the collected data. The alarm threshold is used to compare with the calculation result to determine the alarm. The script information is used by the alarm collection system to execute the corresponding collection script to collect the data.
[0066] For the second alarm rule, the detection period can refer to the period during which alarms are detected. The alarm content range can refer to the range of detection results that trigger the existence of an alarm.
[0067] In some embodiments, to facilitate the distinction between the first alarm rule and the second alarm rule, a rule type for the alarm rule can be generated based on the content of the alarm rule's metadata and added to the alarm rule's metadata. For example, the rule type may include collection type and node type. The alarm rule of the collection type is the first alarm rule, and the alarm rule of the node type is the second alarm rule. The rule type of the alarm rule can be directly read from the metadata to determine whether it is the first or second alarm rule.
[0068] S203. Send the first alarm rule to the alarm collection system.
[0069] S204. Send the second alarm rule to the data node.
[0070] S205. Based on the alarm rule information, perform alarm judgment on the collected data reported by the alarm collection system to obtain the first detection result.
[0071] In an optional embodiment, the step of judging the alarm data reported by the alarm collection system based on the alarm rule information to obtain a first detection result includes: performing corresponding logical operations on the collected data reported by the alarm collection system for each first alarm rule in the alarm rule information to obtain a logical operation result; determining the detection result of the first alarm rule based on the logical operation result and the alarm threshold of the first alarm rule; and determining the detection result of each first alarm rule as the first detection result.
[0072] The alarm acquisition system collects data for each first alarm rule and adds relevant information about the first alarm rule to the collected data. It can retrieve the data corresponding to each first alarm rule from the collected data. For each first alarm rule, logical operations are performed on the corresponding data based on the calculation information in the first alarm rule to obtain the logical operation result. The logical operation result is compared with the alarm threshold to obtain the comparison result. Based on the comparison result and the judgment conditions in the first alarm rule, the detection result of the first alarm rule is determined. The detection results of all first alarm rules are aggregated to obtain the first detection result.
[0073] For example, the first alarm rule: the collected data is the CPU utilization rate over 5 minutes. The calculation information could be the average CPU utilization rate over 5 minutes. Based on this average, the logical operation result is 85%. The alarm threshold is 60%. Comparing the logical operation result with the alarm threshold, the average CPU utilization rate of 85% > the alarm threshold of 60%. The judgment condition for this first alarm rule is: if the logical operation result exceeds the alarm threshold, the detection result is determined to require an alarm. Accordingly, the detection result of the first alarm rule corresponding to this collected data is that an alarm is required.
[0074] For example, the collected data includes CPU utilization and memory utilization. The calculated logical operation results are: CPU utilization 91%, memory utilization 93%. The alarm threshold for CPU utilization is 90%, and the alarm threshold for memory utilization is 95%. The judgment condition for the first alarm rule is: if both CPU utilization and memory utilization exceed the alarm thresholds, the detection result of the first alarm rule corresponding to the collected data is determined to require an alarm; otherwise, the detection result of the first alarm rule corresponding to the collected data is determined to be normal. The average CPU utilization of 91% is greater than the CPU utilization alarm threshold of 90%, while the memory utilization of 93% is less than the memory utilization alarm threshold of 95%. Accordingly, the detection result of the first alarm rule corresponding to the collected data is normal.
[0075] For example, the collected data is CPU utilization. The calculated logical operation result is: CPU utilization is 91%. The alarm thresholds for CPU utilization are 90% and 60%. The judgment conditions for the first alarm rule are: if the CPU utilization exceeds 90%, the detection result of the first alarm rule corresponding to the collected data is determined to require an alarm, and the alarm level is urgent; if the CPU utilization exceeds 60% but does not exceed 90%, the detection result of the first alarm rule corresponding to the collected data is determined to require an alarm, and the alarm level is important; if the CPU utilization does not exceed 60%, the detection result of the first alarm rule corresponding to the collected data is determined to be normal. The average CPU utilization of 91% is greater than the CPU utilization alarm threshold of 90%, and the alarm level is urgent. Accordingly, the detection result of the first alarm rule corresponding to the collected data is determined to require an alarm, and the alarm level is urgent.
[0076] As can be seen, by performing logical operations on the corresponding collected data based on the first alarm rule, obtaining the logical operation result, and comparing the logical operation result with the alarm threshold to determine the detection result, and determining the first detection result based on the detection results of each first alarm rule, composite logical judgment conditions can be used to realize correlation analysis of multi-dimensional data, identify hidden anomalies, and multi-condition weighted judgment can also realize hierarchical alarms, improve the ability to identify missed alarms. At the same time, logical judgment can reduce the resources required for alarm detection, especially in high-frequency data detection scenarios, which can reduce redundant resource consumption and improve the system operation stability in high-frequency data detection scenarios.
[0077] S206. Obtain the second detection result reported by the data node.
[0078] S207. Determine the alarm detection results of the database based on the first detection result and the second detection result.
[0079] This invention, through configuring metadata for alarm rules and configuring different metadata for different alarm processing objects, can enrich alarm processing scenarios. By combining node alarms and collection alarms, it can handle more complex alarm requirements and has multi-dimensional alarm functions.
[0080] In an optional embodiment, the database alarm detection method further includes: sending script information associated with each first alarm rule to the corresponding detection object through the alarm collection system according to the received alarm rule information, so that the detection object executes the associated script information, obtains collected data, and feeds it back; acquiring the collected data fed back by each detection object through the alarm collection system; and integrating and reporting the collected data fed back by each detection object through the alarm collection system according to the alarm rule information.
[0081] The alarm acquisition system includes scripts. The monitored object executes the scripts to obtain collected data. The monitored object then feeds the collected data back to the alarm acquisition system. The alarm acquisition system integrates and reports the collected data fed back by the monitored object. The script information may include the identifier of the executed script. The script information may also include the script itself.
[0082] Deploying an alarm collection system within a database system allows for the use of a rich set of external scripts to periodically monitor various database system metrics in real time, including common performance indicators and component status metrics. The database management module acts as the overall task initiator, sending scripts to each component for execution to collect data. Each component can be configured with timers to periodically and automatically run its collection scripts. The collected results are uploaded to the alarm collection system and then passed through to the database management module. The database management module calculates the data according to the alarm rules specifying the calculation methods for each metric, compares the final result with alarm thresholds, and generates an alarm if the conditions are met.
[0083] The detection targets may include proxy database management modules, proxy data nodes, and compute nodes. The detection scope of the first alarm rule can be to detect performance, faults, and basic component operation issues. In some embodiments, a script is a script for collecting its own memory usage and CPU usage. The detection target executes this script, obtains the memory usage and CPU usage of the detection target, and sends it to the alarm collection system.
[0084] The alarm process of the alarm acquisition system reporting the collected data:
[0085] ① The database management module configures alarm rules, and the alarm collection system performs periodic automatic detection according to the configured first alarm rule by calling system commands, running scripts, and component self-monitoring, and returns the collected data to the database management module.
[0086] ② The database management module processes the collected data returned by the alarm acquisition system according to the operators of the first alarm rule, and compares the logical operation result with the alarm threshold to see if it meets the judgment condition. The judgment condition can be to first check whether the inequality is satisfied, and if so, to check whether the duration meets the duration threshold. For collected data that meets the duration threshold, the alarm details are displayed on the rules page of the database management module.
[0087] ③ For any existing alarm, the database management module can enable automatic cancellation. If the same fault does not occur again within a certain period after the alarm and the automatic cancellation time is met, the alarm will be restored.
[0088] As can be seen, the alarm collection system sends the script information associated with the first alarm rule to the detection object, so that the detection object executes the script and obtains the collected data feedback to the alarm collection system. At the same time, the alarm collection system integrates the collected data of multiple detection objects and reports it to the database management module in the database system. The scope and content of the collection object can be adjusted by configuring the script, which can improve the flexibility of the alarm scope, reduce the complexity of alarm scope adjustment, and increase the diversity of detection content.
[0089] In an optional embodiment, the database alarm detection method further includes: parsing the received alarm rule information through the data node to obtain the detection period of each second alarm rule; merging the detection periods of each second alarm rule through the data node to obtain at least one target period; and performing alarm detection on log data through the data node according to each target period and the second alarm rule associated with each target period to obtain a second detection result and report it.
[0090] In this system, each second alarm rule corresponds to one detection period. If several second alarm rules have the same detection period, these periods can be merged. For example, if second alarm detection rule a has a detection period of 1 second, second alarm detection rule b has a detection period of 2 seconds, and second alarm detection rule c has a detection period of 1 second, then the detection periods of a and c can be merged. The merged detection period is then determined as the target period. Detection periods that cannot be merged are also determined as the target period.
[0091] During runtime, data nodes record operational information and any exceptions encountered during operation in the log. The second alarm rule is used to detect alarms based on the data node's log data.
[0092] Data nodes parse alarm rule information to obtain at least one second alarm rule and extract the detection period from the second alarm rule. The detection periods are merged to obtain at least one target period. Log data is periodically acquired according to the target period. Within the corresponding target period, the second alarm rule corresponding to that target period is selected to perform alarm detection on the log data generated in that target period or the full amount of log data available in that target period.
[0093] As can be seen, by extracting the detection period from the second alarm rule by the data node and merging the detection periods, at least one target period is obtained. Alarm detection is then performed on the log data available within the target period according to the second alarm rule corresponding to the target period. This achieves periodic anomaly detection of the data node, optimizes the period fusion, reduces redundant periods, improves resource utilization, and uses different period detection for different second alarm rules to address targeted alarm detection for different scenarios. This adapts to diverse scenario alarm requirements, increases the scalability of alarm scenarios, and reduces the false negative rate compared to single-point detection.
[0094] In an optional embodiment, the step of performing alarm detection on log data through the data node according to each target period and the second alarm rule associated with each target period includes: obtaining log data through the data node according to the target period; and detecting whether the log data belongs to the alarm content range of each associated second alarm rule according to at least one second alarm rule associated with the target period through the data node.
[0095] The process involves retrieving log data at each target cumulative period and then resetting the timer. This log data can be incremental log data for the target period or the full amount of log data available for the target period.
[0096] When the target period is not the merged period, retrieve one second alarm rule associated with the target period. When the target period is the merged period, retrieve multiple second alarm rules associated with the target period.
[0097] When log data falls within the alarm content range of the second alarm rule associated with the target period, the second alarm rule is determined to have been hit; when log data does not fall within the alarm content range of the second alarm rule associated with the target period, the second alarm rule is determined to have been missed. All second alarm rules associated with the target period need to be determined to have been hit. If a second alarm rule is hit, the detection result of that second alarm rule is determined to require an alarm. If a second alarm rule is missed, the detection result of that second alarm rule is determined to be normal.
[0098] Specifically, in the data node, the database agent parses the metadata of the second alarm rule issued by the database management module, sets a timer for periodic automatic detection based on the acquired metadata, and in order to improve the detection efficiency, multiple second alarm rules using the same detection period are merged by the database agent, that is, multiple second alarm rules are detected at once.
[0099] Data nodes save anomalies from their logs as a record in the system table, including the thread identifier, error code, problem description, occurrence time, and error severity. The database agent compares the alert content range of the second alert rule with the record in the system table. If the error code in the system table falls within the alert content range required by the second alert rule, the database agent reports an alert. For example, if the log data contains the error code 123, and the second alert rule's alert content range includes error codes 123, then the log data falls within the alert content range, the second alert rule is triggered, and the detection result of the second alert rule is determined to require an alert.
[0100] The second alarm rule configured on the rules page of the database management module is an alarm rule related to data node logs. When the database management module configures the second alarm rule, it sends it to the database agent, updates the memory stored in the second alarm rule, and updates the database agent's periodic alarm detection operations according to the updated content. For example, if the detection period of the second alarm rule is modified, the database agent needs to modify the detection period of the original scheduled task.
[0101] After the second alarm rule is edited on the rules page of the database management module, the database management module sends it to the operation and maintenance monitoring management agent, which then forwards the message to the database agent.
[0102] If the database agent restarts, it sends a message to the operations and maintenance monitoring and management agent to proactively request all relevant alarm rules from the data node logs. The operations and maintenance monitoring and management agent then forwards this request to the database management module. The database management module then issues all alarm rules, following the same process as when the database management module edits and issues rules, thus achieving a flow from the database management module to the operations and maintenance monitoring and management agent, and finally to the database agent.
[0103] After receiving the second alarm rule, the database agent sets up a corresponding timer task, which triggers subsequent periodic detection tasks.
[0104] The database agent returns a response message to the operation and maintenance monitoring and management agent when the database agent is set up successfully or unsuccessfully. The operation and maintenance monitoring and management agent then forwards the message to the database management module.
[0105] If the second alarm rule fails to be issued after modification in the database management module, a pop-up window will appear on the page to inform the user that the editing failed.
[0106] The database agent groups alarms with the same detection period into the same timer to improve detection performance. Data nodes write anomalies from their logs to the alarm system table, including thread identifier, anomaly occurrence time, anomaly severity level, anomaly description, and anomaly error code. The database agent compares the system table with the second alarm rule. If data in the system table falls within the scope of the alarm content required by the second alarm rule, an alarm is generated and reported to the operations and maintenance monitoring management agent. The operations and maintenance monitoring management agent reports the alarm to the database management module, which displays it on the rules page. If the database management module does not find any new instances of the same problem in the next period after the alarm is generated, and the elimination period is met, the alarm is restored.
[0107] As can be seen, by having data nodes periodically acquire log data according to the target period and detect whether the log data falls within the alarm content range of the corresponding second alarm rule to determine the second detection result, the complexity of data node alarm detection can be simplified, while reusing the self-monitoring results of data nodes and reducing the resource consumption of data node alarm detection.
[0108] In a scenario, such as Figure 3 As shown, the database system includes a database management module (Insight), a metadata module (RDB), data nodes (DN), and an alarm collection system. Database agent (DBAgent) and operation and maintenance monitoring management agent (OMMAgent) are deployed on the data nodes.
[0109] 1. When the Insight rules page starts, it retrieves all current alarm rules from the RDB table and displays them on the rules page.
[0110] 2. After editing the alarm rules on the Insight rules page, store them in the RDB table to ensure that the metadata in the RDB is consistent with the display on the rules page.
[0111] DBAgent component rule alerting function:
[0112] 3. After editing the second alarm rule related to the DN log on the Insight rules page, it is sent to OMMAgent, which then forwards the message to DBAgent. If DBAgent restarts, it sends a message to OMMAgent to actively request the second alarm rules related to all DN logs, which are then forwarded to Insight by OMMAgent.
[0113] 4. After receiving the second alarm rule related to the DN log, DBAgent sets a corresponding timer task, which triggers subsequent periodic detection tasks. DBAgent returns a response message to OMMAgent for successful or unsuccessful processing of the second alarm rule, which is then passed through to Insight. If Insight modifies and then fails to issue the first alarm rule, a pop-up window on the rule page will indicate to the user that editing failed.
[0114] 5. DBAgent places alarms with the same detection cycle into the same timer to improve detection performance.
[0115] 6. DN writes the exceptions in its own logs to the alarm system table, including thread identifier, time of occurrence of exception, severity level of exception, description of exception and error code.
[0116] 7. DBAgent compares the alarm system table with the second alarm rule. When the system table contains data within the alarm content range that needs to be alarmed according to the second alarm rule, an alarm is generated and reported to OMMAgent.
[0117] 8. OMMAgent will report the alerts to Insight, which will then be displayed on the Insight rules page.
[0118] If no new cases of the same problem are found in the next cycle after an alert is generated by Insight, and the clearance cycle is met, the alert will be restored.
[0119] The alarm process of the alarm acquisition system:
[0120] 9. The first alarm rule configured successfully in Insight, excluding DN log anomalies, will be automatically detected periodically by the alarm collection system through methods such as calling system commands, running scripts, and component self-monitoring, and the detected values will be returned to Insight.
[0121] 10. Insight processes the collected data returned by the alarm acquisition system according to the operators of the first alarm rule, and compares the logical operation result with the alarm threshold to see if it meets the judgment condition. The judgment condition can be to first check whether the inequality is satisfied, and if so, to check whether the duration meets the duration threshold. For collected data that meets the duration threshold, the alarm details are displayed on Insight's rule page.
[0122] When an alarm in Insight is set to automatically clear, if the same fault does not occur again within a certain period after the alarm and the automatic clearing time is met, the alarm will be restored.
[0123] This invention implements database rule-based alerting. By periodically and automatically detecting database faults and performance according to rules, it can automatically detect database faults and performance issues, and report and recover from them. It supports both built-in and custom alert rules. Basic database system monitoring items are defined as built-in alert rules, while users can set custom rules as needed. It supports multiple operators, meaning rules can perform calculations on multiple indicators and generate alerts based on the results. Multiple components are monitored simultaneously, including system commands, monitoring scripts, and the component's own monitoring capabilities, performing periodic checks across multiple dimensions. The technology for alarm reporting and recovery reduces the burden of manually troubleshooting database system faults and performance degradation, improving database stability and performance. The rule-based alerting detection has a wide range, providing comprehensive database fault detection, and supports both built-in and custom alert rules. The rules not only support single-indicator detection and alarms, but also multi-indicator joint calculation alarms, and simultaneous monitoring of multiple components. A data collection system is deployed, utilizing system commands, monitoring collection scripts, and the component's own monitoring capabilities for multi-dimensional periodic detection. This covers various types of faults and performance issues across all components of the database system, classifying issues by alarm level to effectively assist operations and maintenance personnel in risk prevention and fault diagnosis. It also boasts high flexibility, supporting user-defined rule alarms and multiple operators. Users can add or delete alarm rules as needed, and edit alarm rules to create database rule alarm functions tailored to business characteristics. The system provides a method and device for periodic automatic detection and alarming based on built-in rules and user-added rules, ensuring database stability, timely problem exposure, shortening fault location time, and reducing the workload of operations and maintenance personnel.
[0124] Figure 4 This is a schematic diagram of a database alarm detection device provided in an embodiment of the present invention. The present invention is applicable to database alarm detection, and the device can execute a database alarm detection method. The device can be implemented in hardware and / or software.
[0125] See Figure 4 The database alarm detection device shown includes:
[0126] The alarm rule information acquisition module 401 is used to acquire alarm rule information from the database;
[0127] The alarm rule information sending module 402 is used to send the alarm rule information to the alarm processing object; the alarm processing object includes an alarm acquisition system or a data node;
[0128] The first detection module 403 is used to perform alarm judgment on the collected data reported by the alarm collection system according to the alarm rule information, and obtain the first detection result;
[0129] The second detection module 404 is used to obtain the second detection result reported by the data node;
[0130] The alarm detection module 405 is used to determine the alarm detection result of the database based on the first detection result and the second detection result.
[0131] The technical solution of this invention configures alarm rule information for the alarm collection system and data nodes respectively, and performs alarm detection based on the corresponding alarm rule information by the alarm collection system and data nodes, thereby increasing the detection dimension and scope. This solves the problem of difficulty in discovering and quickly locating problems in the prior art, and can quickly and accurately detect database anomalies and issue real-time alarms to prevent the problem from escalating further, thereby improving the stability of database operation.
[0132] Optionally, the alarm rule information sending module 402 is specifically used for:
[0133] Based on the metadata of each alarm rule in the alarm rule information, a first alarm rule and a second alarm rule are determined; the first alarm rule is configured with collection information, calculation information, alarm threshold and script information; the second alarm rule is configured with detection period and alarm content range.
[0134] The first alarm rule is sent to the alarm collection system;
[0135] The second alarm rule is sent to the data node.
[0136] Optionally, the first detection module 403 is specifically used for:
[0137] For each first alarm rule in the alarm rule information, perform corresponding logical operations on the collected data reported by the alarm collection system to obtain the logical operation results;
[0138] Based on the logical operation result and the alarm threshold of the first alarm rule, the detection result of the first alarm rule is determined;
[0139] The detection results of each of the first alarm rules are determined as the first detection results.
[0140] Optionally, the alarm rule information sending module 402 is specifically used for:
[0141] Obtain the rule information entered by the user through the rules page and generate alarm rules;
[0142] The generated alarm rules are stored in the metadata module of the database;
[0143] The stored alarm rules are extracted from the metadata module and used as alarm rules for the database.
[0144] Optionally, the database alarm detection device also includes: an alarm acquisition system used for:
[0145] Based on the received alarm rule information, the script information associated with each first alarm rule is sent to the corresponding detection object, so that the detection object executes the associated script information, obtains the collected data, and provides feedback.
[0146] Obtain the collected data fed back by each of the detection objects;
[0147] Based on the alarm rule information, the collected data fed back by each of the detection objects are integrated and reported.
[0148] Optionally, the database alarm detection device also includes: data nodes used for:
[0149] Parse the received alarm rule information to obtain the detection period for each second alarm rule;
[0150] The detection periods of each of the second alarm rules are merged to obtain at least one target period;
[0151] According to the target period and the second alarm rule associated with each target period, alarm detection is performed on the log data, the second detection result is obtained and reported.
[0152] Optionally, data nodes are used for:
[0153] Obtain log data according to the target period;
[0154] Based on at least one second alarm rule associated with the target period, detect whether the log data falls within the alarm content range of each associated second alarm rule.
[0155] The database alarm detection device provided in this embodiment of the invention can execute the database alarm detection method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of executing the database alarm detection method.
[0156] Figure 5 A schematic diagram of the structure of an electronic device 500 that can be used to implement an embodiment of the present invention is shown.
[0157] like Figure 5As shown, the electronic device 500 includes at least one processor 501 and a memory, such as a read-only memory (ROM) 502 or a random access memory (RAM) 503, communicatively connected to the at least one processor 501. The memory stores computer programs executable by the at least one processor. The processor 501 can perform various appropriate actions and processes based on the computer program stored in the ROM 502 or loaded into the RAM 503 from storage unit 508. The RAM 503 can also store various programs and data required for the operation of the electronic device 500. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0158] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows electronic device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0159] Processor 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 501 performs the various methods and processes described above, such as database alarm detection methods.
[0160] In some embodiments, the database alarm detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by processor 501, one or more steps of the database alarm detection method described above may be performed. Alternatively, in other embodiments, processor 501 may be configured to perform the database alarm detection method by any other suitable means (e.g., by means of firmware).
[0161] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0162] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0163] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0164] To provide interaction with a user, the systems and techniques described herein can be implemented on an operational detection device. This electronic device includes: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0165] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0166] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability.
[0167] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0168] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A database alarm detection method, characterized in that, The method includes: Retrieve alarm rule information from the database; The alarm rule information is sent to the alarm processing object; the alarm processing object includes the alarm acquisition system or data node. Based on the alarm rule information, the alarm data reported by the alarm collection system is used to make alarm judgments, and a first detection result is obtained; Obtain the second detection result reported by the data node; The alarm detection results of the database are determined based on the first detection result and the second detection result.
2. The method according to claim 1, characterized in that, Sending the alarm rule information to the alarm processing object includes: Based on the metadata of each alarm rule in the alarm rule information, a first alarm rule and a second alarm rule are determined; the first alarm rule is configured with collection information, calculation information, alarm threshold and script information; the second alarm rule is configured with detection period and alarm content range. The first alarm rule is sent to the alarm collection system; The second alarm rule is sent to the data node.
3. The method according to claim 2, characterized in that, The step of performing alarm judgment on the collected data reported by the alarm collection system according to the alarm rule information to obtain a first detection result includes: For each first alarm rule in the alarm rule information, perform corresponding logical operations on the collected data reported by the alarm collection system to obtain the logical operation results; Based on the logical operation result and the alarm threshold of the first alarm rule, the detection result of the first alarm rule is determined; The detection results of each of the first alarm rules are determined as the first detection results.
4. The method according to claim 1, characterized in that, The acquisition of alarm rule information from the database includes: Obtain the rule information entered by the user through the rules page and generate alarm rules; The generated alarm rules are stored in the metadata module of the database; The stored alarm rules are extracted from the metadata module and used as alarm rules for the database.
5. The method according to claim 1, characterized in that, Also includes: The alarm collection system sends the script information associated with each first alarm rule to the corresponding detection object based on the received alarm rule information, so that the detection object executes the associated script information, obtains the collected data, and provides feedback. The alarm collection system acquires the collected data fed back by each of the detected objects. The alarm collection system integrates and reports the collected data from each of the detected objects according to the alarm rules.
6. The method according to claim 1, characterized in that, Also includes: The received alarm rule information is parsed through the data node to obtain the detection period of each second alarm rule; The detection periods of each of the second alarm rules are merged through the data nodes to obtain at least one target period; Through the data nodes, alarm detection is performed on the log data according to each target period and the second alarm rule associated with each target period, and the second detection result is obtained and reported.
7. The method according to claim 6, characterized in that, The step of performing alarm detection on log data through the data nodes, according to each target period and the second alarm rule associated with each target period, includes: Log data is acquired through the data nodes according to the target period. Using the data node, based on at least one second alarm rule associated with the target period, it is detected whether the log data belongs to the alarm content range of each associated second alarm rule.
8. A database alarm detection device, characterized in that, The device includes: The alarm rule information acquisition module is used to acquire alarm rule information from the database. An alarm rule information sending module is used to send the alarm rule information to the alarm processing object; the alarm processing object includes an alarm acquisition system or a data node; The first detection module is used to perform alarm judgment on the collected data reported by the alarm collection system according to the alarm rule information, and obtain the first detection result; The second detection module is used to obtain the second detection result reported by the data node; An alarm detection module is used to determine the alarm detection results of the database based on the first detection result and the second detection result.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the database alarm detection method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that are used to cause a processor to execute the database alarm detection method according to any one of claims 1-7.