Method and system for realizing unified monitoring alarm fault self-healing

By acquiring and configuring monitoring metadata, combining task scheduling to achieve unified monitoring, alarms and fault self-healing, the problem of inability to uniformly manage multi-computer room monitoring and alarms in the existing technology is solved, and automatic self-healing and intelligent operation and maintenance of system failures are realized.

CN120492258APending Publication Date: 2025-08-15BAOFOO COM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510476403.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing technology cannot realize unified monitoring, alarm and fault self-healing functions. Especially after the deployment of dual-living in the same city and multiple-living in different places, it is impossible to realize the aggregated cross-contrast monitoring alarm of multiple computer rooms, resulting in many false alarms and the inability to intervene in time for system failure.

Method used

Through external acquisition and storage, monitoring metadata configuration rules are set, including configurations of tenants, monitoring indicators, data sources, personnel management, alarm rules and fault self-healing rules, and combined with task scheduling to achieve unified monitoring, alarm and fault self-healing.

Benefits of technology

It minimizes the impact in the event of system abnormalities or failures, reduces the fault handling time, and realizes intelligent operation and maintenance without manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492258A_ABST
    Figure CN120492258A_ABST
Patent Text Reader

Abstract

The invention provides a method and a system for realizing unified monitoring and alarm fault self-healing, and solves the problem that monitoring, alarm and fault self-healing functions cannot be realized in a unified manner in the prior art. By collecting the monitoring metadata and performing task scheduling according to the preset monitoring metadata configuration rule, unified monitoring, alarm and fault self-recovery are realized, the influence when the system is abnormal or has a fault is controlled to be minimum or no influence, and the fault processing time is greatly shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of Internet technology, and in particular to a method and system for realizing unified monitoring alarm fault self-healing. Background Art

[0002] Currently, most open-source or commercial monitoring and alerting products in the industry collect indicator data separately from dispersed data sources, making it difficult to aggregate and generate alerts from these data, and their functionality is relatively limited. For example, implementing log, Prometheus, and big data alerts may require building three platforms to meet the relevant monitoring and alerting needs, making centralized management impossible. Especially when deploying systems with active-active systems in the same city or multiple active-active systems in multiple data centers across different locations, only single-data center monitoring and alerting are supported, with no ability to aggregate and cross-compare monitoring and alerts across multiple data centers. Furthermore, there is no way to determine whether the system has experienced related faults or anomalies through cross-comparison and initiate self-healing based on relevant policies, thereby achieving the goal of intelligent operations and maintenance.

[0003] At present, the alarm platform mainly includes three major alarm platforms: big data monitoring platform (based on database or real-time data warehouse) composed of open source and commercial products, log alarm platform (based on log keyword tracking) and Nightingale monitoring platform (based on Prometheus tracking and Nightingale platform configuration rules). The alarms of various business systems are mainly realized through the above three alarm platforms, but there are the following defects: the platform is relatively scattered and cannot be centrally managed, and it is impossible to perform aggregate calculation alarms. A single alarm generates a large number of false alarms. Too many alarms can easily lead to numbness and miss important alarms. Failures cannot be intervened in time, and most alarms still require human intervention. Manual intervention is required to determine whether a system failure has occurred. This problem is particularly prominent after implementing dual-active deployment in the same city and the same computer room. After deployment in separate computer rooms, the monitoring and alarm indicator data collection for the same project is separate, and the data sources are also dispersed, making it difficult to achieve aggregated alarms after data aggregation. At the same time, log alarms, Sky Eye alarms, and Nightingale alarms can only achieve simple alarms and cannot achieve data indicator calculation alarms. Multiple monitoring indicators are aggregated and cross-compared, and based on relevant policies, it is determined whether the relevant system is abnormal or faulty, thereby implementing related functions such as system fault or abnormal self-healing based on relevant events and policies. With the development of the business, in order to improve the user service experience, a unified monitoring alarm and fault self-healing platform is urgently needed to achieve unified monitoring, unified alarms, unified automatic fault self-healing, and unified centralized management, so that a large number of system anomalies or faults can be self-healed without manual intervention. Summary of the Invention

[0004] The present invention provides a method and system for realizing unified monitoring, alarming and fault self-healing, so as to solve the problem in the prior art that the monitoring, alarming and fault self-healing functions cannot be realized in a unified manner.

[0005] In a first aspect, the present invention provides a method for realizing unified monitoring alarm fault self-healing, which specifically comprises the following steps:

[0006] Step S1: externally obtain monitoring metadata and store it;

[0007] Step S2: Setting monitoring metadata configuration rules;

[0008] Step S3: Perform task scheduling according to the monitoring metadata and the monitoring metadata configuration rules to complete monitoring, alarming, and fault self-healing.

[0009] Preferably, in step S1, the external access method includes one or more of http, dubbo, socket, mq and sdk.

[0010] Preferably, in step S1, the storage location of the monitoring metadata includes one or more of mongodb (document database), dragonflydb (memory database) and OceanBase (distributed database).

[0011] Preferably, in step S2, the configuration rules specifically include tenant (in this application, the business party of the method or system used to implement unified monitoring alarm fault self-healing is referred to as a tenant) access configuration, monitoring indicator configuration, data source configuration, personnel management configuration, group management configuration, alarm rule configuration and fault self-healing rule configuration.

[0012] Preferably, the tenant access configuration is used for external identity authentication, including tenant code, description, key, enterprise WeChat push key, user name, password and other field configurations.

[0013] Preferably, the monitoring indicator configuration is used to set the indicators to be monitored, so as to facilitate reading the data corresponding to the indicators to be monitored from the metadata, including the configuration of fields such as indicator code, title, keyword, level, indicator description, and monitoring data source type.

[0014] More preferably, the monitoring data source type includes one or more of ES (Elastic Search), Prometheus, SR (SkyWalking Receiver), and MongoDB.

[0015] Preferably, the data source configuration is used to monitor the configuration of the metadata storage data source, including the configuration of fields such as data source code, data source type, data source name, user name, password, service IP address, port, code, connection timeout, read timeout, etc.

[0016] Preferably, the personnel management configuration is used to set the recipients of the alarm information, including the configuration of fields such as personnel code, name, telephone number, and email address.

[0017] Preferably, the group management configuration is used to bind mutually related persons.

[0018] Preferably, the alarm rule configuration is used to set the monitoring scheduling rules, data sources and alarm groups corresponding to each monitoring indicator, including the configuration of fields such as indicator type, indicator code, execution frequency, execution frequency unit, data source, alarm method, application name, namespace (i.e. resource pool name), alarm trigger value, trigger value unit, relational operator, repeated alarm limit frequency, index name, alarm wording, title, data monitoring range, data monitoring range unit, environment (i.e. computer room identification), etc.

[0019] More preferably, the indicator types include global and application levels.

[0020] The execution frequency units include seconds, minutes and hours.

[0021] The trigger value units include percentage and quantity.

[0022] The relational operators include greater than, less than, equal to, greater than or equal to, and less than or equal to.

[0023] Preferably, the fault self-healing rule configuration is used to configure the logical rules for fault self-healing, including the configuration of automatic recovery rule serial number, indicator type, recovery type, title, execution frequency, automatic recovery limit frequency, execution frequency unit, trigger logic, notification method, recovery notification template, level, application name, namespace, environment identifier, automatic recovery prerequisites, multiple (in this application, the "multiple" means at least two) prerequisite logical judgment rules, data monitoring range, data monitoring range unit, alarm indicator code, trigger value, trigger condition, indicator title, application name, namespace, environment and other fields.

[0024] More preferably, the recovery types include F5, configuration, database, data center cut, node removal, node restart, and privilege reduction.

[0025] The trigger logic includes logical AND (&&) and logical OR (||).

[0026] The prerequisite logic judgment rules include logical AND and logical OR.

[0027] The data monitoring range units include seconds, minutes and hours.

[0028] Preferably, in step S3, the task scheduling includes monitoring scheduling and automatic recovery scheduling.

[0029] More preferably, the monitoring scheduling includes setting a monitoring execution frequency, an alarm message push limit frequency, and a monitoring data query time range.

[0030] More preferably, the automatic recovery scheduling includes setting an automatic recovery execution frequency, an automatic recovery limit frequency, and an alarm data query time range.

[0031] Preferably, in step S3, task scheduling is performed according to the monitoring metadata and the monitoring metadata configuration rules to complete monitoring and alarming, specifically including the following steps:

[0032] Step S301a: regularly read the monitoring metadata, and query metadata related to monitoring and alarming in combination with the monitoring indicator configuration and the alarm rule configuration;

[0033] Step S301b: Parse the metadata related to monitoring and alarming obtained from the query to complete metadata monitoring;

[0034] Step S301c: When the monitored metadata meets the conditions of the alarm rule, an alarm message is generated, relevant personnel are notified, and an alarm record is generated and saved.

[0035] Preferably, in step S301c, the alarm record facilitates subsequent problem backtracking and can also be used to trigger fault self-healing.

[0036] Preferably, in step S3, task scheduling is performed according to the monitoring metadata and the monitoring metadata configuration rules to complete fault self-healing, which specifically includes the following steps:

[0037] Step S302a: Read the alarm record regularly and determine whether fault self-healing is required based on the fault self-healing rule configuration;

[0038] Step S302b: When the fault is self-healing, the fault line is switched to the second line (the target normal line) and the precondition of the second line is determined to be consistent with the line switching condition.

[0039] Step S302c: When the precondition of the second line (indicating that the switching of the second line requires the second line to be detected according to the rules configured in the metadata, and the second line will be switched only if the result is normal) meets the line switching condition, a judgment is made as to whether the current switching action is within the switching current limiting period. If it is not within the switching current limiting period, the line switching action is performed; otherwise, the line switching action is performed after the switching current limiting period ends.

[0040] Preferably, in step S302a, when one or more indicators among the alarm indicators in the alarm record reaches an alarm threshold, fault self-healing is required.

[0041] In a second aspect, the present invention further provides a system for realizing unified monitoring, alarming, and fault self-healing, which specifically includes the following modules:

[0042] Monitoring metadata acquisition module, used to externally acquire monitoring metadata and store it;

[0043] Configuration rule setting module, used to set monitoring metadata configuration rules;

[0044] The monitoring alarm and self-healing processing module is used to schedule tasks according to the monitoring metadata and the monitoring metadata configuration rules, and complete monitoring, alarming and fault self-healing.

[0045] Preferably, in the monitoring metadata acquisition module, the external access method includes one or more of http, dubbo, socket, mq and sdk.

[0046] Preferably, in the monitoring metadata acquisition module, the storage location of the monitoring metadata includes one or more of mongodb, dragonflydb and OceanBase.

[0047] Preferably, in the configuration rule setting module, the configuration rules specifically include tenant access configuration, monitoring indicator configuration, data source configuration, personnel management configuration, group management configuration, alarm rule configuration and fault self-healing rule configuration.

[0048] Preferably, the tenant access configuration is used for external identity authentication, including tenant code, description, key, enterprise WeChat push key, user name, password and other field configurations.

[0049] Preferably, the monitoring indicator configuration is used to set the indicators to be monitored, so as to facilitate reading the data corresponding to the indicators to be monitored from the metadata, including the configuration of fields such as indicator code, title, keyword, level, indicator description, and monitoring data source type.

[0050] More preferably, the monitoring data source type includes one or more of ES, Prometheus, SR, and MongoDB.

[0051] Preferably, the data source configuration is used to monitor the configuration of the metadata storage data source, including the configuration of fields such as data source code, data source type, data source name, user name, password, service IP address, port, code, connection timeout, read timeout, etc.

[0052] Preferably, the personnel management configuration is used to set the recipients of the alarm information, including the configuration of fields such as personnel code, name, telephone number, and email address.

[0053] Preferably, the group management configuration is used to bind mutually related persons.

[0054] Preferably, the alarm rule configuration is used to set the monitoring scheduling rules, data sources and alarm groups corresponding to each monitoring indicator, including the configuration of indicator type, indicator code, execution frequency, execution frequency unit, data source, alarm method, application name, namespace, alarm trigger value, trigger value unit, relational operator, repeated alarm limit frequency, index name, alarm wording, title, data monitoring range, data monitoring range unit, environment and other fields.

[0055] More preferably, the indicator types include global and application levels.

[0056] The execution frequency units include seconds, minutes and hours.

[0057] The trigger value units include percentage and quantity.

[0058] The relational operators include greater than, less than, equal to, greater than or equal to, and less than or equal to.

[0059] Preferably, the fault self-healing rule configuration is used to configure the logical rules for fault self-healing, including the configuration of automatic recovery rule serial number, indicator type, recovery type, title, execution frequency, automatic recovery limit frequency, execution frequency unit, trigger logic, notification method, recovery notification template, level, application name, namespace, environment identifier, automatic recovery prerequisites, multiple prerequisite logical judgment rules, data monitoring range, data monitoring range unit, alarm indicator code, trigger value, trigger condition, indicator title, application name, namespace, environment and other fields.

[0060] More preferably, the recovery types include F5, configuration, database, data center cut, node removal, node restart, and privilege reduction.

[0061] The trigger logic includes logical AND (&&) and logical OR (||).

[0062] The prerequisite logic judgment rules include logical AND and logical OR.

[0063] The data monitoring range units include seconds, minutes and hours.

[0064] Preferably, in the monitoring alarm and self-recovery processing module, the task scheduling includes monitoring scheduling and automatic recovery scheduling.

[0065] More preferably, the monitoring scheduling includes setting a monitoring execution frequency, an alarm message push limit frequency, and a monitoring data query time range.

[0066] More preferably, the automatic recovery scheduling includes setting an automatic recovery execution frequency, an automatic recovery limit frequency, and an alarm data query time range.

[0067] Preferably, the monitoring alarm and self-healing processing module specifically includes a monitoring alarm submodule and a self-healing submodule.

[0068] Preferably, the monitoring and alarm submodule is used to perform task scheduling according to the monitoring metadata and the monitoring metadata configuration rules, complete monitoring and alarming, and specifically includes the following units:

[0069] The first unit is used to periodically read the monitoring metadata, and query metadata related to monitoring and alarming in combination with the monitoring indicator configuration and the alarm rule configuration;

[0070] The second unit is used to parse the metadata related to monitoring and alarm obtained from the query and complete metadata monitoring;

[0071] The third unit is used to generate alarm information, notify relevant personnel, and generate and save alarm records when the monitored metadata meets the conditions of the alarm rules.

[0072] Preferably, in the third unit, the alarm record facilitates subsequent problem tracing and can also be used to trigger fault self-healing.

[0073] Preferably, the self-healing submodule performs task scheduling according to the monitoring metadata and the monitoring metadata configuration rules to complete fault self-healing, and specifically includes the following units:

[0074] The fourth unit is used to read alarm records at regular intervals and determine whether fault self-healing is required based on the fault self-healing rule configuration.

[0075] The fifth unit is configured to switch the faulty line when performing fault self-healing, switch the first line (current faulty line) to the second line (target normal line), and determine whether the precondition of the second line meets the line switching condition;

[0076] The sixth unit is used to judge whether the current switching action is within the switching current limiting period when the precondition of the second line meets the line switching condition. If it is not within the switching current limiting period, the line switching action is performed; otherwise, the line switching action is performed after the switching current limiting period ends.

[0077] Preferably, in the fourth unit, when one or more indicators among the alarm indicators in the alarm record reaches an alarm threshold, fault self-healing is required.

[0078] In a third aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for realizing unified monitoring alarm fault self-healing as described in any one of the first aspects of the present application.

[0079] In a fourth aspect, the present invention also provides an electronic device, comprising: a memory storing a computer program; a processor communicating with the memory, and executing a method for realizing unified monitoring alarm fault self-healing as described in any one of the first aspects of the present application when calling the computer program.

[0080] Compared with the prior art, the present invention has the following obvious outstanding substantial features and significant advantages:

[0081] This invention provides a method and system for unified monitoring, alarming, and fault self-healing, resolving the existing problem of inability to uniformly implement monitoring, alarming, and fault self-healing functions. By collecting monitoring metadata and scheduling tasks according to pre-configured monitoring metadata configuration rules, unified monitoring, alarming, and fault self-healing are achieved, minimizing or eliminating the impact of system anomalies or failures, significantly reducing troubleshooting time. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] The accompanying drawings, which constitute part of the present invention, are provided to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the accompanying drawings:

[0083] Figure 1 The present invention is a flowchart of a method for realizing unified monitoring alarm fault self-healing according to a preferred embodiment of the present invention.

[0084] Figure 2 It is a schematic diagram of the structure of a system for realizing unified monitoring, alarming and fault self-healing according to a preferred embodiment of the present invention.

[0085] Description of labels:

[0086] 100. Monitoring metadata acquisition module; 200. Configuration rule setting module; 300. Monitoring alarm and self-healing processing module. DETAILED DESCRIPTION

[0087] The present invention provides a method and system for implementing unified monitoring, alarm, and fault self-healing. To clarify the objectives, technical solutions, and effects of the present invention, the present invention is further described below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention.

[0088] It should be noted that the terms "first," "second," and the like in the specification and claims of the present invention and the accompanying drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances. In addition, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatuses.

[0089] Example 1:

[0090] like Figure 1-Figure 2 As shown, the method for realizing unified monitoring alarm fault self-healing described in this embodiment specifically includes the following steps:

[0091] Step S1, externally obtain monitoring metadata and store it; wherein, the external access method includes one or more of http, dubbo, socket, mq and sdk; the storage location of the monitoring metadata includes one or more of mongodb (document database), dragonflydb (memory database) and OceanBase (distributed database).

[0092] Step S2: Setting monitoring metadata configuration rules;

[0093] In a specific implementation of this embodiment, the configuration rules specifically include tenant access configuration, monitoring indicator configuration, data source configuration, personnel management configuration, group management configuration, alarm rule configuration and fault self-healing rule configuration.

[0094] Among them, the tenant access configuration is used for external identity authentication, including tenant code, description, secret key, enterprise WeChat push key, user name, password and other field configurations.

[0095] Among them, the monitoring indicator configuration is used to set the indicators to be monitored, so as to facilitate reading the data corresponding to the indicators to be monitored from the metadata, including the configuration of fields such as indicator code, title, keyword, level, indicator description, and monitoring data source type; further, the monitoring data source type includes one or more of ES, Prometheus, SR, and MongoDB.

[0096] Among them, the data source configuration is used to monitor the configuration of the metadata storage data source, including the configuration of fields such as data source code, data source type, data source name, user name, password, service IP address, port, encoding, connection timeout, read timeout, etc.

[0097] The personnel management configuration is used to set the recipients of the alarm information, including the configuration of fields such as personnel code, name, telephone number, and email address.

[0098] The group management configuration is used to bind mutually related persons.

[0099] Among them, the alarm rule configuration is used to set the monitoring scheduling rules, data sources and alarm groups corresponding to each monitoring indicator, including the configuration of indicator type (including global and application levels), indicator code, execution frequency, execution frequency unit, data source, alarm method, application name, namespace (ie resource pool name), alarm trigger value, trigger value unit, relational operator, repeated alarm limit frequency, index name, alarm wording, title, data monitoring range, data monitoring range unit, environment (ie computer room identification) and other fields.

[0100] Among them, the execution frequency units include seconds, minutes and hours; the trigger value units include percentage and quantity; the relational operators include greater than, less than, equal to, greater than or equal to and less than or equal to.

[0101] Among them, the fault self-healing rule configuration is used to configure the logical rules for fault self-healing, including the configuration of automatic recovery rule serial number, indicator type, recovery type, title, execution frequency, automatic recovery limit frequency, execution frequency unit, trigger logic, notification method, recovery notification template, level, application name, namespace, environment identifier, automatic recovery prerequisites, multiple prerequisite logical judgment rules, data monitoring range, data monitoring range unit, alarm indicator code, trigger value, trigger condition, indicator title, application name, namespace, environment and other fields.

[0102] Furthermore, the recovery types include F5, configuration, database, data center switching, node removal, node restart, and privilege reduction.

[0103] Among them, the trigger logic includes logical (&&) and logical or (||); the prerequisite logic judgment rules include logical and and logical or; the data monitoring range units include seconds, minutes and hours.

[0104] Step S3: Perform task scheduling according to the monitoring metadata and the monitoring metadata configuration rules to complete monitoring, alarming, and fault self-healing.

[0105] Optionally, the task scheduling includes monitoring scheduling and automatic recovery scheduling.

[0106] Among them, the monitoring scheduling includes setting the monitoring execution frequency, alarm message push limit frequency and monitoring data query time range; the automatic recovery scheduling includes setting the automatic recovery execution frequency, automatic recovery limit frequency and alarm data query time range.

[0107] Furthermore, in step S3, task scheduling is performed according to the monitoring metadata and the monitoring metadata configuration rules to complete monitoring and alarming, which specifically includes steps S301a to S301c.

[0108] Step S301a: regularly read the monitoring metadata, and query metadata related to monitoring and alarming in combination with the monitoring indicator configuration and the alarm rule configuration.

[0109] Step S301b: parse the metadata related to monitoring and alarming obtained through query to complete metadata monitoring.

[0110] Step S301c: When the monitored metadata meets the conditions of the alarm rule, an alarm message is generated, relevant personnel are notified, and an alarm record is generated and saved; wherein, the alarm record facilitates subsequent problem backtracking and can also be used to trigger fault self-healing.

[0111] Furthermore, in step S3, task scheduling is performed according to the monitoring metadata and the monitoring metadata configuration rules to complete fault self-healing, which specifically includes steps S302a to S302c.

[0112] Step S302a: Read the alarm record regularly and determine whether fault self-healing is required based on the fault self-healing rule configuration. When one or more alarm indicators in the alarm record reach the alarm threshold, fault self-healing is required.

[0113] Step S302b: When performing fault self-healing, perform switching processing on the fault line, switch the first line (current fault line) to the second line (target normal line) and determine whether the precondition of the second line meets the line switching condition.

[0114] Step S302c: When the precondition of the second line (indicating that the switching of the second line requires the second line to be detected according to the rules configured in the metadata, and the second line will be switched only if the result is normal, to avoid incorrect line switching) meets the line switching condition, a judgment is made as to whether the current switching action is within the switching current limiting period. If it is not within the switching current limiting period, the line switching action is performed; otherwise, the line switching action is performed after the switching current limiting period ends.

[0115] Taking the automatic offline processing of F5 traffic when system abnormalities are detected as an example, it is listed that F5 traffic switching is a very typical fault self-healing method under the premise of multi-environment and multi-computer room deployment. By detecting abnormal system inlet traffic indicators, internal rpcdubbo cluster communication abnormalities, and a large number of system indicator abnormalities, and complying with the relevant fault self-healing configuration logic, the corresponding abnormal environment traffic can be temporarily offline by calling F5, allowing other normal environments to take over all traffic to quickly achieve the purpose of fault self-healing. When the system detects that the abnormal environment has returned to normal, it can perform grayscale traffic switching back according to relevant rules.

[0116] The overall process of F5 traffic switching fault self-recovery is as follows:

[0117] 1. The monitoring scheduling module triggers relevant monitoring alarm tasks according to the rules. For example, if the relevant alarm rule indicators are set: if the average HTTP time consumed by the system entrance in the last minute exceeds 2 seconds, the alarm condition is met. After reading the relevant metadata, the monitoring alarm task performs relevant logical matching judgment based on the threshold set by the current indicator. If it meets the requirements, it means that the entrance time of the relevant system in the last minute is high, and the system may be blocked or faulty. It is necessary to push relevant messages to the corresponding system person in charge in the corresponding manner according to the alarm method settings as soon as possible to attract their attention, and the alarm record should be stored in the database.

[0118] 2. The fault recovery scheduling module triggers fault recovery tasks based on scheduled rules. For example, if an F5 automatic failover task is configured, automatic recovery can be performed only when one or more conditions are met. For example, in scenario 1, if the alarm counts of two indicators in scenario 2 meet the same conditions, an F5 automatic failover will occur. Preconditions for automatic recovery can also be set, requiring the preconditions and selected monitoring indicators to be met before an F5 automatic failover occurs. For example, if there are two environments, Nanhui 1 and Nanhui 2, you can configure the alarm count of Nanhui 1's relevant indicator to be greater than 2, while comparing the alarm count of Nanhui 2's relevant indicator to be less than or equal to 1 or 0. The Nanhui 2 configuration is the opposite of the Nanhui 1 failover rule: the alarm count of Nanhui 2's relevant indicator is greater than 2, while the alarm count of Nanhui 1's relevant indicator is less than or equal to 1 or 0. By cross-comparing the two environments, it is determined whether preliminary failover conditions are met. If preliminary failover conditions are met, the preconditions for the number of online nodes in the other environment must be determined. If so, the preconditions for the number of online nodes in the other environment must be determined to be within the failover current limit period. This triple guarantee prevents system crashes caused by accidental failovers and frequent failovers.

[0119] Through the alarm records, we can match which environment currently has an alarm and which environment has no alarms and everything is normal. Then, we can take the traffic offline of the corresponding environment to achieve the effect of fault self-healing. When the system detects that the abnormal environment has returned to normal, it can perform grayscale switching back according to relevant rules.

[0120] Example 2:

[0121] like Figure 2 As shown, the system for realizing unified monitoring alarm fault self-healing described in this embodiment specifically includes a monitoring metadata acquisition module 100, a configuration rule setting module 200 and a monitoring alarm and self-healing processing module 300.

[0122] The monitoring metadata acquisition module 100 is used to externally acquire and store monitoring metadata; wherein, the external access method includes one or more of http, dubbo, socket, mq and sdk; the storage location of the monitoring metadata includes one or more of mongodb, dragonflydb and OceanBase.

[0123] The configuration rule setting module 200 is used to set monitoring metadata configuration rules.

[0124] The configuration rules specifically include tenant access configuration, monitoring indicator configuration, data source configuration, personnel management configuration, group management configuration, alarm rule configuration and fault self-healing rule configuration.

[0125] Among them, the tenant access configuration is used for external identity authentication, including tenant code, description, key, enterprise WeChat push key, user name, password and other field configurations.

[0126] The monitoring indicator configuration is used to set the indicators to be monitored, so as to facilitate reading the data corresponding to the indicators to be monitored from the metadata, including the configuration of fields such as indicator code, title, keyword, level, indicator description, and monitoring data source type; the monitoring data source type includes one or more of ES, Prometheus, SR, and MongoDB.

[0127] Among them, the data source configuration is used to monitor the configuration of the metadata storage data source, including the configuration of fields such as data source code, data source type, data source name, user name, password, service IP address, port, encoding, connection timeout, read timeout, etc.

[0128] The personnel management configuration is used to set the recipients of the alarm information, including the configuration of fields such as personnel code, name, telephone number, and email address.

[0129] The group management configuration is used to bind mutually related persons.

[0130] Among them, the alarm rule configuration is used to set the monitoring scheduling rules, data sources and alarm groups corresponding to each monitoring indicator, including the configuration of indicator type, indicator code, execution frequency, execution frequency unit, data source, alarm method, application name, namespace, alarm trigger value, trigger value unit, relational operator, repeated alarm limit frequency, index name, alarm wording, title, data monitoring range, data monitoring range unit, environment and other fields; the indicator types include global and application levels.

[0131] Among them, the execution frequency units include seconds, minutes and hours; the trigger value units include percentage and quantity; the relational operators include greater than, less than, equal to, greater than or equal to and less than or equal to.

[0132] Among them, the fault self-healing rule configuration is used to configure the logical rules for fault self-healing, including the automatic recovery rule serial number, indicator type, recovery type, title, execution frequency, automatic recovery limit frequency, execution frequency unit, trigger logic, notification method, recovery notification template, level, application name, namespace, environment identifier, automatic recovery prerequisites, multiple prerequisite logical judgment rules, data monitoring range, data monitoring range unit, alarm indicator code, trigger value, trigger condition, indicator title, application name, namespace, environment and other fields; the recovery types include F5, configuration, database, cutting computer room, removing nodes, node restarting, demotion and other types.

[0133] Among them, the trigger logic includes logical (&&) and logical or (||); the prerequisite logic judgment rules include logical and and logical or; the data monitoring range units include seconds, minutes and hours.

[0134] The monitoring alarm and self-healing processing module 300 is used to perform task scheduling according to the monitoring metadata and the monitoring metadata configuration rules, and complete monitoring, alarming and fault self-healing.

[0135] Among them, in the monitoring alarm and self-recovery processing module 300, the task scheduling includes monitoring scheduling and automatic recovery scheduling.

[0136] The monitoring scheduling includes setting the monitoring execution frequency, the alarm message push limit frequency and the monitoring data query time range.

[0137] The automatic recovery scheduling includes setting an automatic recovery execution frequency, an automatic recovery limit frequency, and an alarm data query time range.

[0138] The monitoring alarm and self-healing processing module 300 specifically includes a monitoring alarm submodule and a self-healing submodule.

[0139] The monitoring and alarm submodule is used to perform task scheduling according to the monitoring metadata and the monitoring metadata configuration rules, and complete monitoring and alarming, and specifically includes a first unit, a second unit and a third unit.

[0140] The first unit is used to periodically read the monitoring metadata, and query metadata related to monitoring and alarming in combination with monitoring indicator configuration and alarm rule configuration.

[0141] The second unit is used to parse the metadata related to monitoring and alarm obtained from the query and complete the monitoring of the metadata.

[0142] The third unit is used to generate an alarm message, notify relevant personnel, and generate and save an alarm record when the monitored metadata meets the conditions of the alarm rule; wherein, the alarm record facilitates subsequent problem tracing and can also be used to trigger fault self-healing.

[0143] The self-healing submodule is used to perform task scheduling according to the monitoring metadata and the monitoring metadata configuration rules to complete fault self-healing, and specifically includes a fourth unit, a fifth unit, and a sixth unit.

[0144] The fourth unit is used to periodically read the alarm record and, in combination with the fault self-healing rule configuration, determine whether fault self-healing is required; wherein, when one or more of the alarm indicators in the alarm record reaches the alarm threshold, fault self-healing is required.

[0145] The fifth unit is used to switch the fault line when performing fault self-recovery, switch the first line (current fault line) to the second line (target normal line) and determine whether the precondition of the second line meets the line switching condition.

[0146] The sixth unit is used to judge whether the current switching action is within the switching current limiting period when the precondition of the second line meets the line switching condition. If it is not within the switching current limiting period, the line switching action is performed; otherwise, the line switching action is performed after the switching current limiting period ends.

[0147] While the specific embodiments of the present invention have been described in detail above, these are merely exemplary and the present invention is not limited thereto. For those skilled in the art, any equivalent modifications and substitutions to the present invention are also within the scope of the present invention. Therefore, any equivalent changes and modifications made without departing from the spirit and scope of the present invention are intended to be encompassed within the scope of the present invention.

Claims

1. A method for realizing unified monitoring alarm fault self-healing, characterized in that: The specific steps include: Step S1: externally obtain monitoring metadata and store it; Step S2: Setting monitoring metadata configuration rules; Step S3: Perform task scheduling according to the monitoring metadata and the monitoring metadata configuration rules to complete monitoring, alarming, and fault self-healing; In step S1, the external access method includes one or more of http, dubbo, socket, mq and sdk; the storage location of the monitoring metadata includes one or more of mongodb, dragonflydb and OceanBase.

2. A method for realizing unified monitoring alarm fault self-healing according to claim 1, characterized in that: In step S2, the configuration rules specifically include tenant access configuration, monitoring indicator configuration, data source configuration, personnel management configuration, group management configuration, alarm rule configuration and fault self-healing rule configuration.

3. A method for realizing unified monitoring alarm fault self-healing according to claim 2, characterized in that: The tenant access configuration is used for external identity authentication, including tenant code, description, key, enterprise WeChat push key, user name, and password.

4. A method for realizing unified monitoring alarm fault self-healing according to claim 2, characterized in that: The monitoring indicator configuration is used to set the indicators to be monitored to facilitate reading the data corresponding to the indicators to be monitored from the metadata, including the indicator code, title, keyword, level, indicator description, and monitoring data source type; wherein the monitoring data source type includes one or more of ES, Prometheus, SR, and MongoDB.

5. The method for realizing unified monitoring alarm fault self-healing according to claim 2, characterized in that: The data source configuration is used to monitor the configuration of the metadata storage data source, including field configurations such as data source code, data source type, data source name, user name, password, service IP address, port, code, connection timeout, and read timeout. The personnel management configuration is used to set the recipients of alarm information, including personnel code, name, telephone number, and email address. The group management configuration is used to bind related personnel.

6. A method for realizing unified monitoring alarm fault self-healing according to claim 2, characterized in that: The alarm rule configuration is used to set the monitoring scheduling rules, data sources and alarm groups corresponding to each monitoring indicator, including indicator type, indicator code, execution frequency, execution frequency unit, data source, alarm method, application name, namespace, alarm trigger value, trigger value unit, relational operator, repeated alarm limit frequency, index name, alarm wording, title, data monitoring range, data monitoring range unit, and environment; wherein, the indicator type includes global and application level; the execution frequency units include seconds, minutes and hours, and the trigger value units include percentage and quantity; the relational operators include greater than, less than, equal to, greater than or equal to, and less than or equal to.

7. A method for realizing unified monitoring alarm fault self-healing according to claim 2, characterized in that: The fault self-healing rule configuration is used to configure the logical rules for fault self-healing, including the configuration of fields such as the automatic recovery rule serial number, indicator type, recovery type, title, execution frequency, automatic recovery limit frequency, execution frequency unit, trigger logic, notification method, recovery notification template, level, application name, namespace, environment identifier, automatic recovery prerequisite, multiple prerequisite logical judgment rules, data monitoring range, data monitoring range unit, alarm indicator code, trigger value, trigger condition, indicator title, application name, namespace, environment, etc.; wherein, the recovery type includes F5, configuration, database, data center cut, node removal, node restart, and privilege reduction; Among them, the trigger logic includes logical AND and logical OR; the prerequisite logic judgment rules include logical AND and logical OR; the data monitoring range units include seconds, minutes and hours.

8. The method for realizing unified monitoring alarm fault self-healing according to claim 2, characterized in that: In step S3, task scheduling is performed according to the monitoring metadata and the monitoring metadata configuration rules to complete monitoring and alarming, which specifically includes the following steps: Step S301a: regularly read the monitoring metadata, and query metadata related to monitoring and alarming in combination with the monitoring indicator configuration and the alarm rule configuration; Step S301b: Parse the metadata related to monitoring and alarming obtained from the query to complete metadata monitoring; Step S301c: When the monitored metadata meets the conditions of the alarm rule, an alarm message is generated, relevant personnel are notified, and an alarm record is generated and saved.

9. A method for realizing unified monitoring alarm fault self-healing according to claim 2, characterized in that: In step S3, task scheduling is performed according to the monitoring metadata and the monitoring metadata configuration rules to complete fault self-healing, which specifically includes the following steps: Step S302a: Read the alarm record regularly and determine whether fault self-healing is required based on the fault self-healing rule configuration; Step S302b: When the fault is self-healing, the fault line is switched, the first line is switched to the second line, and it is determined whether the precondition of the second line meets the line switching condition; Step S302c: When the precondition of the second line meets the line switching condition, determine whether the current switching action is within the switching current limiting period. If it is not within the switching current limiting period, perform the line switching action; otherwise, wait until the switching current limiting period ends before performing the line switching action.

10. A system for realizing unified monitoring, alarm and fault self-healing, characterized in that: Specifically, it includes the following modules: Monitoring metadata acquisition module, used to externally acquire monitoring metadata and store it; Configuration rule setting module, used to set monitoring metadata configuration rules; The monitoring alarm and self-healing processing module is used to perform task scheduling according to the monitoring metadata and the monitoring metadata configuration rules, and complete monitoring, alarming and fault self-healing.