A health monitoring system and method under a multi-site multi-center architecture
Through the health monitoring system under the multi-site and multi-center architecture, using the abnormal detection and analysis and scoring subsystems, the problem of difficulty in monitoring small faults in the computer room system has been solved, automated health monitoring and scoring have been achieved, and the accuracy and efficiency of monitoring have been improved.
Patent Information
- Application Number
- CN202111226110.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-21
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2041-10-21
AI Technical Summary
In a multi-site and multi-center computer room system, existing technologies are difficult to accurately and timely monitor a small number of or individual machine failures. Manual monitoring has the problems of high labor costs and prone to errors.
A health monitoring system with a multi-site, multi-center architecture is adopted, including a multi-site, multi-center deployment subsystem, an abnormality detection and analysis subsystem, and a scoring subsystem, to achieve automated health monitoring and scoring. The abnormality detection and analysis subsystem detects component abnormalities and performs disaster recovery switching. The scoring subsystem scores and displays the health status based on health information.
It achieves accurate, timely and flexible judgment of abnormal problems in multi-site and multi-center architectures, improves the accuracy and efficiency of monitoring, and supports automatic switching and health status display in major disaster scenarios.
Smart Images

Figure CN113934600B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of system management technology, and in particular to a health monitoring system and method under a multi-site multi-center architecture. Background Art
[0002] To achieve disaster recovery and multi-active data center deployments, enterprises often need to build multiple locations and centers to meet business needs. However, this often leads to challenges. Determining the health of vast data center systems and determining when to switch data centers have become challenging issues for enterprises. Currently, manual monitoring of the operating status of machines in each data center is a common approach. While this approach can detect major disasters like power outages, fires, and network outages, it struggles to detect failures in smaller numbers or even individual machines. Furthermore, manual monitoring increases labor costs and is prone to errors and mismatches, which can negatively impact the enterprise.
[0003] Therefore, there is an urgent need for a technical solution that can overcome the above problems and efficiently monitor the health system of the computer room system. Summary of the Invention
[0004] To address the challenges of existing technologies, this paper proposes a health monitoring system and method for a multi-site, multi-center architecture. This system accurately, promptly, and automatically identifies various potential anomalies, addressing the automated identification of inter-machine room handoffs within a multi-site, multi-center architecture. It also addresses the monitoring needs for major disaster scenarios and equipment and component failures. Furthermore, it demonstrates the health status of each architecture through health scores, effectively improving monitoring accuracy and efficiency.
[0005] In a first aspect of an embodiment of the present invention, a health monitoring system in a multi-site multi-center architecture is proposed, the system comprising:
[0006] Multi-site and multi-center deployment subsystem, deployed in a multi-site and multi-center architecture, is used to monitor the health information of the architecture;
[0007] The anomaly detection and analysis subsystem is used to detect anomalies in components under the architecture and perform disaster recovery switching based on the detection results;
[0008] The scoring subsystem is used to perform health scoring on the multi-site and multi-center architecture based on the set scoring configuration parameters and the health information, and evaluate the health status based on the scoring results.
[0009] In a second aspect of an embodiment of the present invention, a health monitoring method under a multi-site multi-center architecture is proposed. The health monitoring method under the multi-site multi-center architecture is implemented based on a multi-site multi-center deployment subsystem, an abnormality detection and analysis subsystem, and a scoring subsystem; wherein,
[0010] Deploy multi-site and multi-center deployment subsystems in a multi-site and multi-center architecture to monitor the health information of the architecture;
[0011] The anomaly detection and analysis subsystem detects component anomalies in the architecture and performs disaster recovery switching based on the detection results.
[0012] The scoring subsystem performs health scoring on the multi-site and multi-center architecture according to the set scoring configuration parameters and the health information, and evaluates the health status based on the scoring results.
[0013] In a third aspect of an embodiment of the present invention, a computer device is proposed, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a health monitoring method is implemented under a multi-site and multi-center architecture.
[0014] In a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is proposed, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a health monitoring method under a multi-site multi-center architecture is implemented.
[0015] The health monitoring system and method under the multi-site multi-center architecture proposed in the present invention deploys the multi-site multi-center deployment subsystem under the multi-site multi-center architecture to monitor the health information of the architecture. The abnormality detection and analysis subsystem is used to detect component abnormalities under the architecture and perform disaster recovery switching according to the detection results. The scoring subsystem is used to perform health scoring on the multi-site multi-center architecture based on the set scoring configuration parameters and the health information, and evaluate the health status based on the scoring results. The overall solution can accurately, timely and flexibly judge abnormal problems that may occur under the multi-site multi-center architecture, meet the monitoring of major disaster scenarios and failures of equipment, components, etc., and display the health status of each architecture through health scoring, effectively improving the accuracy and efficiency of monitoring. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0017] Figure 1 2 is a schematic diagram of a health monitoring system architecture under a multi-site multi-center architecture according to an embodiment of the present invention.
[0018] Figure 2 FIG. 1 is a schematic diagram of the architecture of an abnormality detection and analysis subsystem according to an embodiment of the present invention.
[0019] Figure 3FIG. 4 is a schematic diagram of a scoring subsystem architecture according to an embodiment of the present invention.
[0020] Figure 4 2 is a schematic diagram of a health monitoring system architecture under a multi-site multi-center architecture according to a specific embodiment of the present invention.
[0021] Figure 5 It is a schematic diagram of the relationship of distributed detection in a specific embodiment of the present invention.
[0022] Figure 6 It is a schematic diagram of the relationship between the customized combination configuration of the registration module in a specific embodiment of the present invention.
[0023] Figure 7 1 is a flow chart of a health monitoring method under a multi-site multi-center architecture according to an embodiment of the present invention.
[0024] Figure 8 It is a schematic diagram of the structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided solely to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. Rather, these embodiments are provided to make this disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0026] Those skilled in the art will appreciate that the embodiments of the present invention may be implemented as a system, apparatus, device, method, or computer program product. Therefore, the present disclosure may be implemented in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software.
[0027] According to an embodiment of the present invention, a health monitoring system and method under a multi-site multi-center architecture are proposed, which relate to the field of system management technology.
[0028] The principles and spirit of the present invention are explained in detail below with reference to several representative embodiments of the present invention.
[0029] Figure 1 FIG is a schematic diagram of a health monitoring system architecture under a multi-site multi-center architecture according to an embodiment of the present invention. Figure 1 As shown, the system includes:
[0030] The multi-site multi-center deployment subsystem 110 is deployed in a multi-site multi-center architecture and is used to monitor the health information of the architecture;
[0031] The abnormality detection and analysis subsystem 120 is used to detect abnormalities in components under the architecture and perform disaster recovery switching based on the detection results;
[0032] The scoring subsystem 130 is used to perform health scoring on the multi-site and multi-center architecture according to the set scoring configuration parameters and the health information, and evaluate the health status according to the scoring results.
[0033] In order to explain the health monitoring system under the multi-site multi-center architecture more clearly, an embodiment is described below.
[0034] The multi-site multi-center deployment subsystem 110 is composed of multi-site multi-center deployment modules, which are deployed in a multi-site multi-center architecture.
[0035] refer to Figure 2 , is a schematic diagram of the abnormal detection and analysis subsystem architecture of an embodiment of the present invention. Figure 2 As shown, the abnormality detection and analysis subsystem 120 includes:
[0036] The personalized monitoring module 121 is used to obtain monitoring configuration information and configure the monitoring components according to the monitoring configuration information;
[0037] The participant modules 122 are deployed in a multi-site and multi-center architecture;
[0038] The global coordinator module 123 is deployed under one of the architectures;
[0039] When the global coordinator module 123 obtains the health detection task, it calls the participant module 122, which aggregates the metadata set;
[0040] The abnormality detection module 124 is used to detect component abnormalities based on the metadata set and perform disaster recovery switching based on the detection results.
[0041] Further reference Figure 2 The abnormality detection and analysis subsystem 120 further includes:
[0042] The registration module 125 is deployed in multiple locations and centers, and is used to initialize the component registration information under the architecture and obtain the scoring configuration parameters set by the user; wherein, the component registration information at least identifies information including the computer room and shard, and supports cross-computer room writing.
[0043] In this embodiment, the abnormal activity detection module 124 is specifically used to:
[0044] When multiple nodes go down, the detection result is abnormal and triggers an alarm, prompting monitoring personnel to switch to other shards to continue providing services.
[0045] Among them, the multi-location and multi-center architecture includes components in dimensions such as cities, centers, computer rooms, and shards, as well as links between logical objects.
[0046] In this embodiment, the personalized monitoring module 121 is specifically used to:
[0047] The user independently configures the monitored components and the weights of the monitored components. The monitored components include at least the middleware, services, physical hardware, and network in the city, center, computer room, and shard.
[0048] refer to Figure 3 , is a schematic diagram of the scoring subsystem architecture of an embodiment of the present invention. Figure 3 As shown, the scoring subsystem 130 includes:
[0049] User module 131, used to record the user's operation content and permissions each time;
[0050] The service module 132 is used to annotate components, clusters, and services that need to be monitored and to initialize the registration of health information including the initial reporting of metadata;
[0051] The permission module 133 is used to set the user's permissions;
[0052] The scoring module 134 is used to weight the marked components, clusters, and services according to the health information to obtain a health score, evaluate the health status according to the score result, and display it through a visual interface.
[0053] It should be noted that although the above detailed description mentions several subsystems and modules of the health monitoring system under the multi-site multi-center architecture, this division is merely exemplary and not mandatory. In fact, according to embodiments of the present invention, the features and functions of two or more subsystems and modules described above can be concretized in one subsystem or module. Conversely, the features and functions of one subsystem or module described above can be further divided to be concretized by multiple subsystems or modules.
[0054] In order to better illustrate the health monitoring system under the multi-site multi-center architecture of the present invention, a specific embodiment is described below.
[0055] refer to Figure 4 , is a schematic diagram of a health monitoring system architecture under a multi-site multi-center architecture according to a specific embodiment of the present invention. Figure 4 As shown, the health monitoring system can be deployed in a multi-site multi-center architecture. In addition to meeting the basic health detection functions of various components, services, networks, physical machines, routing, storage devices, etc., it can also adapt to the "multi-site multi-center" (for example, four sites and eight centers) unitized architecture.
[0056] Every city, center, data center, shard, component, and the links between logical objects are considered monitoring and statistical targets. The detected logical objects are collectively linked. That is, a city-level detection includes the monitoring and statistical results of all centers within that city. Similarly, each center also includes the monitoring information of all its data centers. Overall health monitoring can be viewed as a tree-like topology.
[0057] Supports multi-site, multi-datacenter monitoring, encompassing every city, center, datacenter, unit, and component. It provides macro-level visibility into global health, as well as micro-level visibility into the health of each node, component, and link. It adapts to virtualization and containerization for flexible scaling and capacity expansion. It also supports grayscale health monitoring. The health monitoring system itself supports high availability, as well as synchronized backup of monitoring data across different city centers.
[0058] This system consists of three subsystems: a multi-site, multi-center deployment subsystem, an anomaly detection and analysis subsystem, and a scoring subsystem. Within a multi-site, multi-center architecture, this system implements automatic disaster recovery failover, automatically determines whether to initiate inter-center failover, and can be deployed across multiple sites and multiple centers. The anomaly detection and analysis subsystem accurately, promptly, and automatically identifies various potential anomalies, providing reliable input for inter-computer room failover. The scoring subsystem allows for autonomous configuration of modules to be monitored and calculates a final health score.
[0059] Specifically, refer to Figure 4 ,The multi-site multi-center deployment subsystem is composed of multi-site multi-center ,deployment modules, and is deployed under a multi-site multi-center architecture.
[0060] The abnormal activity detection and analysis subsystem includes personalized monitoring module, global coordinator module, participant module, registration module, and abnormal activity detection module;
[0061] The personalized monitoring module supports customized monitoring. This module comprehensively covers all system components and possesses platform capabilities, supporting all middleware, services, physical components, networks, etc.
[0062] In the personalized monitoring module, users can independently configure the components they want to monitor, such as middleware, services, physical hardware, and networks. They can also select the nodes, services, components, units, computer rooms, and centers they want to monitor. They can also configure weights for different components to prioritize the content they want to monitor, achieving personalized monitoring.
[0063] Use the global coordinator module and the participating module to realize distributed detection, refer to Figure 5 , is a diagram showing the relationship between distributed detection. Figure 5As shown, in a multi-site, multi-active unitized architecture (for example, GZong1, GZone2, RZone1, and RZone2), service activity detection is relatively complex. In addition to considering the high availability of multiple sites and multiple computer rooms, it is also necessary to consider each computer room, unit, and shard.
[0064] This system uses a global coordinator module to implement health detection in batches. Through distributed scheduling, a global coordinator is established to initiate health detection tasks for each batch. Each of the other partitions has a participant module responsible for collecting and initially summarizing all the meta-health information for its area of responsibility.
[0065] The registration module provides the unit service initialization customized registration information and configuration information. The purpose of the participant module is to aggregate the metadata collection of the detection, such as Figure 6 As shown, the registration module can provide users with customized configuration combinations (such as components a, b, c, and d) to facilitate customized summary display of the scoring subsystem. Each data center deploys a registration module, and the two registration modules should share a database. Component registration information should identify information such as the data center and shard, and support cross-data center writes.
[0066] The abnormal activity detection module provides accurate decision-making for disaster recovery switching, especially for cluster, shard, and system failures caused by multi-node outages. If a shard fails entirely and can no longer provide full service, the abnormal activity detection module triggers an alarm, prompting monitoring personnel to switch to another shard to resume service. Furthermore, the system possesses sufficient platform support capabilities, adapts to a modular architecture, and provides a unified, standardized activity detection interface that accommodates all components and services within the platform.
[0067] The scoring subsystem includes at least user module, service module, scoring module, and permission module, and provides a clear and friendly visual interface.
[0068] The service module must first annotate the components, clusters, and services to be detected, such as databases, caches, computer rooms, and microservices. Secondly, it must include metadata for initial registration when health information is first reported, and support both horizontal and vertical scaling. The scoring subsystem provides self-configuration capabilities, weighting the middleware, services, storage, and routing components required for registration based on different business entities and health concerns, ultimately generating an overall health score.
[0069] In specific application scenarios, users can perform personalized operations in the health scoring subsystem. The user module records the content and permissions of each user's operation. In the service module, users can select the components, clusters, services, and other types to be detected, such as databases, caches, computer rooms, and microservices. The permissions module allows users to set operational user permissions or revoke permissions. The scoring module calculates a total score based on the user's configuration and the current health status of the computer room.
[0070] In summary, the health monitoring system under the multi-site multi-center architecture proposed in the present invention can accurately, timely and flexibly judge the abnormal problems that may occur in the multi-site multi-center architecture, meet the monitoring of major disaster scenarios and failures of equipment and components, and display the health status of each architecture through health scores, effectively improving the accuracy and efficiency of monitoring.
[0071] After introducing the system of the exemplary embodiment of the present invention, next, reference is made to Figure 7 A health monitoring method under a multi-site multi-center architecture according to an exemplary embodiment of the present invention is introduced.
[0072] Figure 7 This is a flow chart of a health monitoring method under a multi-site multi-center architecture according to an embodiment of the present invention. Figure 7 As shown, the health monitoring method under the multi-site multi-center architecture is implemented based on the multi-site multi-center deployment subsystem, the abnormality detection and analysis subsystem, and the scoring subsystem; wherein, the method includes:
[0073] S101, deploys multi-site and multi-center deployment subsystems in a multi-site and multi-center architecture to monitor the health information of the architecture;
[0074] S102: Detect component anomalies in the architecture through the anomaly detection and analysis subsystem, and perform disaster recovery switching based on the detection results.
[0075] S103, using the scoring subsystem to perform health scoring on the multi-site and multi-center architecture according to the set scoring configuration parameters and the health information, and evaluating the health status according to the scoring results.
[0076] In one embodiment, S102, the specific process of detecting component anomalies in the architecture through the anomaly detection and analysis subsystem and performing disaster recovery switching based on the detection results is as follows:
[0077] Obtain monitoring configuration information and configure monitoring components based on the monitoring configuration information;
[0078] Deploy the participant modules in multiple locations and multiple centers, and deploy the global coordinator module in one of the architectures;
[0079] When the global coordinator module obtains the health detection task, it calls the participant module, which summarizes the metadata set;
[0080] Detect component anomalies based on metadata collection and perform disaster recovery switching based on the detection results.
[0081] In one embodiment, the method further comprises:
[0082] Deploy registration modules in multiple locations and centers, initialize component registration information under the architecture, and obtain user-set scoring configuration parameters; the component registration information at least identifies information including the computer room and shard, and supports cross-computer room writing.
[0083] In one embodiment, the specific process of obtaining monitoring configuration information and configuring the monitored components according to the monitoring configuration information is as follows:
[0084] The user independently configures the monitored components and the weights of the monitored components. The monitored components include at least the middleware, services, physical hardware, and network in the city, center, computer room, and shard.
[0085] In one embodiment, the specific process of disaster recovery switching is as follows:
[0086] When multiple nodes go down, the detection result is abnormal and triggers an alarm, prompting monitoring personnel to switch to other shards to continue providing services.
[0087] Among them, the multi-location and multi-center architecture includes components in dimensions such as cities, centers, computer rooms, and shards, as well as links between logical objects.
[0088] In one embodiment, S103, the scoring subsystem performs a health score on the multi-site and multi-center architecture based on the set scoring configuration parameters and the health information. The specific process of evaluating the health status based on the scoring results is as follows:
[0089] Record the user's operation content and permissions each time;
[0090] Label the components, clusters, and services that need to be monitored, and initialize the registration of health information including the initial reporting of metadata;
[0091] Set user permissions;
[0092] Based on the health information, the labeled components, clusters, and services are weighted to obtain a health score. The health status is evaluated based on the score result and displayed through a visual interface.
[0093] It should be noted that although the operations of the method of the present invention are described in a specific order in the above embodiments and drawings, this does not require or imply that these operations must be performed in this specific order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0094] Based on the above invention concept, Figure 8 As shown, the present invention also proposes a computer device 800, including a memory 810, a processor 820, and a computer program 830 stored in the memory 810 and executable on the processor 820. When the processor 820 executes the computer program 830, the health monitoring method under the aforementioned multi-site multi-center architecture is implemented.
[0095] Based on the aforementioned inventive concept, the present invention proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the aforementioned health monitoring method under the multi-site multi-center architecture.
[0096] The health monitoring system and method under the multi-site multi-center architecture proposed in the present invention deploys the multi-site multi-center deployment subsystem under the multi-site multi-center architecture to monitor the health information of the architecture. The abnormality detection and analysis subsystem is used to detect component abnormalities under the architecture and perform disaster recovery switching according to the detection results. The scoring subsystem is used to perform health scoring on the multi-site multi-center architecture based on the set scoring configuration parameters and the health information, and evaluate the health status based on the scoring results. The overall solution can accurately, timely and flexibly judge abnormal problems that may occur under the multi-site multi-center architecture, meet the monitoring of major disaster scenarios and failures of equipment, components, etc., and display the health status of each architecture through health scoring, effectively improving the accuracy and efficiency of monitoring.
[0097] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0098] The present invention is described with reference to flowcharts and / or block diagrams of methods and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0099] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0100] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0101] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A health monitoring system under a multi-site multi-center architecture, characterized in that: The system includes: Multi-site and multi-center deployment subsystem, deployed in a multi-site and multi-center architecture, is used to monitor the health information of the architecture; The anomaly detection and analysis subsystem is used to detect anomalies in components under the architecture and perform disaster recovery switching based on the detection results; A scoring subsystem, configured to perform a health score on the multi-site and multi-center architecture based on the set scoring configuration parameters and the health information, and to evaluate the health status based on the scoring results; The abnormal detection and analysis subsystem includes: Participant modules are deployed in a multi-site and multi-center architecture; The global coordinator module is deployed in one of the architectures; When the global coordinator module obtains a health detection task, it calls the participant module, which summarizes the metadata set. The global coordinator module is used to implement detection in batches. Through the distributed coordination concept, a global coordinator is established to initiate the health detection task for each batch. Each other partition has a participant module, which is responsible for collecting all the meta-health information for which it is responsible and performing preliminary aggregation. The abnormality detection module is used to detect component abnormalities based on metadata sets and perform disaster recovery switching based on the detection results. The abnormality detection module provides a decision basis for disaster recovery switching and handles cluster, shard, and system failures that may be caused by multi-node failures. When a shard is completely down and can no longer provide complete services, the abnormality detection module triggers an alarm, prompting monitoring personnel to switch to other shards to continue providing services. The system has sufficient platform support capabilities, adapts to the unitized architecture, and provides a unified and standardized detection interface to adapt to various components and services in the platform. The scoring subsystem includes: User module, used to record the user's operation content and permissions each time; The service module is used to annotate components, clusters, and services that need to be monitored, and to initialize the registration of health information including the initial reporting of metadata; Permission module, used to set user permissions; The scoring module is used to weight the labeled components, clusters, and services according to the health information to obtain a health score, evaluate the health status according to the score result, and display it through a visual interface.
2. The health monitoring system under the multi-site multi-center architecture according to claim 1, characterized in that: The abnormality detection and analysis subsystem includes: The personalized monitoring module is used to obtain monitoring configuration information and configure the monitoring components according to the monitoring configuration information.
3. The health monitoring system under the multi-site multi-center architecture according to claim 2, characterized in that: The abnormality detection and analysis subsystem shown also includes: The registration module is deployed in multiple locations and centers to initialize the component registration information under the architecture and obtain the scoring configuration parameters set by the user. Among them, the component registration information at least identifies information including the computer room and shard, and supports cross-computer room writing.
4. The multi-site multi-center health monitoring system according to claim 2, characterized in that: The abnormal activity detection module is specifically used to: When multiple nodes go down, the detection result is abnormal and triggers an alarm, prompting monitoring personnel to switch to other shards to continue providing services.
5. The health monitoring system under the multi-site multi-center architecture according to claim 2, characterized in that: The multi-location and multi-center architecture includes components in the dimensions of cities, centers, computer rooms, and shards, as well as links between logical objects.
6. The health monitoring system under the multi-site multi-center architecture according to claim 5, characterized in that: The personalized monitoring module is specifically used for: The user independently configures the monitored components and the weights of the monitored components. The monitored components include at least the middleware, services, physical hardware, and network in the city, center, computer room, and shard.
7. A health monitoring method under a multi-site multi-center architecture, characterized in that: The health monitoring method under the multi-site multi-center architecture is implemented based on the multi-site multi-center deployment subsystem, the abnormality detection and analysis subsystem and the scoring subsystem; wherein, Deploy multi-site and multi-center deployment subsystems in a multi-site and multi-center architecture to monitor the health information of the architecture; The anomaly detection and analysis subsystem detects component anomalies in the architecture and performs disaster recovery switching based on the detection results. The scoring subsystem performs health scoring on the multi-site and multi-center architecture according to the set scoring configuration parameters and the health information, and evaluates the health status according to the scoring results; The method further includes: Deploy the participant modules in multiple locations and multiple centers, and deploy the global coordinator module in one of the architectures; When the global coordinator module obtains a health detection task, it calls the participant module, which summarizes the metadata set. The global coordinator module is used to implement detection in batches. Through the distributed coordination concept, a global coordinator is established to initiate the health detection task for each batch. Each other partition has a participant module, which is responsible for collecting all the meta-health information for which it is responsible and performing preliminary aggregation. Component anomalies are detected based on metadata collections, and disaster recovery switching is performed based on the detection results. The abnormal liveness detection module provides a basis for decision-making for disaster recovery switching, and handles cluster, shard, and system failures that may be caused by multi-node failures. When a shard is completely down and can no longer provide complete services, the abnormal liveness detection module triggers an alarm, prompting monitoring personnel to switch to other shards to continue providing services. The system has sufficient platform support capabilities, adapts to the unitized architecture, and provides a unified and standardized liveness detection interface to accommodate various components and services in the platform. The scoring subsystem performs health scoring on the multi-site and multi-center architecture based on the set scoring configuration parameters and the health information, and evaluates the health status based on the scoring results, including: Record the user's operation content and permissions each time; Label the components, clusters, and services that need to be monitored, and initialize the registration of health information including the initial reporting of metadata; Set user permissions; Based on the health information, the labeled components, clusters, and services are weighted to obtain a health score. The health status is evaluated based on the score result and displayed through a visual interface.
8. The health monitoring method under a multi-site multi-center architecture according to claim 7, characterized in that: The anomaly detection and analysis subsystem detects component anomalies in the architecture and performs disaster recovery switching based on the detection results, including: Get monitoring configuration information and configure monitoring components based on the monitoring configuration information.
9. The health monitoring method under a multi-site multi-center architecture according to claim 8, characterized in that: The method further includes: Deploy registration modules in multiple locations and centers, initialize component registration information under the architecture, and obtain user-set scoring configuration parameters; the component registration information at least identifies information including the computer room and shard, and supports cross-computer room writing.
10. The health monitoring method under a multi-site multi-center architecture according to claim 8, characterized in that: Detect component anomalies based on metadata collection and perform disaster recovery switching based on the detection results, including: When multiple nodes go down, the detection result is abnormal and triggers an alarm, prompting monitoring personnel to switch to other shards to continue providing services.
11. The health monitoring method under a multi-site multi-center architecture according to claim 8, characterized in that: The multi-location and multi-center architecture includes components in the dimensions of cities, centers, computer rooms, and shards, as well as links between logical objects.
12. The health monitoring method under a multi-site multi-center architecture according to claim 11, characterized in that: Obtain monitoring configuration information and configure monitoring components based on the monitoring configuration information, including: The user independently configures the monitored components and the weights of the monitored components. The monitored components include at least the middleware, services, physical hardware, and network in the city, center, computer room, and shard.
13. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 7 to 12 is implemented.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 7 to 12 is implemented.
Citation Information
Patent Citations
Cluster disaster recovery management system based on ETCD
CN111371599A
Data disaster recovery method, device and system
US20190095293A1