Monitoring alarm method and system based on cross-regional hybrid cloud
By classifying and defining the basic information of cloud nodes in a cross-regional hybrid cloud environment, the problem of large and scattered data is solved, unified data management and precise alarms are realized, maintenance complexity and entry threshold are reduced, and architecture is simplified.
Patent Information
- Application Number
- CN202510695642.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-18
AI Technical Summary
In a cross-regional hybrid cloud environment, the amount of monitoring data in cloud nodes is large and scattered, resulting in complex data management, high entry threshold, high maintenance costs, and inability to effectively link services to alert.
Through the data management center node, the basic information of each cloud node is classified and defined, the cloud configuration information is collected, the tag information and associated flag information are issued, and the automatic detection node is used to push alarms to the business line group to realize a cross-regional hybrid cloud monitoring and alarm system.
It realizes unified data management in cross-regional hybrid cloud, precise alarm push, reduces construction and maintenance costs, improves scalability and maintainability, simplifies architecture, and facilitates use and maintenance.
Smart Images

Figure CN120342834A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of operation and maintenance monitoring, and particularly to a monitoring and alarming method and system based on a cross-region hybrid cloud. Background Art
[0002] For data monitoring and alarming in the cloud, there are two common practices: 1. Uniformly collect data from each region and each cloud architecture and summarize it into a database or platform, and then separately configure specified business line groups to send alarms and monitor the data. Because the data interaction and transmission are very low, the data is scattered, and the setup is simple, it is often applied to small-scale monitoring teams with small data volumes and low entry thresholds, such as the early design of zabbix monitoring. 2. Manage and monitor data separately on each region and each cloud to achieve a self-closed loop of monitoring and alarming. This involves data collaboration and management among multiple teams. Due to many rules, high management costs, large data volumes, complex setup, high interaction and transmission costs, and complex links, it is often used by large-scale monitoring teams with large monitoring data volumes and high entry thresholds.
[0003] The above content is only used to assist in understanding the technical solution of the present application, and does not represent an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of the present application is to provide a monitoring and alarming method and system based on a cross-region hybrid cloud, aiming to solve the technical problem that it is impossible to link services for alarming due to a large amount of monitoring data of cloud nodes.
[0005] To achieve the above purpose, the present application proposes a monitoring and alarming system based on a cross-region hybrid cloud. The monitoring and alarming system based on a cross-region hybrid cloud includes a data management center node and multiple cloud nodes. Monitoring components, alarm nodes, and automatic detection nodes are deployed on each cloud node. Each cloud node is at least one type of cloud server deployed in multiple regions. The data management center node is respectively connected to each cloud node;
[0006] The data management center node is used to classify and define the basic information of each cloud node to obtain acquisition classification information;
[0007] The cloud node is used to collect data through the monitoring component, and perform automatic alarming based on the acquisition classification information through the automatic detection node and the alarm node.
[0008] Optionally, the method is applied to the data management center node of the monitoring and alarming system based on a cross-region hybrid cloud as described in claim 1;
[0009] The method includes:
[0010] Classify and define the basic information of each cloud node based on the cloud configuration information collected by the monitoring components corresponding to each node to obtain the collected classification information;
[0011] Send label information and association flag information to each cloud node based on the collected classification information, so as to push alarms to the corresponding business line groups through the automatic detection nodes of the cloud nodes.
[0012] Optionally, before the step of classifying and defining the basic information of each cloud node based on the cloud configuration information collected by the monitoring components corresponding to each node to obtain the collected classification information, the following steps are further included:
[0013] Obtain the address information and type information of each cloud node collected by the monitoring components according to the monitoring rules and the target scripts stored in the time series databases arranged on each cloud node.
[0014] Optionally, before the step of obtaining the address information and type information of each cloud node by calling the target scripts stored in the time series databases arranged on each cloud node according to the monitoring rules, the following steps are further included:
[0015] Generate and store monitoring rules, and send the monitoring rules to the specified paths corresponding to the monitoring components on each cloud node.
[0016] Optionally, the step of obtaining the cloud configuration information of each cloud node based on the address information and the type information includes:
[0017] Determine the data source address information and configuration file information of each cloud node based on the address information and the type information;
[0018] Obtain the cloud configuration information of each cloud node according to the data source address information and the configuration file information.
[0019] Optionally, the step of classifying and defining the basic information of each cloud node based on the cloud configuration information to obtain the collected classification information includes:
[0020] Call the application programming interfaces or automatic scripts corresponding to each cloud node to collect the basic information of each cloud node;
[0021] Classify and define based on the basic information and the cloud configuration information to obtain the collected classification information.
[0022] Optionally, the step of classifying and defining based on the basic information and the cloud configuration information to obtain the collected classification information includes:
[0023] Determine label information according to the cloud configuration information;
[0024] Classify and define the basic information based on the tag information to obtain the acquisition classification information.
[0025] Optionally, the method is applied to the automatic detection nodes in each cloud node of the monitoring and warning system based on cross-region hybrid cloud as described in claim 1.
[0026] The method includes:
[0027] Perform activity detection on the data of each cloud node.
[0028] Request associated flag information from the data management center node according to the tag information sent by the data management center node and the detection result.
[0029] Determine the person in charge information and business development information based on the associated flag information.
[0030] Determine the alarm level based on the person in charge information and the business development information, and perform alarm push according to the alarm level.
[0031] Optionally, the cloud node further includes a dashboard node, and the dashboard node includes a cluster node composed of a recent view page and a long-term data query page. The method further includes:
[0032] Divide the acquisition classification information into target cycle data and complete data.
[0033] Update the target cycle data to the recent view page of the dashboard node.
[0034] Update the complete data to the long-term data query page of the dashboard node.
[0035] Optionally, the cloud node further includes a plurality of storage nodes. The method further includes:
[0036] Determine the corresponding relationship between the storage nodes and each cloud node.
[0037] Correspondingly store the cloud configuration information of each cloud node in each storage node based on the corresponding relationship, and perform backup storage through the target cluster structure in the storage node.
[0038] In addition, to achieve the above object, the present application also proposes a monitoring and warning system based on cross-region hybrid cloud. The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the monitoring and warning method based on cross-region hybrid cloud as described above.
[0039] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the monitoring and alarm method based on the cross-regional hybrid cloud as described above are implemented.
[0040] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the monitoring and alarm method based on a cross-region hybrid cloud as described above.
[0041] One or more technical solutions proposed in this application have at least the following technical effects:
[0042] This application classifies and defines the basic information of each cloud node through the cloud configuration information collected by the monitoring components corresponding to each node, and obtains the collected classification information; based on the collected classification information, label information and associated flag information are sent to each cloud node, so as to push alarms to the corresponding business line group through the automatic detection node of the cloud node. In this way, in the monitoring and alarm system based on the cross-regional hybrid cloud under the cross-regional hybrid cloud architecture, it is possible to monitor the configuration and basic information of different cloud nodes, and the business system can be linked to alarm and notify when an alarm is triggered. It solves the problems of large data volume and scattered data under the management of cross-regional hybrid clouds, as well as the high learning entry threshold and complex maintenance, which is convenient for users and maintainers to get started. At the same time, the architecture is simple, the construction cost is low, and it is maintainable and scalable. Finally, the unified management of basic data, automatic classification and monitoring of data, and accurate push of alarms to people can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0045] Figure 1 A schematic diagram of the structure provided for the monitoring and alarm system based on a cross-region hybrid cloud for this application;
[0046] Figure 2 A flow chart of the first embodiment of the monitoring and alarm method based on cross-region hybrid cloud of the present application;
[0047] Figure 3 This is a schematic diagram of system deployment in an embodiment of the monitoring and alarm method based on cross-region hybrid cloud of the present application;
[0048] Figure 4 This is a schematic flow diagram provided by Embodiment 2 of the monitoring and alarm method based on cross-region hybrid cloud of the present application;
[0049] Figure 5 This is a schematic flow diagram provided by Embodiment 3 of the monitoring and alarm method based on cross-region hybrid cloud of the present application;
[0050] Figure 6 This is a schematic diagram of the implementation process in an embodiment of the monitoring and alarm method based on cross-region hybrid cloud of the present application.
[0051] The realization of the purpose, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments
[0052] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0053] In order to better understand the technical solutions of the present application, the following will be described in detail with reference to the accompanying drawings of the specification and specific embodiments.
[0054] The main solution of the embodiment of the present application is that the monitoring component collects the cloud configuration information of each cloud node; classifies and defines the basic information of each cloud node based on the cloud configuration information to obtain the collected classification information; and issues label information and associated flag information to each cloud node based on the collected classification information, so as to push alarms to the corresponding business line groups through the automatic detection nodes of the cloud nodes.
[0055] In this embodiment, for the convenience of description, the following will be described with the monitoring and alarm system based on cross-region hybrid cloud as the execution subject.
[0056] Since there are two common practices for data monitoring and data alarm in the prior art: 1. Uniformly collect the data of each region and each cloud architecture and summarize them into a database or platform, and then separately configure the specified business line groups to send alarms and monitor the data. Because the data interaction and transmission are very low, the data is scattered, and the setup is simple, it is often applied to small-scale monitoring teams with small data volumes and has a low entry threshold, such as the early zabbix monitoring design. 2. Separate management and monitoring of data on each region and each cloud respectively to achieve a self-closed loop of monitoring and alarm. This involves data collaboration and management of multiple teams. Due to many rules, high management costs, large data volumes, complex setup, high interaction and transmission costs, and complex links, it is often used for large-scale monitoring teams with large monitoring data volumes and has a high entry threshold.
[0057] This application provides a solution. In this way, in a monitoring and alerting system based on cross-region hybrid cloud under a cross-region hybrid cloud architecture, it is possible to monitor the configurations and basic information of different cloud nodes, and when an alert is triggered, the business system can be linked for alerting and notification. This solves the problems of large data volume and data dispersion under cross-region hybrid cloud management, as well as the high learning threshold for getting started and the complexity of maintenance, making it convenient for users and maintainers to get started. At the same time, the architecture is simple, the setup cost is low, and it has good maintainability and scalability. Finally, it can achieve unified management of basic data, automatic classification and monitoring of data, and accurate pushing of alerts to people.
[0058] It should be noted that the execution entity of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a cloud system, etc. that can implement the above functions. Hereinafter, taking the monitoring and alerting system based on cross-region hybrid cloud as an example, this embodiment and the following embodiments will be described.
[0059] As Figure 1 shown is the structural schematic diagram of the monitoring and alerting system based on cross-region hybrid cloud in the embodiment of this application.
[0060] It should be noted that the system in this embodiment includes a data management center node and multiple cloud nodes connected to the data management center node. Each cloud node further includes a monitoring component, an alert node, and an automatic detection node.
[0061] The data management center node is used to classify and define the basic information of each cloud node to obtain the acquisition classification information;
[0062] The cloud node is used to collect data through the monitoring component, and perform automatic alerting based on the acquisition classification information through the automatic detection node and the alert node.
[0063] The embodiment of this application provides a monitoring and alerting method based on cross-region hybrid cloud. Refer to Figure 2 , Figure 2 which is the flowchart of the first embodiment of the monitoring and alerting method based on cross-region hybrid cloud of this application applied to the data management center node.
[0064] In this embodiment, the monitoring and alerting method based on cross-region hybrid cloud includes steps S10 to S30:
[0065] Step S10, classify and define the basic information of each cloud node through the cloud configuration information collected by the monitoring component corresponding to each node to obtain the acquisition classification information;
[0066] It should be noted that asFigure 3 The deployment diagram of the solution of this embodiment is shown. The core components of the monitoring and alerting system based on cross-region hybrid cloud in this embodiment include: Core component description: 1. Use open-source Prometheus as the basic monitoring component (replaceable). Under the current trend, the data of indicators is relatively simple for operation and maintenance and development personnel to get started. 2. The storage uses the InfluxDB time-series database and the highly scalable VictoriaMetrics time-series database. 3. The unified data management center node is built with Flask-Python and is used to read and store basic resource information, application association information, and personnel management information. 4. The webhook is built with Flask-Python and is used to distribute the data obtained from the data management center node for data processing and responsible person distribution.
[0067] It should be understood that none of the current monitoring data software on the market is associated with the actual business. The data tagging of this solution and CMDB realizes the combination of the data source and the business line, solves the problem of automatic classification monitoring, and can also accurately push the aggregation alert to the minimum of the responsible person. The realization of automatic classification also solves the problem of cumbersome configuration of alert rules. It manages the problems of large data volume and data dispersion under cross-region hybrid cloud, and performs secondary calculation after unified aggregation. Using storage and open-source components, its architecture is simple, the construction cost is low, and the maintainability and scalability are good. Adopting the solution of this embodiment can solve the pain points of large monitoring data volume, data dispersion, and the need to accurately configure each business line group for rules, reduce the entry construction cost, enable fast and effective alert push, reduce the management cost of monitoring, reduce the architecture complexity of the monitoring system, have low later expansion cost, and high compatibility. Thus, it can solve: 1. Solve the problems of large data volume and data dispersion under cross-region hybrid cloud management. A unified panel can be used. 2. Solve the problems of high learning entry threshold and complex maintenance. It is convenient for users and maintainers to get started. 3. Simple architecture, low construction cost, good maintainability and scalability. 4. Unified management of basic data, automatic classification monitoring of data, and accurate pushing of alerts to people. 5. Multiple channels for sending, improving the reliability of alert data sending.
[0068] In specific implementation, first, the configuration information and data in each cloud node are collected through the monitoring component, which is convenient for subsequent monitoring. Among them, the cloud configuration information includes, but is not limited to, data such as the IP of each cloud node. And each cloud node in the monitoring and alerting system based on cross-region hybrid cloud in this embodiment can be any type of cloud node, such as private cloud, public cloud, etc., and can also be cross-region hybrid cloud.
[0069] It should be noted that after obtaining the cloud configuration information, the basic information of each cloud resource collected is classified and defined for monitoring. The basic information refers to the basic data of each cloud node and the data uploaded to the data management center node.
[0070] In a feasible implementation, in order to accurately implement the classification definition of basic information, step S20 includes: obtaining the address information and type information of each cloud node collected by the monitoring component according to the monitoring rules and the target scripts stored in the time series databases arranged on each cloud node; performing classification definition based on the basic information and the cloud configuration information to obtain the acquisition classification information.
[0071] It should be understood that the basic information is collected through the APIs or shell scripts of each cloud. After obtaining the basic information, classification definition is performed based on the basic information, so that the acquisition classification information can be obtained.
[0072] In a feasible implementation, in order to accurately perform classification definition, the step of performing classification definition based on the basic information and the cloud configuration information to obtain the acquisition classification information includes: determining tag information according to the cloud configuration information; performing classification definition on the basic information based on the tag information to obtain the acquisition classification information.
[0073] In specific implementation, there is a unified data management center node. The basic information is collected through the APIs or shell scripts of each cloud, and then the unique ID, name (identifier), component type (such as mysq / redis / kafka / ecs / pg / consul / mycat / mq, etc.), domain, importance level, unique link address, region, cloud provider, environment, and monitoring address are classified and defined based on the cloud configuration information. The tags here are extensible. A read interface is provided according to signature authentication. At the same time, a unified alarm rule file transfer controller is also made. The data management center will also store the monitoring storage information of each cloud for the classified distribution of data source information.
[0074] It should be noted that the resources data of cloud providers is called through API encapsulation. When there is a resource change, data pulling is triggered, written to the database, and tagged. The information of the export nodes and ports is automatically allocated from the monitoring node pool, and timed pulling can also be set. The unique ID, name (identifier), component type (such as mysq / redis / kafka / ecs / pg / consul / mycat / mq, etc.), domain, importance level, unique link address, region, cloud provider, environment, and monitoring address in the relevant information are used as tags.
[0075] In step S20, based on the acquisition classification information, tag information and association flag information are sent to each cloud node, so as to perform alarm push to the corresponding business line group through the automatic detection nodes of the cloud node.
[0076] It should be noted that the high-availability mode (gossip) of the open-source alertmanager is adopted here for the alarm node to bear the concurrent alarm pressure.
[0077] In the specific implementation, after obtaining the collected classification information, the data of cloud nodes with different labels are monitored, and whether there is changed information is monitored. After a change occurs or abnormal node data is found, an alarm will be issued, and the alarm is specifically combined with the responsible persons and business development information in the business system.
[0078] This embodiment provides a monitoring and alarming method based on a cross-region hybrid cloud. The basic information of each cloud node is classified and defined through the cloud configuration information collected by the monitoring components corresponding to each node to obtain the collected classification information; based on the collected classification information, label information and associated flag information are sent to each cloud node, so as to push alarm messages to the corresponding business line group through the automatic detection nodes of the cloud node. In this way, in the monitoring and alarming system based on the cross-region hybrid cloud under the cross-region hybrid cloud architecture, the configuration and basic information of different cloud nodes can be monitored, and when an alarm is triggered, the business system can be linked for alarm and notification. It solves the problems of large data volume and data dispersion under cross-region hybrid cloud management, as well as high learning threshold and complex maintenance, and is convenient for users and maintainers to get started. At the same time, the architecture is simple, the construction cost is low, and it has maintainability and scalability. Finally, the unified management of basic data, automatic classification monitoring of data, and accurate alarm pushing to people can be realized.
[0079] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as the above-mentioned embodiment 1 can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 4 , step S10 includes steps S101 to S102:
[0080] Step S101, call the target script stored in the time series database arranged on each cloud node, and obtain the address information and type information of each cloud node collected by the monitoring component according to the monitoring rules and the target script stored in the time series database arranged on each cloud node;
[0081] It should be noted that the open-source Prometheus is adopted as the basic monitoring component (replaceable) as the monitoring and collection aggregation component.
[0082] It should be understood that since different types of data sources are classified, multiple nodes of Prometheus can be extended, and a python script is stored on each Prometheus node.
[0083] In a specific implementation, Python is used to request data from the data center by taking the local IP and type as parameters, and then obtain cloud configuration information based on the local IP and type.
[0084] Step S102: Obtain the cloud configuration information of each cloud node based on the address information and the type information.
[0085] In a specific implementation, after obtaining the address information and the type information, obtain the cloud configuration information based on the local IP and type.
[0086] In a feasible implementation manner, in order to facilitate the generation and use of monitoring rules, before the step of obtaining the address information and the type information of each cloud node according to the monitoring rules by calling the target script stored in the time series database arranged on each cloud node, it further includes: generating and storing monitoring rules, and sending the monitoring rules to the specified paths corresponding to the monitoring components on each cloud node.
[0087] It should be noted that the monitoring rules uniformly go through the data center. The storage directory rule is: / storage directory / data type / ; the file name rule is: Id-timestamp-rule name.rules. Then, use the data center to synchronize and distribute and copy files to the specified paths of the corresponding Prometheus using the scp copy command (see the rule matching address in Architecture Description 3: / address where the rules are stored / *.rules). It is also possible to set the rsync host's built-in synchronization tool to synchronize the entire directory rule (because the data sources are different, synchronizing the entire rule will not cause alarm chaos).
[0088] In a feasible implementation manner, in order to accurately obtain the cloud configuration information, the step of obtaining the cloud configuration information of each cloud node based on the address information and the type information includes: determining the data source address information and the configuration file information of each cloud node based on the address information and the type information; obtaining the cloud configuration information of each cloud node according to the data source address information and the configuration file information.
[0089] It should be understood that the label information is obtained by calling the interface for obtaining the data monitoring address attribute of the data center through the API, and the Prometheus source data acquisition configuration is obtained according to the type (configuration general matching address: / storage address / type / *json) and written into the json data. Synchronously set it as a scheduled task (executed once every 5 minutes). Write the data returned by the data center into a json file, where the returned monitoring address is the data source address pulled by the Prometheus component. If different regions are involved, the obtained data source address is the data source address of the corresponding region. The configuration general matching address in the configuration file is: / storage address / type / *json, and the rule matching general matching address is: / rule storage address / *.rules. In this way, the component characteristics can be used to achieve dynamic real-time automatic monitoring refresh without changing the configuration and reloading. Enable remote writing and then double-write to the influxdb and victoriametrics storages.
[0090] In this embodiment, the address information and type information of each cloud node are obtained by calling the target script stored in the time series database arranged on each cloud node according to the monitoring rules; based on the address information and the type information, the cloud configuration information of each cloud node is obtained. In this way, the acquisition of the address and type is realized on each cloud node through an automatic script, which facilitates subsequent data labeling and classification management.
[0091] It should be noted that the above examples are only for understanding the present application and do not constitute a limitation on the monitoring and warning method of the present application based on cross-regional hybrid clouds. Based on this technical concept, more forms of simple transformations are within the protection scope of the present application.
[0092] The embodiment of the present application also provides a monitoring and warning method based on cross-regional hybrid clouds, referring to Figure 5 , Figure 5 which is a schematic flowchart of the first embodiment of the monitoring and warning method of the present application based on cross-regional hybrid clouds.
[0093] This embodiment is applied to cloud nodes and includes steps S01 to S04.
[0094] Step S01: Perform activity detection through the automatic detection nodes of each cloud node;
[0095] It should be noted that first, activity detection is performed through the automatic detection nodes in the cloud nodes to determine the abnormal nodes and data. The automatic detection node auto-webhook will be deployed on each cloud for alarm distribution. Among them, it will perform liveness detection with the auto-webhook node in the cluster where it is located.
[0096] Step S02, request associated flag information from the data management center node according to the tag information sent by the data management center node and the detection result;
[0097] It should be understood that since the alarm configuration is a general alarm, the tags of the data management center will be obtained here, and then the data management center will be requested.
[0098] Step S03, determine the person in charge information and business development information based on the associated flag information;
[0099] In a specific implementation, the application responsible persons are associated with different IDs to obtain the corresponding person in charge information and business development information.
[0100] Step S04, determine the alarm level based on the person in charge information and the business development information, and perform alarm push through the automatic detection node of the corresponding cloud node according to the alarm level.
[0101] It should be noted that as Figure 6 shown in the implementation flowchart of the solution of this embodiment, alarm push is performed according to the alarm level, supporting multiple channels. In order to reduce the traffic transmission of pulling data, the associated information is cached locally, and the data validity is one day, and this value can be adjusted in the configuration.
[0102] It should be understood that when the alarm is triggered, the tags passed through the data center are recognized, and the data center is requested according to the type and unique ID. The obtained returned person in charge information and the information on the APP logic side associated therewith are combined, so that the information of all personnel responsible for this component and its surrounding components can be pushed and notified in place. Among them, the alarm rule is a general matching rule.
[0103] In a specific implementation, Auto-webhook itself will also communicate with other nodes. If it is recognized that the local sending fails, a request for forwarding will be initiated to other nodes. If it is detected that a certain node is out of contact, an out-of-contact alarm will be issued. The purpose is for high availability and effective sending of alarm data.
[0104] In a feasible implementation manner, in order to realize the distributed storage of data, the monitoring and alarm system based on the cross-region hybrid cloud further includes a storage node, and the method further includes: determining the corresponding relationship between the storage node and each cloud node; storing the cloud configuration information of each cloud node correspondingly in each storage node based on the corresponding relationship.
[0105] It should be noted that InfluxDB is used for storage because Prometheus has read and write characteristics for it, and the open-source version does not have high availability and has a poor compression ratio. Therefore, VictoriaMetrics cluster mode is used for backup storage. It has high availability and high scalability, supports data interaction of more than one million levels, is friendly to backup data storage, and is suitable for scenarios with high read requirements. Here, vmselect can be made into multiple nodes, which is beneficial to scenarios with high-concurrency read requirements. Vminsert is suitable for high-concurrency write scenarios. Here, it can be used to supply inspection scripts for data calculation. Vmstorage is based on file storage, has the characteristic of high data compression ratio, and has good scalability.
[0106] In a feasible implementation manner, in order to achieve node-based storage of data, the monitoring and alerting system based on cross-region hybrid cloud further includes a dashboard node. The method further includes: dividing the collected classification information into target-period data and complete data; updating the target-period data to the recent view page of the dashboard node; and updating the complete data to the long-term data query page of the dashboard node.
[0107] It should be understood that Grafana and vminsert are used for the dashboard. The former is used for the recent view page of the dashboard for viewing recent time, and the latter is used as the query page for long-term data, that is, the long-term data query page. In this way, the real-time performance and scalability of reading in the cluster architecture are better, and the pressure brought by long-term large-volume data reading to the architecture is shared. Other clouds are also deployed separately.
[0108] In this embodiment, activity detection is performed through the automatic detection nodes of each cloud node; association flag information is requested from the data management center node according to the label information and the detection result issued by the data management center node; the person in charge information and business development information are determined based on the association flag information; the alert level is determined based on the person in charge information and the business development information, and alert push is performed through the automatic detection nodes of the corresponding cloud node according to the alert level. In this way, the linkage between data detection and alerting and the business system is realized, and direct distribution and confirmation of the corresponding responsible person and business system can be achieved for accurate alerting.
[0109] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by a processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.
[0110] The monitoring and alerting system based on cross-region hybrid cloud provided by the present application adopts the monitoring and alerting method based on cross-region hybrid cloud in the above embodiments, and can solve the technical problem that it is impossible to link services for alerting due to a large amount of monitoring data of cloud nodes. Compared with the prior art, the beneficial effects of the monitoring and alerting system based on cross-region hybrid cloud provided by the present application are the same as those of the monitoring and alerting method based on cross-region hybrid cloud provided in the above embodiments, and other technical features in the monitoring and alerting system based on cross-region hybrid cloud are the same as the features disclosed in the method of the previous embodiment, which will not be elaborated here.
[0111] It should be understood that each part disclosed in the present application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0112] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0113] The present application provides a computer-readable storage medium, having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the monitoring and alerting method based on cross-region hybrid cloud in the above embodiments.
[0114] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0115] The above computer-readable storage medium can be included in a cross-regional hybrid cloud-based monitoring and alerting system; or it can exist independently without being assembled into a cross-regional hybrid cloud-based monitoring and alerting system.
[0116] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by a cross-regional hybrid cloud-based monitoring and alerting system, the cross-regional hybrid cloud-based monitoring and alerting system collects the cloud configuration information of each cloud node through the monitoring component; classifies and defines the basic information of each cloud node based on the cloud configuration information to obtain the collected classification information; and sends the label information and associated flag information to each cloud node based on the collected classification information, so as to push alerts to the corresponding business line groups through the automatic detection nodes of the cloud nodes.
[0117] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0118] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0119] The modules involved in the embodiments described in this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0120] The readable storage medium provided in this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned monitoring and alarming method based on cross-region hybrid cloud, and can solve the technical problem that it is impossible to link services for alarming due to a large amount of cloud node monitoring data. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the monitoring and alarming method based on cross-region hybrid cloud provided in the above embodiments, and will not be elaborated here.
[0121] The present application also provides a computer program product, including a computer program, which when executed by a processor implements the steps of the above-mentioned monitoring and alarming method based on a cross-region hybrid cloud.
[0122] The computer program product provided by the present application can solve the technical problem that it is impossible to link services for alarming due to a large amount of monitoring data of cloud nodes. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the monitoring and alarming method based on a cross-region hybrid cloud provided by the above-mentioned embodiments, and will not be elaborated here.
[0123] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A monitoring and alarming system based on cross - regional hybrid cloud, the monitoring and alarming system based on cross - regional hybrid cloud includes a data management center node and multiple cloud nodes. Monitoring components, alarming nodes and automatic detection nodes are deployed on each cloud node. Each cloud node is at least one type of cloud server deployed in multiple regions, and the data management center node is connected to each cloud node respectively; The data management center node is used for classifying and defining the basic information of each cloud node to obtain collection classification information; The cloud node is used for collecting data through the monitoring component, and performing automatic alarming through the automatic detection node and the alarming node based on the collection classification information.
2. A monitoring and alarming method based on cross-region hybrid cloud, characterized in that The method is applied to the data management center node in the monitoring and alarming system based on cross - regional hybrid cloud as claimed in claim 1; The method includes: Classifying and defining the basic information of each cloud node through the cloud configuration information collected by the monitoring component corresponding to each node to obtain collection classification information; Based on the collection classification information, sending label information and associated flag information to each cloud node, so as to push alarm to the corresponding business line group through the automatic detection node of the cloud node.
3. The method according to claim 2, wherein Before the step of classifying and defining the basic information of each cloud node through the cloud configuration information collected by the monitoring component corresponding to each node to obtain collection classification information, it further includes: Obtaining the address information and type information of each cloud node collected by the monitoring component according to the monitoring rules and the target script stored in the time - series database arranged on each cloud node; Obtaining the cloud configuration information of each cloud node based on the address information and the type information.
4. The method according to claim 3, wherein Before the step of obtaining the address information and type information of each cloud node by calling the target script stored in the time - series database arranged on each cloud node according to the monitoring rules, it further includes: Generating and storing monitoring rules, and sending the monitoring rules to the specified path corresponding to the monitoring component on each cloud node.
5. The method according to claim 3, characterized in that, The step of obtaining the cloud configuration information of each cloud node based on the address information and the type information includes: Determining the data source address information and configuration file information of each cloud node based on the address information and the type information; Obtaining the cloud configuration information of each cloud node according to the data source address information and the configuration file information.
6. The method according to claim 2, wherein The step of classifying and defining the basic information of each cloud node based on the cloud configuration information to obtain collection classification information includes: Calling the application programming interface or automatic script corresponding to each cloud node to collect the basic information of each cloud node; Classifying and defining based on the basic information and the cloud configuration information to obtain collection classification information.
7. The method according to claim 6, wherein The step of classifying and defining based on the basic information and the cloud configuration information to obtain collection classification information includes: Determining label information according to the cloud configuration information; Classifying and defining the basic information based on the label information to obtain collection classification information.
8. A monitoring and alarming method based on cross-region hybrid cloud, characterized in that, The method is applied to the cloud node in the monitoring and alarming system based on cross - regional hybrid cloud as claimed in claim 1; The method includes: Performing activity detection through the automatic detection node of each cloud node; Request the associated flag information from the data management center node according to the tag information sent by the data management center node and the detection result; Determine the person-in-charge information and business development information based on the associated flag information; Determine the alarm level based on the person-in-charge information and the business development information, and perform alarm push through the automatic detection node of the corresponding cloud node according to the alarm level.
9. The method according to claim 8, wherein The cloud node further includes a dashboard node, and the dashboard node includes a cluster node composed of a recent view page and a long-term data query page. The method further includes: Divide the collected classification information into target cycle data and complete data; Update the target cycle data to the recent view page of the dashboard node; Update the complete data to the long-term data query page of the dashboard node.
10. The method according to claim 8, characterized in that, The cloud node further includes a plurality of storage nodes. The method further includes: Determine the corresponding relationship between the storage nodes and each cloud node; Correspondingly store the cloud configuration information of each cloud node in each storage node based on the corresponding relationship, and perform backup storage through the target cluster structure in the storage node.