Edge Cluster Monitoring Method and System

The described method and system address the complexity of edge cluster monitoring by dynamically creating and deploying tailored monitoring components, reducing network traffic and ensuring secure, efficient monitoring of diverse edge clusters.

CN113986662BActive Publication Date: 2025-07-15JINAN INSPUR DATA TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111231265.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-07-15
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

The prior art is difficult to achieve flexible, real-time and global monitoring of edge clusters, especially when there are a wide variety of equipment and loads, monitoring configurations are difficult to flexibly distribute, and there is a problem of high monitoring costs.

Method used

Dynamically obtain the status information of the edge cluster through the central cluster, identify changing loads or devices, create targeted monitoring component configuration information, and send it to the edge cluster through the envoy gateway. The edge cluster deploys monitoring components based on the configuration information, and combines a distributed time series database and monitoring alarm components to realize the intelligent deployment and data collection of monitoring components.

Benefits of technology

It realizes targeted monitoring component deployment for different edge clusters, reduces network transmission, prevents the deployment of useless components, and separates policy formulation and execution, ensuring the security of the central cluster and the real-time and global monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113986662B_ABST
    Figure CN113986662B_ABST
Patent Text Reader

Abstract

This application relates to an edge cluster monitoring method and system; the method includes: the central cluster dynamically obtains the status information of all edge clusters; when it is found that the load or device has changed, obtains the characteristic information of the target edge cluster; the target edge cluster is the edge cluster where the changed load or device is located; creates the configuration information of the monitoring component for the target edge cluster according to the characteristic information; sends the configuration information to the target edge cluster; after receiving the configuration information, the target edge cluster deploys the corresponding monitoring component according to the configuration information to perform data collection and monitoring. The solution of this application realizes the targeted deployment of monitoring components for different edge clusters, reduces network transmission, prevents the deployment of useless monitoring components, and realizes the separation of policy formulation and policy execution. It not only solves the monitoring of edge clusters with a large variety of loads and devices, but also ensures the security of the central cluster.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cloud computing technology, and in particular to an edge cluster monitoring method and system. Background Art

[0002] With the advent of the 5G and Internet of Things era, sinking the capabilities of cloud computing to the edge side and managing them through the cloud center will be an important development trend of cloud computing. An edge cluster is a small-scale cloud data center distributed on the edge side of the network, providing real-time data processing and analysis and decision-making.

[0003] Monitoring is an important means to ensure the availability of edge services. Efficient and comprehensive monitoring coverage is the cornerstone and guarantee for the reliable operation of edge services. It not only needs to support the monitoring of the performance of edge clusters and edge nodes, but also needs to collect service data of edge clusters and edge nodes. In addition, it is necessary to monitor the information of edge devices.

[0004] In the related art, there are many types of workloads in edge clusters, many types of edge devices, and the number and scale of clusters are also constantly expanding. These situations bring great difficulties to the monitoring of the entire cluster. How to ensure the flexible distribution of monitoring configurations, how to ensure the real-time and global nature of monitoring data, and how to reduce the cost and difficulty of cluster monitoring have become technical challenges faced by edge cluster monitoring. Summary of the Invention

[0005] To overcome at least to some extent the problems existing in the related art, this application provides an edge cluster monitoring method and system.

[0006] According to the first aspect of the embodiments of this application, an edge cluster monitoring method is provided, including:

[0007] The central cluster dynamically obtains the status information of all edge clusters;

[0008] When it is found that the load or device has changed, obtain the characteristic information of the target edge cluster; the target edge cluster is the edge cluster where the changed load or device is located;

[0009] Create configuration information of monitoring components for the target edge cluster according to the characteristic information;

[0010] Send the configuration information to the target edge cluster;

[0011] After receiving the configuration information, the target edge cluster deploys corresponding monitoring components according to the configuration information for data collection and monitoring.

[0012] Further, the method further includes:

[0013] The central cluster creates or modifies the scheduling policy for edge monitoring through a configurator, and the scheduling policy includes a monitoring installation package and a distribution policy.

[0014] Furthermore, the central cluster dynamically obtains the status information of all edge clusters, including:

[0015] The resource analyzer periodically polls the monitoring data of all edge clusters in the database;

[0016] Compares the polled data with historical data to determine the changed loads or devices;

[0017] Provides an interface for the controller to poll.

[0018] Furthermore, the creation of the configuration information of the monitoring component for the target edge cluster according to the feature information includes:

[0019] The controller formulates a monitoring policy for the changed loads or devices according to the scheduling policy;

[0020] The rule converter converts the monitoring policy into configuration information according to the feature information of the target edge cluster;

[0021] Among them, the feature information is the operating system of the edge cluster or the type and version of the cloud products of the edge cluster; the configuration information includes an execution method and an execution policy.

[0022] Furthermore, the distribution of the configuration information to the target edge cluster includes:

[0023] The configuration information is distributed to the target edge cluster in the form of an API through the envoy gateway.

[0024] Furthermore, the deployment of the corresponding monitoring component according to the configuration information includes:

[0025] After receiving the configuration information sent by the central cluster, the edge control node uses the execution policy as an execution parameter according to the operation type and execution method to perform the deployment operation.

[0026] Furthermore, the method further includes:

[0027] After the edge control node performs the deployment operation, it distributes the monitoring component to the edge node where the load or device is located, so that the edge node installs and deploys the corresponding monitoring component.

[0028] Furthermore, each edge cluster is configured with a set of time series databases and monitoring and alerting components, and independently completes the alerting task;

[0029] The central cluster uses a distributed time series database to receive the monitoring data of each edge cluster.

[0030] Furthermore, the edge cluster uses the proxy component for data distribution and controls the data sent to the edge cluster and the data of the central cluster according to the formulated distribution strategy.

[0031] According to the second aspect of the embodiments of the present application, an edge cluster monitoring system is provided, including: a central cluster and several edge clusters;

[0032] The central cluster is used for: dynamically obtaining the status information of all edge clusters; when it is found that the load or device changes, obtaining the feature information of the target edge cluster; the target edge cluster is the edge cluster where the changed load or device is located; creating the configuration information of the monitoring component for the target edge cluster according to the feature information; and sending the configuration information to the target edge cluster;

[0033] The edge cluster is used for: after receiving the configuration information, deploying the corresponding monitoring component according to the configuration information for data collection and monitoring.

[0034] The technical solutions provided by the embodiments of the present application have the following beneficial effects:

[0035] The solution of the present application realizes the targeted deployment of monitoring components for different edge clusters, reduces network transmission, prevents the deployment of useless monitoring components, and realizes the separation of policy formulation and policy execution. It not only solves the monitoring of edge clusters with a large variety of loads and devices, but also ensures the security of the central cluster.

[0036] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0038] Figure 1 It is a schematic structural diagram of an edge cluster monitoring system in an embodiment of the present invention.

[0039] Figure 2 It is a flowchart of an edge cluster monitoring method in an embodiment of the present invention.

[0040] Figure 3 It is a schematic diagram of policy formulation and policy distribution of the control plane central cluster in an embodiment of the present invention.

[0041] Figure 4 It is a schematic diagram of the execution of the edge cluster policy of the data plane in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of methods and systems consistent with some aspects of the present application as detailed in the appended claims.

[0043] The edge cluster monitoring method provided by the present application can be applied to a system as shown in Figure 1 The system includes a central cluster and several edge clusters. The present application creatively proposes the concept of a monitoring grid, which can well complete the monitoring of edge clusters. The grid divides services, without adding new functions, but extracts the service - to - service communication for logical management from each service and abstracts it into an infrastructure layer. The present application is applicable to large - scale edge clusters with a large number of devices or loads and can achieve good results; however, for small - scale clusters, this method is too complex and has no advantages.

[0044] Figure 2 is a flowchart of an edge cluster monitoring method shown according to an exemplary embodiment. The method may include the following steps:

[0045] Step S1: The central cluster dynamically obtains the status information of all edge clusters;

[0046] Step S2: When it is found that the load or device has changed, obtain the characteristic information of the target edge cluster; the target edge cluster is the edge cluster where the changed load or device is located;

[0047] Step S3: Create the configuration information of the monitoring component for the target edge cluster according to the characteristic information;

[0048] Step S4: Send the configuration information to the target edge cluster;

[0049] Step S5: After the target edge cluster receives the configuration information, deploy the corresponding monitoring component according to the configuration information to perform data collection and monitoring.

[0050] The solution of the present application realizes the targeted deployment of monitoring components for different edge clusters, reduces network transmission, prevents the deployment of useless monitoring components, and realizes the separation of policy formulation and policy execution. It not only solves the monitoring of edge clusters with a large variety of loads and devices, but also ensures the security of the central cluster.

[0051] It should be understood that although Figure 2The steps in the flowchart are shown in sequence according to the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 2 at least some of the steps in Figure 2 may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least some of the sub-steps or stages of other steps or other steps.

[0052] The solution of this application will be further described below in combination with specific application scenarios.

[0053] The monitoring grid of this solution consists of two parts: the data plane and the control plane. The control plane is responsible for managing and configuring monitoring policies, discovering edge devices and edge loads, and formulating and distributing deployment policies. The data plane relies on the authentication, policy execution, and application deployment of the edge cluster itself to achieve targeted deployment of monitoring components. The solution of this application realizes the separation of policy formulation and policy execution, which not only solves the monitoring of edge clusters with a large variety of loads and devices, but also ensures the security of the central cluster.

[0054] (1) Control plane of monitoring: Configuration of policies, policy formulation, and policy distribution

[0055] The cloud center with control management as the main goal serves as the control plane of the solution. It can dynamically perceive the cluster itself, the types of workloads of the cluster, and changes in various monitoring resources according to monitoring metrics. For example, when load A is deployed on an edge node, the edge metrics will obtain information such as the node where the load is located, resource occupancy, and load type.

[0056] After obtaining the information of load A, the control management center creates configuration information for monitoring components for each edge cluster specifically, distributes this configuration information, and then the specific edge cluster realizes the automated deployment of the monitoring system according to the configuration information, so as to achieve the automatic generation of monitoring policies to the intelligent deployment of monitoring components.

[0057] As Figure 3 shown, as the control plane of monitoring, the acquisition and distribution of configurations by the central cluster specifically include the following steps:

[0058] a. Create or modify the scheduling policy for edge monitoring through a configurator, including monitoring installation packages and distribution policies. The monitoring installation package can be a container image or a binary installation package. It mainly focuses on the correspondence between the installation package and the edge load or device. For example, when load A is deployed at the edge, A-exporter version 1.1 is used for monitoring. The distribution policy mainly includes the distribution method and target of the image. For example, A-exporter is pulled into the image repository of the edge cluster in the dockerpull manner.

[0059] b. The resource analyzer periodically polls the data in the distributed time series database (monitoring data of all edge clusters), compares it with historical data, and obtains newly added or deleted edge loads or edge devices, and provides an interface for the controller to poll.

[0060] c. Based on the list / watch technology, the controller calls the resource analyzer interface. When a load or device is newly added or deleted, and according to the monitoring scheduling policy, specific monitoring policies are formulated for the changed load or device, which can be stored in the form of an xml file.

[0061] d. The rule converter changes the monitoring policies formulated by the controller into specific execution methods and execution policies according to the edge cluster operating system or the types and versions of cloud products in the edge cluster. For example, the execution method for binary packages is bash, the execution method for containers is dockerrun, and the execution method for k8s clusters is kubectl. The execution policy can be stored in the form of an xml file.

[0062] e. Through the envoy gateway, the xml file, operation type, execution method, etc. are sent to the corresponding edge cluster in the form of an API. The choice of the envoy gateway method mainly supports two-way SSL authentication and role-based access control (RBAC) to provide a security solution from the central cluster to the edge cluster.

[0063] (2) Data plane for monitoring: Execution of the deployment of monitoring components

[0064] All edge clusters form the data plane of the solution and are the actual executors of the policy. The data plane continuously listens for information sent from the control plane, including the type of load, the location of the load, application information of the load, deployment configuration of application monitoring, etc.

[0065] After receiving the configuration information sent from the control plane, the data plane pulls the image or installation package according to the configuration policy and deploys the corresponding monitoring components to achieve the purpose of data collection. When a certain edge cluster receives the API, it will execute the installation operation of the monitoring client, as shown in a certain edge cluster Figure 4 as shown.

[0066] a. After the edge control node receives the API and xml file sent by the central cluster, it performs the execution with the xml as the execution parameter according to the operation type and execution method in the API. For example: the operation type is installation and the execution method is bash.

[0067] b. After the edge control node executes the deployment operation, the installation and deployment execution will be sent to the load or the edge node where the device is located through SSH (Secure Shell Protocol) or the container's own mechanism, and the edge node will perform the installation and deployment.

[0068] (3) Storage and processing of monitoring data

[0069] As Figure 1 shown, the main processes of distributed data storage and processing are as follows:

[0070] a. Each edge cluster has its own set of TSDB (Time-Series Database) and monitoring components, which independently complete the alarm tasks spontaneously to ensure the real-time nature of alarms. At the same time, the entire data plane uses distributed TSDB as the persistent storage to ensure the comprehensiveness and globality of data. These data can also be shared with the cloud center for the control plane to achieve dynamic perception. The time-series database is mainly used to process data with time tags (changing in the order of time, that is, time serialization). Data with time tags is also called time-series data; it is a new type of non-relational database and has great significance in the era of big data.

[0071] b. The central cluster uses distributed TSDB to receive the monitoring data of each edge cluster to ensure the comprehensiveness and globality of data. These data can also be shared with the cloud center for the control plane to achieve dynamic perception.

[0072] c. To facilitate the control of monitoring data, the monitoring data collected by the monitoring collection client is not directly sent back to the edge cluster TSDB and the central cluster TSDB, but a proxy component is added to the edge cluster as data distribution, and distribution policies can be formulated to control the data sent to the edge cluster TSDB and the distributed TSDB of the central cluster.

[0073] A method for monitoring edge clusters based on a monitoring grid in the present invention realizes targeted deployment of monitoring components for different edge clusters, reduces network transmission, prevents the deployment of useless monitoring components, and achieves the separation of policy formulation and policy execution. It not only solves the monitoring of edge clusters with a wide variety of loads and devices, but also ensures the security of the central cluster.

[0074] As Figure 1As shown in the figure, the present application further provides an edge cluster monitoring system, including: a central cluster and several edge clusters;

[0075] The central cluster is used for: dynamically obtaining the status information of all edge clusters; when it is found that the load or device has changed, obtaining the characteristic information of the target edge cluster; the target edge cluster is the edge cluster where the changed load or device is located; creating the configuration information of the monitoring component for the target edge cluster according to the characteristic information; and sending the configuration information to the target edge cluster;

[0076] The edge cluster is used for: after receiving the configuration information, deploying the corresponding monitoring component according to the configuration information to perform data collection and monitoring.

[0077] Regarding the system in the above embodiments, the specific steps for each part to perform operations have been described in detail in the embodiments related to the method, and will not be elaborated here. Each part in the above edge cluster monitoring system can be implemented in whole or in part by software, hardware, and their combination. Each of the above parts can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above parts.

[0078] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and the content not detailed in some embodiments can be referred to the same or similar content in other embodiments.

[0079] It should be noted that in the description of the present application, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality" refers to at least two.

[0080] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a manner that is not shown or discussed in sequence, including in a substantially simultaneous manner according to the functions involved or in a reverse order, which should be understood by those skilled in the art of the embodiments of the present application.

[0081] It should be understood that each part of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logic functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0082] Those of ordinary skill in the art can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0083] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0084] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.

[0085] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0086] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limitations to the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

Claims

1. An edge cluster monitoring method, characterized in that, Including: Using the central cluster as the control plane for monitoring, the central cluster dynamically obtains the status information of all edge clusters, and dynamically senses the changes of the cluster itself, the workload types of the cluster, and various monitoring resources according to the monitoring metrics; When it is found that the load or device has changed, obtain the characteristic information of the target edge cluster; the target edge cluster is the edge cluster where the changed load or device is located; Create the configuration information of the monitoring component for the target edge cluster according to the characteristic information, including: the controller formulates a monitoring strategy for the changed load or device according to the scheduling strategy, and the rule converter converts the monitoring strategy into configuration information according to the characteristic information of the target edge cluster; Send the configuration information to the target edge cluster; Using all edge clusters as the data plane, continuously listen to the information sent by the control plane. After the target edge cluster receives the configuration information, deploy the corresponding monitoring components according to the configuration information for data collection and monitoring, including: after the edge control node receives the configuration information sent by the central cluster, according to the operation type and execution method, use the execution strategy as the execution parameter to execute the deployment operation. After the edge control node executes the deployment operation, send the monitoring component to the edge node where the load or device is located, so that the edge node installs and deploys the corresponding monitoring component.

2. The method according to claim 1, characterized in that It also includes: The central cluster creates or modifies the scheduling strategy for edge monitoring through a configurator, and the scheduling strategy includes a monitoring installation package and a sending strategy.

3. The method according to claim 2, characterized in that, The central cluster dynamically obtains the status information of all edge clusters, including: The resource analyzer periodically polls the monitoring data of all edge clusters in the database; Compare the polled data with the historical data to determine the changed load or device; Provide an interface for the controller to poll.

4. The method according to claim 3, characterized in that, The creating the configuration information of the monitoring component for the target edge cluster according to the characteristic information includes: The characteristic information is the operating system of the edge cluster or the type and version of the cloud product of the edge cluster; the configuration information includes an execution method and an execution strategy.

5. The method according to claim 4, wherein The sending the configuration information to the target edge cluster includes: Send the configuration information to the target edge cluster in the form of an API through the envoy gateway.

6. The method according to any one of claims 1-5, characterized in that: Each edge cluster is configured with a set of time series databases and monitoring alarm components, and independently completes the alarm task; The central cluster uses a distributed time series database to receive the monitoring data of each edge cluster.

7. The method according to any one of claims 1-5, characterized in that: The edge cluster uses the proxy component for data distribution, and controls the data sent to the edge cluster and the data of the central cluster according to the formulated distribution strategy.

8. An edge cluster monitoring system for implementing the method according to any one of claims 1-7, characterized in that, Including: a central cluster and several edge clusters; The central cluster is used for: dynamically obtaining the status information of all edge clusters; when it is found that the load or device has changed, obtaining the characteristic information of the target edge cluster; the target edge cluster is the edge cluster where the changed load or device is located; creating the configuration information of the monitoring component for the target edge cluster according to the characteristic information; sending the configuration information to the target edge cluster; The edge cluster is used to: after receiving the configuration information, deploy corresponding monitoring components according to the configuration information for data collection and monitoring.

Citation Information

Patent Citations

  • Digital twinborn monitoring modeling system and modeling method for cloud-edge collaborative plant

    CN113255170A