Distributed monitoring architecture and monitoring method for cloud computing
By designing a distributed monitoring architecture in the inter-cloud computing environment, including monitoring sensors, collectors, distribution managers and regulatory agents, the problem of existing technologies not being able to provide elastic and scalable monitoring is solved, and flexible and scalable monitoring of inter-cloud computing system resources is achieved, and service quality is ensured.
Patent Information
- Application Number
- CN202510162914.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art cannot provide flexible and scalable monitoring mechanisms in inter-cloud computing environments, is difficult to adapt to the needs of large distributed systems, and is unable to be directly applied to multi-cloud alliance environments.
A distributed monitoring architecture for inter-cloud computing is designed, including monitoring sensors, collectors, distribution managers and regulatory agents. The architecture organizes monitoring components in a modular way, supports on-demand startup, and provides load balancing and fault-tolerant services.
It realizes flexible and scalable monitoring of resource information of inter-cloud computing system, avoids data concentration, provides load balancing and fault tolerance, and ensures the rational use of monitoring resources and service quality.
Smart Images

Figure CN119996455A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cloud computing and resource monitoring, and more specifically, to a distributed monitoring architecture and a monitoring method for cloud computing. Background Art
[0002] Cloud computing is based on collaboration between cloud service providers. It integrates the resources of multiple cloud service providers to provide users with powerful performance and sufficient services. The safe and stable operation of each cloud resource in the cloud is a basic condition for business development. It is necessary to monitor and observe the operation quality of cloud resources in real time to ensure safe operation.
[0003] Information such as the performance of cloud computing resources and the load generated by users can be reviewed, monitored, and managed through cloud monitoring tools. On the one hand, existing monitoring methods focus on more efficient and accurate monitoring of the physical and virtual resources of cloud infrastructure, but do not provide elastic and scalable mechanisms and are not suitable for large distributed systems; on the other hand, there are distributed monitoring methods for cloud computing, which can provide adaptive and scalable monitoring, but cannot be directly applied to the multi-cloud alliance environment of cloud computing. Summary of the invention
[0004] The inventors of this application have found through extensive research and practice that cloud computing is characterized by the sharing of a large number of physical and virtualized resources in a highly dynamic environment. Each cloud service provider may be equipped with hundreds of physical hosts and thousands of virtual machines, and each virtual machine is equipped with multiple physical and logical sensors to collect monitoring data. These components will eventually generate a large amount of network traffic, resulting in increased system processing and bandwidth overhead, and reducing the scalability of the architecture. At the same time, monitoring data needs to be delivered in a timely manner, so the monitoring system should adopt lightweight processing and communication solutions to limit additional overhead. In addition, new cloud service providers may join the alliance at any time in the cloud environment, and monitoring tools should keep monitoring tasks running in failure scenarios and be flexible enough to continuously monitor cloud resources.
[0005] Based on the above considerations, in view of the technical problems existing in the prior art, the present invention provides a distributed monitoring architecture and monitoring method for cloud computing, which can monitor resource information in a cloud computing system, organize monitoring components in a modular way, and support on-demand startup to provide dynamic monitoring capabilities. The present invention can avoid data concentration and provide load balancing and fault tolerance services.
[0006] To achieve the above-mentioned object, the first aspect of the present invention provides a distributed monitoring architecture for cloud computing, which is applied to a cloud computing system. The cloud computing system includes a physical machine, a virtual machine and a cloud service provider. The distributed monitoring architecture includes:
[0007] A monitoring sensor is located in each running physical machine and virtual machine in the cloud computing system, and is used to obtain monitoring information from specific resources on the physical machine and / or virtual machine, establish communication with the target collector, and send the obtained monitoring information to the target collector when receiving a request from the target collector;
[0008] The collector is located in the cloud service provider in the cloud computing system and is used to collect and process the monitoring information sent by the monitoring sensors;
[0009] The distribution manager is located in the cloud service provider in the cloud computing system and is used to allocate available collectors to a group of monitoring sensors. It is also responsible for acquiring, integrating and storing the monitoring information collected and processed by all collectors in the cloud service provider.
[0010] The monitoring agent is located outside the cloud service provider and is used to provide cloud users with access points to cloud monitoring data. It is also responsible for detecting abnormal events and generating alarm notifications.
[0011] In one embodiment, the monitoring sensor is further used to locate the distribution manager to obtain the target collector, pass the located distribution manager by using broadcast messaging, and accept the target collector provided by the distribution manager.
[0012] In one embodiment, the collector is pre-registered on the distribution manager, and the collector has a built-in data collection module and a virtual machine instantiation control module.
[0013] Wherein, the data collection module is used to collect monitoring information from a group of monitoring sensors;
[0014] The virtual machine instantiation control module is used to plan the virtual machine instantiation operation according to the monitoring information obtained by the data collection module.
[0015] In one embodiment, a receive-request hybrid communication mode is adopted between the monitoring sensor and the collector, in which the monitoring sensor periodically provides monitoring reports to the collector according to a determined threshold, and the collector requests the monitoring sensor to provide monitoring information within a determined time interval.
[0016] In one embodiment, the distribution manager has a built-in load distribution module and an information integration module.
[0017] Among them, the load distribution module is used to distribute the available collectors to a group of monitoring sensors using a load balancing strategy;
[0018] The information integration module is used to obtain, integrate and store the information collected and processed by all collectors in the cloud service provider. At the same time, according to the query requirements of the supervision agent, it sends general or specific monitoring information queries to the supervision agent.
[0019] In one embodiment, the load distribution module is used to distribute available collectors to a group of monitoring sensors using a load balancing strategy, including:
[0020] When a monitoring sensor asks the distribution manager for available target collectors, the distribution manager assigns a registered collector according to a specific load balancing policy; when a collector fails, the distribution manager redistributes the monitoring sensors associated with the failed collector to all available collectors.
[0021] In one embodiment, the supervisory agent has a built-in alarm module for notifying users of abnormal processes during infrastructure operations.
[0022] In one embodiment, the alarm module includes an alarm database and an alarm generator;
[0023] Among them, the alarm database is used to maintain a set of data describing the alarms issued by the fault detection mechanism;
[0024] The alert generator is used to generate alerts and send notifications to cloud users and cloud service providers.
[0025] Based on the same inventive concept, the second aspect of the present invention provides a monitoring method for cloud computing, which is implemented based on the distributed monitoring architecture for cloud computing described in the first aspect. The monitoring method includes:
[0026] Acquire monitoring information from specific resources on a physical machine and / or a virtual machine through a monitoring sensor, establish communication with a target collector, and send the acquired monitoring information to the target collector when receiving a request from the target collector;
[0027] Collect and process the monitoring information sent by the monitoring sensors through the collector;
[0028] The distribution manager allocates available collectors to a group of monitoring sensors and is responsible for acquiring, integrating and storing monitoring information collected and processed by all collectors in the cloud service provider.
[0029] The supervisory agent provides cloud users with an access point to cloud monitoring data and is responsible for detecting abnormal events and generating alarm notifications.
[0030] In one embodiment, the monitoring information sent by the monitoring sensor is collected and processed by a collector, including:
[0031] The collector collects monitoring information from a group of monitoring sensors through the data collection module, and analyzes the resource usage status of physical machines and virtual machines based on the monitoring information obtained by the data collection module in the collector through the virtual machine instantiation control module, and plans related virtual machine instantiation operations;
[0032] When the collector receives the information request from the allocation manager, it responds with the monitoring data from the data collection module and the status results of all physical machines and virtual machines analyzed by the virtual machine instantiation module.
[0033] Compared with the prior art, the advantages and beneficial technical effects of the present invention are as follows:
[0034] The present invention provides a distributed monitoring architecture for cloud computing, including: monitoring sensors, collectors, distribution managers, and supervision agents. The monitoring sensors are located in each running physical machine and virtual machine in the cloud computing system, the collectors and distribution managers are located in the cloud service provider in the cloud computing system, and the supervision agent is located outside the cloud service provider. The architecture can monitor resource information in the cloud computing system, organize monitoring components in a modular manner, and support on-demand startup to provide dynamic monitoring capabilities. The present invention can avoid data concentration and provide load balancing and fault tolerance services.
[0035] Furthermore, the collector has a built-in data collection module and a virtual machine instantiation control module, the distribution manager has a built-in load distribution module and an information integration module, and the supervision agent has a built-in user access interface and an alarm module, which can share monitoring tasks according to the load balancing strategy, and can avoid data concentration. The distributed monitoring component of the present invention also supports on-demand startup according to changes in the scale of the architecture infrastructure, so that the architecture has self-organization ability, elasticity, scalability and fault tolerance, while avoiding single point failures that cause monitoring data loss and monitoring function failure. At the same time, the present invention considers the service quality of cloud computing, integrates the virtual machine instantiation control method into the monitoring component, and is used to monitor the resource consumption of the virtual machine instance, maintain the resource utilization rate of the virtual machine instance, and can effectively avoid continuous saturation of resource status to prevent system congestion, response loss and even infrastructure damage, and also avoid the waste of monitoring resources caused by continuous idle resource status. The present invention can ensure that monitoring resources are used reasonably, thereby improving resource utilization efficiency and service quality. The present invention can be applied to provide monitoring services for cloud computing systems in different application scenarios, filling the gap in the field of distributed monitoring in the context of cloud computing in China today. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0037] Figure 1Schematic diagram of a distributed monitoring architecture for cloud computing in an embodiment of the present invention;
[0038] Figure 2 A schematic diagram of a virtual machine instantiation process in an embodiment of the present invention;
[0039] Figure 3 Schematic diagram of the alarm module flow in an embodiment of the present invention. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] Embodiment 1
[0042] This embodiment discloses a distributed monitoring architecture for cloud computing, which is applied to a cloud computing system. The cloud computing system includes physical machines, virtual machines, and cloud service providers. Figure 1 , the distributed monitoring architecture includes:
[0043] A monitoring sensor is located in each running physical machine and virtual machine in the cloud computing system, and is used to obtain monitoring information from specific resources on the physical machine and / or virtual machine, establish communication with the target collector, and send the obtained monitoring information to the target collector when receiving a request from the target collector;
[0044] The collector is located in the cloud service provider in the cloud computing system and is used to collect and process the monitoring information sent by the monitoring sensors;
[0045] The distribution manager is located in the cloud service provider in the cloud computing system and is used to allocate available collectors to a group of monitoring sensors. It is also responsible for acquiring, integrating and storing the monitoring information collected and processed by all collectors in the cloud service provider.
[0046] The monitoring agent is located outside the cloud service provider and is used to provide cloud users with access points to cloud monitoring data. It is also responsible for detecting abnormal events and generating alarm notifications.
[0047] During the specific implementation process, specific resources on the physical machine and / or virtual machine include CPU, memory, network, disk, etc. When receiving the request from the collector, the monitoring sensor sends the monitoring information to the corresponding collector, and the collector performs the monitoring information processing and storage tasks.
[0048] The distribution manager manages a set of distributed collectors that obtain information from a set of meta-sensors running on physical and virtual machines. The supervision agent is responsible for integrating monitoring information across the entire cloud environment and is also the user's access point to the cloud monitoring data. It provides a common API that allows users to read local monitoring data stored in each cloud infrastructure node. In this architecture, system administrators can start a set of distribution server instances associated with a set of collectors and meta-sensors to isolate different parts of each large cloud infrastructure. The distribution manager and collectors are implemented in each cloud instance of the cloud and can be dynamically started based on the scale of the cloud architecture and the monitoring service configuration.
[0049] The monitoring sensor provided by the present invention has the following specific functions and implementation process:
[0050] A101: Run monitoring sensors on physical and / or virtual machines in each service provider's infrastructure in the cloud.
[0051] A102: Get every monitoring metric you need from specific resources (CPU, memory, network, disk, etc.) on physical and / or virtual machines;
[0052] A103: Locate the distribution manager through broadcast messages and obtain the target collector;
[0053] A104: Establish communication with the target collector. Call a subset of specific sensors based on the monitoring indicators requested by the collector and send the monitoring information to the corresponding collector for processing and storage tasks.
[0054] The collector provided by the present invention has a built-in data collection module and a virtual machine instantiation control module, which are specifically implemented as follows:
[0055] B101: Register the collector on the distribution manager;
[0056] B102: Establish contact with corresponding monitoring sensors;
[0057] B103: collects monitoring information from a group of monitoring sensors through the data collection module, and communicates with the monitoring sensors in a receive-request hybrid mode;
[0058] B104: The virtual machine instantiation control module analyzes the resource usage status of the physical machine and virtual machine based on the monitoring values and available resources of the physical machine obtained by the data collection module in the collector, and plans the related virtual machine instantiation operations (allocation, suspension, and recovery). When the collector receives the information request from the allocation manager, it responds with the analysis results of all physical and virtual machine states.
[0059] During the virtual machine instantiation control process, the resource usage status of all physical and virtual machines monitored by the monitoring sensors managed by the collector will be analyzed. When the collector receives the information request from the distribution manager, it will respond with the analysis results of the status of all physical and virtual machines. In this way, the distribution manager can fully understand the status of the cloud service provider.
[0060] like Figure 2 As shown, step B104 is specifically as follows:
[0061] B201: Retrieve the list of running virtual machines;
[0062] B202: Obtain resource indicator information of each VM from the monitoring sensor. Each running VM needs to use a tuple (CPU, memory, hard disk, network, maximum / minimum threshold) to determine the resource usage status of the VM.
[0063] B203: Calculate the virtual machine resource status score VM , and its calculation formula is:
[0064]
[0065] CPU, Mem, Net, and Disk are the utilization rates of the virtual machine CPU, memory, network, and hard disk respectively;
[0066] B204: Compare the score with the upper threshold. If the score is less than the upper threshold, execute step B205; if the score is greater than the upper threshold, execute step B206;
[0067] B204: Compare the score with the upper threshold. If the score is less than the upper threshold, execute step B205; if the score is greater than the upper threshold, execute step B206;
[0068] B205: Compare the score with the lower threshold. If the score is less than the lower threshold, suspend the virtual machine. If the score is greater than the lower threshold, do not perform any operation on the virtual machine.
[0069] B206: Check whether the suspended virtual machine is available, and if there is a virtual machine that meets the conditions, perform a recovery operation on the virtual machine;
[0070] B207: Obtain resource indicator information of each physical machine from the monitoring sensor to determine whether the resource status of the physical machine is saturated. To accept the allocation of a new virtual machine instance, you must check the resource status of the physical machine, that is, whether the physical machine is saturated. The resource status of the physical machine is mainly based on available resources, namely memory, CPU, disk, and network. Therefore, there are four possibilities for judging the saturation of the physical machine: 1) CPU usage exceeds the upper threshold; 2) memory usage exceeds the previously determined threshold; 3) disk usage exceeds the previously determined threshold; 4) network usage exceeds the previously determined threshold. As long as the maximum usage of one of the indicators (memory, CPU, disk, and network) does not exceed the maximum threshold, it is determined that the physical machine has not reached the saturation state. Conversely, as long as the maximum usage of one indicator exceeds the threshold, the physical machine is determined to be saturated.
[0071] B208: Select an unsaturated physical machine to create a new virtual machine instance;
[0072] The distribution manager provided by the present invention has a built-in load distribution module and an information integration module, which are specifically implemented as follows:
[0073] C101: Receive registered collectors;
[0074] C102: When a monitoring sensor requests a collector to be assigned to the distribution manager, the load distribution module distributes the available collectors to a set of monitoring sensors (installed on physical and / or virtual machines). When a collector fails, the load distribution module redistributes the monitoring sensors associated with the failed collector to all available collectors, thereby maintaining a fair load balancing strategy.
[0075] C103: Acquire, integrate and store the information collected and processed by all collectors through the information integration module, that is, the current workload status information of the cloud service provider's infrastructure.
[0076] C104: Receive information query requests from regulatory agents, and send general or specific monitoring information queries to regulatory agents based on the regulatory agent's query requirements.
[0077] The present invention provides a built-in alarm module for a supervisory agent, which is specifically implemented as follows:
[0078] D101: Maintain a list of (user id, cloud service provider, cloud host) and provide a multi-function API as an access point for users to access monitoring data;
[0079] D102: When receiving a user's request to view monitoring information, according to the user ID, search for the cloud service provider and cloud host providing services for the user in the (user ID, cloud service provider, cloud host) list, and send an information query request to the distribution manager in the corresponding cloud service provider;
[0080] D103: Receive monitoring information returned by the distribution manager and push it to the user;
[0081] D104: Regularly receives monitoring information from all distribution managers. When the collected indicator data violates the set rules, the alarm module will trigger an alarm and notify the cloud user. The alarm module consists of an alarm database and an alarm generator.
[0082] The alarm database is used to maintain a set of data describing the alarms issued by the fault detection mechanism. In addition to recording the working history of the fault detection mechanism, the alarm database also includes a certain number of groups for classifying each alarm. Before issuing an alarm to the relevant parties, the alarm module searches the alarm database for similar alarm categories. This is very important to reduce the number of messages sent, because multiple alarms can be replaced by a single alarm with a greater impact.
[0083] Alert Generator: The alert generator is primarily responsible for generating alerts and sending notifications to cloud users and cloud service providers. Alerts are set up using alert policies. The cloud administrator's alert policy specifies under what circumstances and how he wants to receive alerts. This embodiment uses an indicator-based alert policy to track indicator data obtained by the monitoring architecture. Alert policies provide the ability to define standards and conditions, and these standards are based on indicators. Alert policy conditions can track situations where indicators reach specific values (thresholds) or begin to change rapidly. Indicators are associated with resources and can measure certain aspects of resources, such as average CPU utilization of instances, virtual machines in use, total storage, memory, swap memory, and network.
[0084] like Figure 3 As shown, step D104 is specifically as follows:
[0085] D201: Check the monitoring indicators in the current frame,
[0086] D202: Find and classify errors based on monitoring indicators. Can detect and provide different alert mechanisms, including 1) abnormal alerts, 2) intrusion alerts, 3) billing alerts, 4) SLA alerts, etc.
[0087] D203: Generate an alarm notification and add it to the notification queue. The thresholds for each alarm are preset. When the indicator matches the trigger specified by the alarm, the alarm generator will generate an alarm notification to notify the user. The parameters of the alarm notification include the project related to the alarm, the related user ID, the virtual machine related to the alarm, the time when the failure occurred, the maximum number of returned items, and the alarm type.
[0088] D204: Read unsent alarm notifications and select notification users based on the relevant users defined in the alarm.
[0089] D205: Create notification message based on user ID and embed alarm notification.
[0090] D206: Send the notification message to the user and check whether there are any remaining users who have not been notified. If all users have been notified, update the notification queue status.
[0091] The beneficial effects provided by the present invention are: the present invention can provide flexible and scalable distributed monitoring for large-scale cloud computing infrastructure, the monitoring components are organized in a distributed manner, and the monitoring tasks are shared according to the load balancing strategy, which can avoid data concentration. The distributed monitoring components of the present invention also support on-demand startup according to the changes in the scale of the architecture infrastructure, so that the architecture has self-organization ability, elasticity, scalability and fault tolerance, while avoiding single point failures that cause monitoring data loss and monitoring function failure. At the same time, the present invention considers the service quality of cloud computing, integrates the virtual machine instantiation control method into the monitoring component, and is used to monitor the resource consumption of the virtual machine instance, maintain the resource utilization rate of the virtual machine instance, and can effectively avoid the continuous saturation of the resource state to prevent system congestion, response loss and even infrastructure damage, and also avoid the waste of monitoring resources caused by the continuous idleness of the resource state. The present invention can ensure that the monitoring resources are used reasonably, thereby improving resource utilization efficiency and service quality. The present invention can be applied to provide monitoring services for cloud computing systems in different application scenarios, filling the gap in the field of distributed monitoring in the context of cloud computing in China today.
[0092] Embodiment 2
[0093] Based on the same inventive concept, this embodiment discloses a monitoring method for cloud computing, which is implemented based on the distributed monitoring architecture for cloud computing described in Example 1. The monitoring method includes:
[0094] Acquire monitoring information from specific resources on a physical machine and / or a virtual machine through a monitoring sensor, establish communication with a target collector, and send the acquired monitoring information to the target collector when receiving a request from the target collector;
[0095] Collect and process the monitoring information sent by the monitoring sensors through the collector;
[0096] The distribution manager allocates available collectors to a group of monitoring sensors and is responsible for acquiring, integrating and storing monitoring information collected and processed by all collectors in the cloud service provider.
[0097] The supervisory agent provides cloud users with an access point to cloud monitoring data and is responsible for detecting abnormal events and generating alarm notifications.
[0098] The monitoring information sent by the monitoring sensor is collected and processed by the collector, including:
[0099] The collector collects monitoring information from a group of monitoring sensors through the data collection module, and analyzes the resource usage status of physical machines and virtual machines based on the monitoring information obtained by the data collection module in the collector through the virtual machine instantiation control module, and plans related virtual machine instantiation operations;
[0100] When the collector receives the information request from the allocation manager, it responds with the monitoring data from the data collection module and the status results of all physical machines and virtual machines analyzed by the virtual machine instantiation module.
[0101] More specifically, another monitoring method for cloud computing provided in this embodiment includes the following specific steps:
[0102] A101: Run monitoring sensors on physical machines and / or virtual machines in each cloud service provider infrastructure and obtain each required monitoring metric from specific resources (CPU, memory, network, disk, etc.) on the physical machines and / or virtual machines;
[0103] A102: The monitoring sensor locates the distribution manager through broadcast messages, obtains the target collector, and establishes communication with the target collector. It calls a subset of specific sensors based on the monitoring indicators requested by the collector and sends the monitoring information to the corresponding collector to perform processing and storage tasks.
[0104] A103: The collector collects monitoring information from a set of monitoring sensors through the data collection module, and analyzes the resource usage status of the physical machine and virtual machine based on the monitoring values obtained by the data collection module in the collector and the available resources of the physical machine through the virtual machine instantiation control module, and plans related virtual machine instantiation operations (allocation, suspension, and recovery).
[0105] A104: When the collector receives the information request from the allocation manager, it responds with the monitoring data of the data collection module and the status results of all physical machines and virtual machines analyzed by the virtual machine instantiation module.
[0106] A105: When a monitoring sensor requests a collector assignment from the distribution manager, the load distribution module assigns available collectors to a set of monitoring sensors (installed on PMs and / or VMs). When a collector fails, the load distribution module redistributes the monitoring sensors associated with the failed collector to all available collectors, thereby maintaining a fair load balancing strategy.
[0107] A106: The distribution manager obtains, integrates and stores the information collected and processed by all collectors through the information integration module, that is, the current workload status information of the cloud service provider infrastructure. It receives the information query request of the supervision agent and sends general or specific monitoring information query to the supervision agent according to the query request of the supervision agent.
[0108] A107: The monitoring agent maintains a list of (user id, cloud service provider, cloud host) and provides a multi-functional API as an access point for users to access monitoring data;
[0109] A108: When the supervision agent receives a request from a user to view monitoring information, it searches for the cloud service provider and cloud host providing services to the user in the list of (user ID, cloud service provider, cloud host) according to the user ID, and sends an information query request to the distribution manager in the corresponding cloud service provider;
[0110] A109: Send a query request to the distribution manager, receive the monitoring information returned by the distribution manager, and push it to the user;
[0111] A110: Regularly receives monitoring information from all distribution managers. When the collected indicator data violates the set rules, the alarm module will trigger an alarm and notify the cloud user. The alarm module consists of an alarm database and an alarm generator.
[0112] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0113] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0114] Although preferred embodiments of the present invention have been described, additional changes and modifications may be made to these embodiments by those skilled in the art once the basic creative concepts are known. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention. Obviously, those skilled in the art may make various changes and modifications to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if these modifications and variations of the embodiments of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
Claims
1. A distributed monitoring architecture for cloud computing, characterized in that: Applied to a cloud computing system, the cloud computing system includes physical machines, virtual machines and cloud service providers, and the distributed monitoring architecture includes: A monitoring sensor is located in each running physical machine and virtual machine in the cloud computing system, and is used to obtain monitoring information from specific resources on the physical machine and / or virtual machine, establish communication with the target collector, and send the obtained monitoring information to the target collector when receiving a request from the target collector; The collector is located in the cloud service provider in the cloud computing system and is used to collect and process the monitoring information sent by the monitoring sensors; The distribution manager is located in the cloud service provider in the cloud computing system and is used to allocate available collectors to a group of monitoring sensors. It is also responsible for acquiring, integrating and storing the monitoring information collected and processed by all collectors in the cloud service provider. The monitoring agent is located outside the cloud service provider and is used to provide cloud users with access points to cloud monitoring data. It is also responsible for detecting abnormal events and generating alarm notifications.
2. The distributed monitoring architecture for cloud computing as claimed in claim 1, characterized in that: The monitoring sensor is also used to locate the distribution manager to obtain the target collector, by using broadcast messaging to pass the located distribution manager and accept the target collector offered by the distribution manager.
3. The distributed monitoring architecture for cloud computing as claimed in claim 1, characterized in that: The collector is pre-registered on the distribution manager, and has a built-in data collection module and a virtual machine instantiation control module. Wherein, the data collection module is used to collect monitoring information from a group of monitoring sensors; The virtual machine instantiation control module is used to plan the virtual machine instantiation operation according to the monitoring information obtained by the data collection module.
4. The distributed monitoring architecture for cloud computing as claimed in claim 1, characterized in that: A receive-request hybrid communication mode is adopted between the monitoring sensor and the collector. In this communication mode, the monitoring sensor regularly provides monitoring reports to the collector according to a certain threshold, and the collector requests the monitoring sensor to provide monitoring information within a certain time interval.
5. The distributed monitoring architecture for cloud computing as claimed in claim 1, characterized in that: The distribution manager has a built-in load distribution module and an information integration module. Among them, the load distribution module is used to distribute the available collectors to a group of monitoring sensors using a load balancing strategy; The information integration module is used to obtain, integrate and store the information collected and processed by all collectors in the cloud service provider. At the same time, according to the query requirements of the supervision agent, it sends general or specific monitoring information queries to the supervision agent.
6. The distributed monitoring architecture for cloud computing as claimed in claim 5, characterized in that: The load distribution module is used to distribute available collectors to a group of monitoring sensors using a load balancing strategy, including: When a monitoring sensor asks the distribution manager for available target collectors, the distribution manager assigns a registered collector according to a specific load balancing policy; when a collector fails, the distribution manager redistributes the monitoring sensors associated with the failed collector to all available collectors.
7. The distributed monitoring architecture for cloud computing as claimed in claim 1, characterized in that: The supervisory agent has a built-in alarm module for notifying users of abnormal processes during infrastructure operations.
8. The distributed monitoring architecture for cloud computing as claimed in claim 7, characterized in that: The alarm module includes an alarm database and an alarm generator; Among them, the alarm database is used to maintain a set of data describing the alarms issued by the fault detection mechanism; The alert generator is used to generate alerts and send notifications to cloud users and cloud service providers.
9. A monitoring method for cloud computing, characterized in that: Based on the distributed monitoring architecture for cloud computing described in any one of claims 1 to 8, the monitoring method includes: Acquire monitoring information from specific resources on a physical machine and / or a virtual machine through a monitoring sensor, establish communication with a target collector, and send the acquired monitoring information to the target collector when receiving a request from the target collector; Collect and process the monitoring information sent by the monitoring sensors through the collector; The distribution manager allocates available collectors to a group of monitoring sensors and is responsible for acquiring, integrating and storing monitoring information collected and processed by all collectors in the cloud service provider. The supervisory agent provides cloud users with an access point to cloud monitoring data and is responsible for detecting abnormal events and generating alarm notifications.
10. The distributed monitoring method for cloud computing according to claim 9, characterized in that: The collector collects and processes the monitoring information sent by the monitoring sensors, including: The collector collects monitoring information from a group of monitoring sensors through the data collection module, and analyzes the resource usage status of physical machines and virtual machines based on the monitoring information obtained by the data collection module in the collector through the virtual machine instantiation control module, and plans related virtual machine instantiation operations; When the collector receives the information request from the allocation manager, it responds with the monitoring data from the data collection module and the status results of all physical machines and virtual machines analyzed by the virtual machine instantiation module.
Citation Information
Patent Citations
Charging method and charging system for cloud computing
CN103152393A
Cloud monitoring system and method for private cloud
CN104113596A
Elastic scaling method and system for power system
CN107515809A
Cloud service management system based on multi-cloud application intelligent operation and maintenance
CN118585372A
High Frequency Skin Care Device
KR1020210117603A