Method and system for monitoring abnormity of cloud platform

Through automatic monitoring and alarming of monitoring units independent of the cloud platform, the problem of difficult to detect abnormalities in the cloud platform message middleware is solved, efficient and accurate abnormal detection and processing is achieved, and the stability of the cloud platform is ensured.

CN120295864APending Publication Date: 2025-07-11BEIJING URBAN CONSTR INTELLIGENT CONTROL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510427660.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the prior art, the message middleware exceptions of the cloud platform are difficult to be discovered and processed in time, resulting in unstable operation of the cloud platform, and relying on manual monitoring efficiency and easy to miss abnormal information.

Method used

The monitoring unit is adopted independently deployed from the cloud platform, and the operation information of the message middleware is collected through components such as Prometheus, and the abnormal alarm information is automatically monitored and pushed according to the set alarm strategy, including indicators such as memory usage, process status and disk space.

Benefits of technology

Automatic monitoring and timely alarm of message middleware exceptions is realized, which reduces the impact on the cloud platform, improves the reliability and accuracy of the monitoring system, and ensures the stable operation of the cloud platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295864A_ABST
    Figure CN120295864A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an exception monitoring method and system for a cloud platform, the exception monitoring method for the cloud platform is applied to a monitoring unit, the monitoring unit is deployed independently of the cloud platform, and the exception monitoring method for the cloud platform comprises the steps that operation information of message middleware in the cloud platform is collected, and the message middleware is a message transmission component of the cloud platform; determining whether the message-oriented middleware in the cloud platform is abnormal or not according to the operation information and a set alarm strategy of the message-oriented middleware; and if the message-oriented middleware is abnormal, pushing abnormal alarm information of the message-oriented middleware in the cloud platform. According to the invention, the abnormity of the message-oriented middleware can be automatically monitored and alarmed in time without depending on artificial experience, the missing of abnormal information is avoided, the abnormity of the message-oriented middleware can be found in time, the abnormity can be solved in time when the abnormity occurs, and the stable and safe operation of the whole cloud platform is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the technical field of cloud computing, and particularly to an abnormal monitoring method and system for a cloud platform. Background Art

[0002] With the rapid development of computer technology and cloud computing technology, various cloud platforms have emerged. To ensure the stable operation of the cloud platform, continuous monitoring of the cloud platform is required during its operation. The cloud platform is a large-scale distributed architecture, and the communication between components is implemented through a message middleware. From the perspective of the actual deployment environment, the message middleware is a common place where exceptions occur, and an exception in the message middleware will cause unknown errors in the entire cloud platform and is difficult to troubleshoot.

[0003] In the prior art, it is often the operation and maintenance personnel who analyze and troubleshoot the operation status of the cloud platform to monitor whether the message middleware has an exception. However, the method of manual monitoring and troubleshooting relies heavily on the experience of the operation and maintenance personnel, consumes a large amount of labor costs, and is prone to missing exception information, resulting in the inability to detect the exceptions of these message middleware in a timely manner, and naturally unable to respond to such exceptions in a timely manner, bringing risks to the stable operation of the cloud platform. Therefore, there is an urgent need for a more efficient and accurate abnormal monitoring solution for the cloud platform. Summary of the Invention

[0004] In view of this, the embodiments of this specification provide an abnormal monitoring method for a cloud platform. One or more embodiments of this specification also relate to an abnormal monitoring system for a cloud platform, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.

[0005] According to the first aspect of the embodiments of this specification, an abnormal monitoring method for a cloud platform is provided, which is applied to a monitoring unit that is independently deployed from the cloud platform. The method includes: Collect the operation information of the message middleware in the cloud platform, where the message middleware is the message transfer component of the cloud platform; Determine whether there is an abnormality in the message middleware in the cloud platform according to the operation information and the set alarm policy of the message middleware; If an abnormality occurs in the message middleware, push the abnormal alarm information of the message middleware in the cloud platform.

[0006] According to the second aspect of the embodiments of this specification, an abnormal monitoring system for a cloud platform is provided, including a monitoring unit and a cloud platform, where the monitoring unit is independently deployed from the cloud platform; The monitoring unit is configured to collect the operation information of the message middleware in the cloud platform, where the message middleware is the message transmission component of the cloud platform; determine whether there is an abnormality in the message middleware in the cloud platform according to the operation information and the set alarm policy of the message middleware; if there is an abnormality in the message middleware, push the abnormal alarm information of the message middleware in the cloud platform.

[0007] According to the third aspect of the embodiments of this specification, a computing device is provided, including: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned abnormal monitoring method of the cloud platform are implemented.

[0008] According to the fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by the processor, the steps of the above-mentioned abnormal monitoring method of the cloud platform are implemented.

[0009] According to the fifth aspect of the embodiments of this specification, a computer program product is provided, including computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned abnormal monitoring method of the cloud platform are implemented.

[0010] An embodiment of this specification provides an abnormal monitoring method for a cloud platform, which is applied to a monitoring unit. The monitoring unit is deployed independently of the cloud platform and collects the operation information of the message middleware in the cloud platform, where the message middleware is the message transmission component of the cloud platform; determine whether there is an abnormality in the message middleware in the cloud platform according to the operation information and the set alarm policy of the message middleware; if there is an abnormality in the message middleware, push the abnormal alarm information of the message middleware in the cloud platform.

[0011] The embodiments of this specification achieve the deployment of the monitoring unit independent of the cloud platform to collect the operation information of the message middleware in the cloud platform, automatically analyze the operation information of the message middleware according to the set alarm policy of the message middleware, determine whether the message middleware has an abnormality, and if it is determined that there is an abnormality, automatically push the abnormal alarm information of the message middleware in the cloud platform, which can help the operation and maintenance personnel to timely discover the abnormality of the message middleware and solve the abnormality before major problems occur in the message middleware. In this way, the abnormality of the message middleware can be automatically monitored and alarmed in a timely manner, without relying on manual experience, avoiding the omission of abnormal information, and can timely discover the abnormality of the message middleware and solve the abnormality in a timely manner when the abnormality occurs, ensuring the stable and safe operation of the entire cloud platform. Moreover, the monitoring unit is deployed and operates independently of the cloud platform, and the abnormal monitoring of the message middleware has a relatively low impact on the cloud platform itself, making other modules of the cloud platform insensitive to the abnormal alarm information of the message middleware, having better versatility, and can be applied to various real production environments, improving the reliability of the abnormal monitoring system of the entire cloud platform. Description of the Drawings

[0012] Figure 1 is a flowchart of an abnormal monitoring method for a cloud platform provided by an embodiment of this specification; Figure 2 is a schematic structural diagram of a cloud platform provided by an embodiment of this specification; Figure 3 is an overall structural diagram of an abnormal monitoring method for a cloud platform provided by an embodiment of this specification; Figure 4 is a schematic diagram of the pushing process of abnormal alarm information of a message middleware provided by an embodiment of this specification; Figure 5 is a schematic structural diagram of an abnormal monitoring system for a cloud platform provided by an embodiment of this specification; Figure 6 is a flowchart of the processing process of an abnormal monitoring method for a cloud platform provided by an embodiment of this specification; Figure 7 is a schematic structural diagram of the monitoring unit in an abnormal monitoring system for a cloud platform provided by an embodiment of this specification; Figure 8 is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed Embodiments

[0013] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific embodiments disclosed below.

[0014] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0015] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0016] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select authorization or rejection.

[0017] First, the noun terms involved in one or more embodiments of this specification are explained.

[0018] Message middleware: It is a software or service that supports asynchronous sending and receiving of messages for communication between different applications in a distributed system. This type of middleware is usually used to build distributed systems and can help decouple different applications, services, or system components, improve the flexibility, reliability, and scalability of the system. Message middleware can handle asynchronous communication, so that the producer (the sender of the message) and the consumer (the receiver of the message) do not need to be online at the same time, and the two parties do not have to interact directly.

[0019] Message Queue: A message queue is a common asynchronous communication mechanism used to transfer data between different applications or services in a distributed system. It can help applications reliably exchange and process data. The producer (the application that sends messages) sends messages to the queue, and the consumer (the application that receives and processes messages) retrieves and processes messages from the queue. The processing of both is asynchronous and does not affect each other. The producer and consumer are decoupled through the queue, so they do not need to directly know about each other's existence, reducing their dependence on each other and improving the flexibility, scalability, and reliability of the system. The message queue can buffer the traffic during peak periods, preventing the system from being overwhelmed by instantaneous high concurrency. It will be slowly consumed when the system is idle, achieving off-peak processing of traffic. The message queue provides mechanisms such as message re-delivery and persistent storage to ensure that messages are delivered at least once, improving the reliability of the system. In short, the message queue is a simple and efficient asynchronous communication mode, widely used in scenarios such as microservice architectures and big data processing, and is an important component for building reliable and high-performance distributed systems.

[0020] Prometheus: An open-source monitoring and alerting system with the following main features: Time series database: Prometheus uses its own built-in time series database, which can effectively store and query various monitoring metric data. The data is stored in a time dimension, facilitating time series analysis; Multi-dimensional data model: Prometheus uses a multi-dimensional data model based on labels, enabling fine-grained filtering and aggregation of monitoring data; Flexible query language: Prometheus provides the PromQL (Prometheus Query Language) query language, allowing users to perform complex monitoring metric queries and calculations to meet various monitoring analysis requirements; Highly available architecture: Prometheus supports horizontal scaling and can achieve data aggregation of multiple Prometheus instances through the Federation mechanism (Federation refers to a mechanism that allows one Prometheus server to scrape metric data from another Prometheus server. This setup is very useful for large-scale distributed environments as it helps address the scale limitations that a single Prometheus instance might encounter and enables cross-data center or regional data aggregation), enhancing the system's availability and scalability; Multiple integration methods: Prometheus can be easily integrated with tools such as Grafana (an open-source analytics and monitoring platform that allows users to visualize time series data through charts, dashboards, and alerts and supports multiple data sources), Alertmanager, etc., providing a complete monitoring solution. It also supports accessing monitoring data from various applications and systems through exporters. Generally speaking, Prometheus is a powerful and flexible open-source monitoring system, widely used in monitoring practices in cloud-native environments and is the de facto standard for container orchestration platforms such as Kubernetes (commonly abbreviated as K8s, an open-source container orchestration platform that allows for automated deployment, scaling, and management of containerized applications).

[0021] Alertmanager: A key component in the Prometheus ecosystem, responsible for handling alerts generated by Prometheus servers. Prometheus itself detects abnormal situations through defined rules and generates alerts when specific conditions are met. These alerts are then sent to Alertmanager, which is responsible for deduplicating, grouping, suppressing, and finally notifying the alerts.

[0022] Websocket: It is a full-duplex communication protocol based on TCP (Transmission Control Protocol). It provides a persistent communication channel that allows for two-way real-time communication between the client and the server. It enables the server to actively push data to the client without the client having to initiate a request. Different from the traditional HTTP (HyperText Transfer Protocol), which is based on the request-response mode and requires the client to send a request first to receive data from the server, WebSocket makes real-time communication more efficient and simple, and is more resource-saving compared to the HTTP protocol.

[0023] Webhook: It is an HTTP callback mechanism. When a specific event occurs, the source application sends an HTTP POST request to a pre-configured URL to notify the target application of the event, enabling the recipient to immediately respond to the event and realizing instant and automated data exchange and integration between applications. Compared with the traditional polling method, Webhook is more efficient and real-time because it does not require the recipient to regularly check for new data.

[0024] Openstack: It is a free and open-source cloud computing management platform used to build and manage public clouds, private clouds, and hybrid cloud environments, providing a set of software for managing and controlling large-scale computing, storage, and network resources. Openstack provides an extensible framework that allows users to deploy and manage virtual machines (VMs), storage resources, and network services through a web-based control panel, command-line tools, or RESTful API (a web service interface designed based on the REST (Representational State Transfer) architectural style. REST is a design style rather than a specific protocol or technology. It defines a set of constraints and principles for creating extensible, stateless, and easy-to-understand web services).

[0025] Rabbitmq: It is an open-source message queue system that supports multiple messaging protocols, providing highly reliable, highly scalable, and highly available asynchronous messaging services. It implements the Advanced Message Queuing Protocol (AMQP). It is widely used to implement asynchronous messaging and decouple application components in distributed systems. Rabbitmq provides reliable message delivery, flexible routing mechanisms, and support for multiple message patterns.

[0026] Exporter: A component in the Prometheus monitoring system. In the field of monitoring and observability, an Exporter generally refers to a tool or service that is responsible for collecting metrics from applications, systems, or other data sources and exposing these metrics to the Prometheus service in a format recognizable by Prometheus. Prometheus fetches these metrics via the HTTP protocol and then stores and uses them to generate alerts, visualization charts, etc.

[0027] It should be noted that OpenStack, as a cloud computing platform, provides solutions for cloud computing infrastructure services. With its fully open-source and easy-to-expand characteristics, OpenStack has attracted more and more attention in the industry. In actual applications, OpenStack is prone to various instability problems and even system crashes. However, the OpenStack native system in the development stage does not provide monitoring and alerting for the message queue Rabbitmq. However, from the perspective of the actual deployment environment, Rabbitmq is a very common place where exceptions occur, and exceptions in Rabbitmq can lead to unknown errors in the entire OpenStack and are difficult to troubleshoot.

[0028] In the embodiments of this specification, a monitoring unit (Prometheus) independent of the OpenStack deployment is used to monitor and alert Rabbitmq, which can cooperate with operation and maintenance personnel to prevent exceptions in the message queue Rabbitmq and resolve exceptions in a timely manner when problems occur. Moreover, the alerting module of OpenStack (cloud platform) itself can be used to notify the abnormality of the Rabbitmq cluster, and the use of the monitoring unit (Prometheus) improves the reliability of the exception monitoring system of the entire cloud platform.

[0029] In this specification, an exception monitoring method for a cloud platform is provided. This specification also relates to an exception monitoring system for a cloud platform, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.

[0030] See Figure 1 , Figure 1 which shows a flowchart of an exception monitoring method for a cloud platform provided according to an embodiment of this specification. The method is applied to a monitoring unit that is independent of the cloud platform deployment and specifically includes the following steps.

[0031] Step 102: Collect the running information of the message middleware in the cloud platform, where the message middleware is the message passing component of the cloud platform.

[0032] Among them, the cloud platform is a collection of Internet-based computing resources and services that allows users to access and use various computing resources, such as servers, storage, databases, networks, software, etc., through the network. The cloud platform provides flexible, scalable, and on-demand resources, enabling enterprises and individuals to quickly deploy and manage applications without investing in expensive hardware infrastructure. The infrastructure capabilities of the cloud platform are highly elastic (increase and decrease) and can be dynamically scaled and configured according to needs. For example, enterprises and institutions no longer need to plan their own data centers or overly concern themselves with IT management unrelated to their core business. They only need to give instructions to the "cloud platform" to obtain information services of different levels and types, and the time and resources saved can be invested in enterprise operations. For individual users, there is no need to spend a large amount of money to purchase software, and the services in the cloud platform can provide various required functions.

[0033] In the embodiments of this specification, the cloud platform is taken as Openstack as an example for illustration. Of course, in actual implementation, the cloud platform can also be services provided by other service providers, such as IaaS (Infrastructure as a Service), PaaS (Platform as a Service), and SaaS (Software as a Service), etc. The embodiments of this specification do not limit this.

[0034] Message middleware is a software or service that supports asynchronous sending and receiving of messages between different applications in a distributed system for communication. This type of middleware is commonly used to build distributed systems. Message middleware can be common message queues, such as Kafka, Rabbitmq. Or, message middleware can also be other types of functions and services to adapt to different application scenarios and technical requirements, such as publish / subscribe (Pub / Sub) components: Different from point-to-point message queues, publish / subscribe components allow a producer to broadcast messages to multiple consumers, and consumers can subscribe to specific topics to receive messages of interest; event-driven architecture (EDA): This is a design pattern in which the interaction between system components is triggered by events, and message middleware acts as an event bus in this architecture to facilitate the generation, transmission, and processing of events.

[0035] The monitoring unit is a service deployed independently of the cloud platform. The monitoring unit can be Prometheus. Components in the monitoring unit are used to implement the abnormal monitoring and alarming of the message middleware in the cloud platform. Of course, in actual implementation, other open-source monitoring systems can also be used for the monitoring unit, such as Zabbix, Nagios, etc. Zabbix is an enterprise-level open-source network monitoring solution that supports real-time monitoring of networks, servers, applications, etc. Nagios is an open-source computer system, network, and infrastructure monitoring tool that can monitor the status of hosts and services, send alarms when problems occur, and provide a powerful alarming and notification mechanism. The embodiments of this specification do not limit this.

[0036] The running information of the message middleware refers to various key indicators and status data collected and displayed during the operation of the message middleware in the cloud platform. These information can reflect the health status, performance, and resource usage of the message middleware in the cloud platform.

[0037] It should be noted that the cloud platform is a large-scale distributed architecture software. A large number of RPCs (Remote Procedure Call) are used for communication between components. These numerous RPCs are implemented based on the message middleware in the cloud platform. Taking the cloud platform Openstack as an example, it uses Rabbitmq to provide the message queue infrastructure. As the message middleware, the message middleware is the most core message-passing component in the cloud platform. Although the message middleware does not provide the relevant functions of the cloud platform, when unexpected abnormalities occur to it, the impact on the overall reliability of the cloud platform is relatively large. Among them, RPC is a technology that allows a program to execute a procedure or method located on another computer. Through RPC, developers can write distributed applications, where the client can directly call the procedure located on the server side as simply as calling a local procedure. This method hides the complexity of network communication and enables developers to focus on the implementation of business logic.

[0038] Exemplarily, Figure 2 is a schematic structural diagram of a cloud platform provided by an embodiment of this specification, as Figure 2As shown, taking the cloud platform OpenStack as an example, the cloud platform includes Heat, Horizon, Neutron, Nova, Glance, VM, Ceilometer, Cinder, Swift, and Keystone. Among them, Heat is a component in OpenStack, mainly used to provide template-based orchestration services. It allows users to define and deploy complex cloud infrastructures through description files, including computing, storage, and network resources, etc. Horizon is the official dashboard of OpenStack, providing UI services. It provides a web-based graphical user interface (GUI), allowing users to manage and operate the OpenStack cloud environment through a browser. Horizon enables non-technical users to easily use various OpenStack services without directly interacting with complex command-line interfaces or APIs. Neutron is the network service component of OpenStack. It is responsible for managing and providing virtual network infrastructures and providing network connections. Neutron enables users to create and manage complex network topologies in the OpenStack environment, including virtual networks, subnets, routers, load balancers, etc. Through Neutron, OpenStack can support multiple network configurations and services, providing flexible network solutions for various applications in the cloud environment. Nova is the computing service component of OpenStack, responsible for managing and scheduling virtual machine instances. Nova provides a set of APIs that allow users to create, manage, and delete virtual machine instances and control the life cycle of these instances. Nova is tightly integrated with other core services of OpenStack (such as the Neutron network service and the Cinder block storage service) to jointly provide users with a complete cloud computing environment. Glance is the image service component of OpenStack, providing image services for virtual machines VM and also providing storage image services for Swift. It provides functions for discovering, registering, and retrieving virtual machine images. Glance enables users to store and manage virtual machine images in various formats and can be integrated with other OpenStack services (such as Nova and Cinder) to support the creation of new virtual machine instances or volumes. VM usually refers to "Virtual Machine" (virtual machine). A virtual machine is a computer system simulated by software. It can run on physical hardware and execute programs like a real computer. Ceilometer is an important component in OpenStack, used to monitor and meter the usage of cloud resources. Ceilometer provides comprehensive monitoring capabilities for OpenStack, capable of collecting, storing, and querying various metric data, such as CPU usage, memory usage, network traffic, etc.Cinder is a component in OpenStack responsible for block storage services. It provides scalable and persistent block storage volumes (volumes) that can be mounted on virtual machine instances as independent disks. Cinder enables users to create, manage, attach, and detach block storage volumes and supports multiple backend storage solutions. Swift is the object storage service of OpenStack, which is a backup volume. It offers a highly scalable, highly available, and distributed storage solution. Swift is suitable for storing large amounts of unstructured data such as pictures, videos, backup files, etc. Different from traditional file systems or block storage, Swift uses a flat namespace to store objects and provides access through a RESTful API. Keystone is the Identity Service of OpenStack. It provides functions such as user authentication, service catalog, and policy-based access control. Keystone plays a crucial role in the OpenStack environment as it is responsible for managing all authentication information and providing a unified authentication mechanism for other OpenStack services.

[0039] As Figure 2 shown, it presents the overall architecture diagram of Openstack. Figure 2 In Openstack, the communication within each component is achieved through a message middleware. For example, if the message middleware is a message queue, the Nova component includes the Nova-api service and the Nova-compute service, and the communication between these services is realized through the message queue.

[0040] Therefore, in the embodiments of this specification, it is not necessary to install a collection plugin and a monitoring plugin on the nodes of the cloud platform to collect the running information of the message middleware in the cloud platform, so as to monitor the exceptions of the message middleware. Instead, an exception monitoring system for the cloud platform is constructed. The exception monitoring system for the cloud platform includes the cloud platform and a monitoring unit deployed independently of the cloud platform. The monitoring unit is used to collect the running information of the message middleware in the cloud platform, and the running situation of the message middleware in the cloud platform is reflected through this running information, so as to monitor the exceptions of the message middleware in a timely manner. The impact of the exception monitoring of the message middleware on the cloud platform itself is relatively low, and there is no need to change the native structure of the cloud platform, making other modules of the cloud platform insensitive to the exception alarms of the message middleware, having better generality, and being applicable to various real production environments, thus improving the reliability of the entire exception monitoring system of the cloud platform.

[0041] In an optional implementation manner of this embodiment, the monitoring unit includes a metric collection component and a monitoring service; collecting the running information of the message middleware in the cloud platform includes: Collect the running information of the message middleware from the cloud platform through the metric collection component, where the running information includes at least one of the memory size occupied by the message middleware, the process running status, and the available disk size; Push the running information to the monitoring service in a set format, where the set format is a data format recognizable by the monitoring service.

[0042] In actual implementation, the monitoring unit may include a metric collection component and a monitoring service. The metric collection component is used to collect the running information of the message middleware from the cloud platform and push it to the monitoring service in a data format recognizable by the monitoring service to enable subsequent data analysis. As an example, the metric collection component may be Rabbitmq-Exporter, and the monitoring service is a Prometheus server. Based on Rabbitmq-Exporter, collect the running information of the message middleware in the cloud platform and expose it to the Prometheus server in a format recognizable by Prometheus. The Prometheus server can fetch this running information through the HTTP protocol and then store it and use it to generate exception alarm information, visualization charts, etc.

[0043] In addition, the running information of the message middleware may include at least one of the memory size occupied by the message middleware, the process running status, and the available disk size. Of course, in actual implementation, the running information may also include the number of messages, the number of consumers, message traffic (such as the message sending and receiving rate per second), performance metrics (such as CPU and memory usage rates), log records (for diagnosis and debugging), and configuration information (such as the maximum length of the queue and the message expiration time), etc., to reflect the running situation of the message middleware.

[0044] In the embodiments of this specification, the running information of the message middleware in the cloud platform can be collected through the metric collection component (Rabbitmq-Exporter) and the monitoring service (Prometheus server), so as to realize the abnormal monitoring of the message middleware in the cloud platform. Through the collected running information, potential problems of the message middleware can be discovered and solved in time to ensure the reliability of message transmission in the cloud platform and the efficient operation of the system. In addition, the metric collection component (Rabbitmq-Exporter) and the monitoring service (Prometheus server) cooperate to monitor the running information, without changing the native structure of the cloud platform, having less impact on the cloud platform, and improving the reliability of the abnormal monitoring system of the entire cloud platform.

[0045] Step 104: Determine whether there is an abnormality in the message middleware in the cloud platform according to the running information and the set alarm policy of the message middleware.

[0046] Among them, the set alarm policy is a rule configured in the monitoring unit to determine whether there is an abnormality in the message middleware. For example, the set alarm policy may include at least one of the following: the memory size exceeds the memory threshold, the process running status is abnormal, and the available disk size exceeds the disk threshold.

[0047] Implement, the monitoring unit can be pre-configured with the set alarm policy of the message middleware, and then analyze the running information of the message middleware according to the set alarm policy of the message middleware to determine whether there is an abnormality in the message middleware in the cloud platform.

[0048] In an optional implementation manner of this embodiment, the method further includes: Deploy a monitoring unit in the exception monitoring system of the cloud platform, and configure the set alarm policy of the message middleware in the monitoring unit through a configuration file during the deployment process of the monitoring unit.

[0049] It should be noted that the set alarm policy of the message middleware can be configured in the monitoring unit through a configuration file during the deployment process of the monitoring unit, so as to monitor the running information of the message middleware based on the pre-configured set alarm policy.

[0050] As an example, the monitoring unit may include a monitoring service Prometheus and a push service Alertmanager, and the set alarm policy is configured through a YAML configuration file. Specifically, it is necessary to prepare the environment first to ensure that Prometheus and Alertmanager have been installed and configured. These two components are usually deployed through Kubernetes or other container orchestration tools; then, create a file for the set alarm policy, and use a file in YAML format in Prometheus to define the set alarm policy. These files are usually stored in the configuration directory of Prometheus and are referenced through the prometheus.yml file; then, in the main configuration file prometheus.yml of Prometheus, add a reference to the file for the set alarm policy. Alertmanager is responsible for processing the exception alarm information sent by Prometheus and sending the exception alarm information to different receivers (such as email, Slack, PagerDuty, etc.) according to the configuration.

[0051] In the embodiments of this specification, when deploying the monitoring unit, the set alarm policy of the message middleware can be configured in the monitoring unit through a configuration file. The configuration process of the set alarm policy is simple and does not require additional interaction or complex calculations. After the monitoring unit collects the running information of the message middleware, it can analyze the running information of the message middleware based on the configured set alarm policy to monitor whether the message middleware has an abnormality, enabling the monitoring unit independent of the cloud platform to have the function of monitoring the message middleware and being able to monitor the abnormality of the message middleware in a timely manner.

[0052] Of course, in practical applications, in addition to configuring the set alarm policy of the message middleware based on the configuration file as described above, the monitoring unit can also provide an alarm policy configuration interface, and the operation and maintenance personnel can configure or edit the set alarm policy based on this alarm policy configuration interface. This specification does not limit this.

[0053] In an optional implementation manner of this embodiment, the set alarm policy includes at least one of the following: the memory size exceeds the memory threshold, the process running status is abnormal, and the available disk size is lower than the disk threshold; Determining whether there is an abnormality in the message middleware in the cloud platform according to the running information and the set alarm policy of the message middleware includes at least one of the following: If the memory size occupied by the message middleware exceeds the memory threshold, it is determined that there is an abnormality in the message middleware in the cloud platform; If the process running status is abnormal, it is determined that there is an abnormality in the message middleware in the cloud platform; If the available disk size is lower than the disk threshold, it is determined that there is an abnormality in the message middleware in the cloud platform.

[0054] It should be noted that the running information may include at least one of the memory size occupied by the message middleware, the process running status, and the available disk size, and the set alarm policy may include at least one of the following: the memory size exceeds the memory threshold, the process running status is abnormal, and the available disk size exceeds the disk threshold.

[0055] In the embodiments of the present specification, if the memory size occupied by the message middleware exceeds the memory threshold, it indicates that there is excessive message accumulation in the message middleware, resulting in the memory occupancy exceeding the set threshold. At this time, it can be determined that the message middleware in the cloud platform is abnormal; if the process running status is abnormal, it indicates that the message middleware may have crashed. At this time, it can be determined that the message middleware in the cloud platform is abnormal; if the available disk size is lower than the disk threshold, it indicates that the available disk space is relatively small. If the message middleware continues to operate, it may cause system failures. Therefore, at this time, it can be determined that the message middleware in the cloud platform is abnormal. In this way, it is possible to monitor the running conditions of the message middleware in multiple aspects, such as message accumulation, running status, and available disk size, perform multi-faceted monitoring on the message middleware, improve the monitoring coverage of the message middleware, avoid missing abnormal information, and improve the monitoring accuracy rate.

[0056] In an optional implementation manner of this embodiment, determining whether the message middleware in the cloud platform is abnormal according to the running information and the set alarm policy of the message middleware includes: Aggregating the running information through a multi-dimensional data model in the monitoring unit to obtain the running information of at least one dimension; Determining whether the message middleware in the cloud platform is abnormal according to the running information of at least one dimension and the set alarm policy corresponding to each dimension.

[0057] It should be noted that the running information may be messy data including various types of information. A multi-dimensional data model can be deployed in the monitoring unit. For example, a multi-dimensional data model based on labels (label) is used in Prometheus, which can perform fine-grained screening and aggregation on the running information of the message middleware to obtain the running information of at least one dimension. Each dimension is pre-configured with a corresponding set alarm policy, and then it is determined whether the message middleware in the cloud platform is abnormal dimension by dimension.

[0058] As an example, the monitoring unit aggregates the running information based on the multi-dimensional data model to obtain dimensions such as performance metrics dimension and message consumption dimension. For the performance metrics dimension, the corresponding set alarm policy can be at least one of the memory size exceeding the memory threshold, the process running status being abnormal, and the available disk size being lower than the disk threshold, or it can also be that the CPU and memory usage rates exceed the usage threshold; for the message consumption dimension, the corresponding set alarm policy can be that the number of messages, the number of consumers, and / or the message traffic exceed the set number threshold.

[0059] In the embodiments of this specification, for the running information of the message middleware, dimensional monitoring can be performed to determine whether there are any abnormalities in different dimensions, more comprehensively monitor the message middleware, avoid missing abnormal information, improve the monitoring accuracy and comprehensiveness of the message middleware, timely detect abnormalities in each dimension of the message middleware, and promptly resolve the abnormalities when they occur, ensuring the stable and secure operation of the entire cloud platform.

[0060] Step 106: If an abnormality occurs in the message middleware, push the abnormal alarm information of the message middleware in the cloud platform.

[0061] It should be noted that if the monitoring unit monitors that an abnormality occurs in the message middleware of the cloud platform, it can push the abnormal alarm information of the message middleware in the cloud platform to the recipient, thereby alarming the abnormality of the message middleware and facilitating timely adoption of corresponding measures to resolve the abnormality of the message middleware.

[0062] In an optional implementation manner of this embodiment, pushing the abnormal alarm information of the message middleware in the cloud platform includes: Pushing the abnormal alarm information of the message middleware to the operation and maintenance personnel of the cloud platform through network communication; and / or, Pushing the abnormal alarm information of the message middleware to the cloud platform so that the cloud platform feeds back the abnormal alarm information of the message middleware to the front end through its own alarm module.

[0063] Specifically, the monitoring unit may include a push service, such as Alertmanager in the Prometheus ecosystem. Based on this push service, the abnormal alarm information of the message middleware can be pushed to the operation and maintenance personnel of the cloud platform through network communication, and / or the abnormal alarm information of the message middleware can be pushed to the cloud platform so that the cloud platform feeds back the abnormal alarm information of the message middleware to the front end through its own alarm module.

[0064] Among them, Alertmanager is a component in the Prometheus ecosystem, mainly used to process alarm information sent by Prometheus or other monitoring tools. Prometheus is an open-source system monitoring and alerting toolkit. It works by scraping the running information of the message middleware in the cloud platform and can generate abnormal alarm information based on this data. When Prometheus detects an abnormality in the message middleware, it will trigger an alarm and generate abnormal alarm information to send to Alertmanager.

[0065] Specifically, the network communication method refers to the process of transmitting data between different devices through network infrastructure (such as the Internet, local area network, wide area network, etc.). The network communication method can be divided into multiple types, and each method has its specific application scenarios and technical characteristics. For example, the network communication method can be email.

[0066] In an optional implementation, the monitoring unit's push service Alertmanager can push the exception warning information of the message middleware to the mailbox of the mailbox operation and maintenance personnel through the email address. The operation and maintenance personnel can learn that there is an exception in the message middleware in the cloud platform by viewing the email, and then take corresponding handling measures. In another optional implementation, the monitoring unit's push service Alertmanager can push the exception warning information of the message middleware to the business layer (Portal) of the cloud platform based on Webhook, that is, the self-alarm module of the cloud platform. Then, this self-alarm module feeds back the exception warning information of the message middleware to the front end through websocket. That is, the front end of the cloud platform can display the operation and maintenance interface, and display the exception warning information of the message middleware in the operation and maintenance interface. The operation and maintenance personnel can learn that there is an exception in the message middleware in the cloud platform by entering the operation and maintenance interface of the cloud platform, and then take corresponding handling measures.

[0067] Among them, the exception warning information of the message middleware can include the dimension and / or specific exception data when the message middleware has an exception, so as to facilitate the operation and maintenance personnel to further locate the cause of the exception and take corresponding countermeasures.

[0068] Exemplarily, Figure 3 is the overall structure diagram of an exception monitoring method for a cloud platform provided by an embodiment of this specification, as Figure 3As shown, taking the Nova component in Openstack as an example, the Nova component includes the Nova-api service, the Nova-compute service, the Nova-conductor service, and the Nova-scheduler service. The communication between the above services is achieved through a message queue. Rabbitmq-Exporter (metric collection component) collects the running information of the message queue, and uses the Prometheus service to determine whether there is an abnormality in the message middleware in the cloud platform according to the running information and the set alarm policy of the message middleware. If an abnormality occurs, it uses Alertmanager to send an email to the mailbox of the operation and maintenance personnel. The email carries the abnormal alarm information of the message middleware, and pushes the abnormal alarm information of the message middleware to the Portal (business layer) of the cloud platform, that is, the self-alarm module of the cloud platform. The business layer (Portal) of the cloud platform feeds back the abnormal alarm information of the message middleware to the front end. The front end can also receive interaction information and feedback it to the business layer to achieve interaction, and the Nova-api service can also interact with the business layer of the cloud platform to perform interface calls to achieve corresponding business functions.

[0069] In another example, Figure 4 is a schematic diagram of the pushing process of the abnormal alarm information of a message middleware provided by an embodiment of this specification. As Figure 4 shown, the pushing service Alertmanager creates an alarm email. The alarm email carries the abnormal alarm information of the message middleware and sends the alarm email to the mailbox address of the operation and maintenance personnel through the email server. And, the pushing service Alertmanager pushes the abnormal alarm information of the message middleware to the business layer (Portal) of the cloud platform based on Webhook. The business layer (Portal) stores the abnormal alarm information of the message middleware in the database and feeds back the abnormal alarm information of the message middleware to the front end through websocket. That is, the front end of the cloud platform can display the operation and maintenance interface and display the abnormal alarm information of the message middleware in the operation and maintenance interface.

[0070] In the implementation of this specification, the monitoring unit can use emails to push the abnormal alarm information of the message middleware to the operation and maintenance personnel of the cloud platform, and / or display the abnormal alarm information of the message middleware in the operation and maintenance interface of the cloud platform through Alertmanager and websocket to perform abnormal alarm information on the message middleware. The operation and maintenance personnel can learn the alarm details of the message middleware by viewing the email and / or the operation and maintenance interface of the cloud platform to adopt corresponding processing strategies, which helps the operation and maintenance personnel to detect the abnormality of the message middleware in time and solve the abnormality before major problems occur in the message middleware, ensuring the operation safety of the entire cloud platform.

[0071] In addition, the push service Alertmanager pushes the exception warning information of the message middleware to the cloud platform, and the cloud platform can return the exception warning information of the message middleware to the front end through its own warning module. The entire monitoring unit is deployed and operates independently of the cloud platform, and the monitoring process has a low impact on the cloud platform itself, making other modules insensitive to the warning, and having better versatility.

[0072] In an optional implementation manner of this embodiment, the method further includes: Receiving a monitoring metric query request, where the monitoring metric query request carries a query dimension parameter; Querying and returning corresponding target operation information according to the query dimension parameter.

[0073] It should be noted that the monitoring unit can receive a monitoring metric query request. This monitoring metric query request can be directly initiated by an operation and maintenance personnel to the monitoring unit, or can be initiated by the operation and maintenance personnel through the cloud platform. The monitoring metric query request carries a query dimension parameter, which is the dimension to be queried currently, such as the dimension where an exception occurs, the performance metric dimension, the message consumption dimension, etc. The monitoring unit can query and return corresponding target operation information based on the query dimension parameter. In this way, the monitoring unit can also provide a query function, which is convenient for querying the operation information corresponding to the specific exception dimension required, or the detailed operation information under a certain dimension, facilitating the management of the cloud platform.

[0074] An embodiment of this specification provides an exception monitoring method for a cloud platform, which realizes automatic monitoring and timely warning of exceptions in the message middleware, does not rely on manual experience, avoids missing exception information, can timely detect exceptions in the message middleware and resolve exceptions in a timely manner when an exception occurs, ensuring the stable and secure operation of the entire cloud platform. Moreover, the monitoring unit is deployed and operates independently of the cloud platform, and the exception monitoring of the message middleware has a low impact on the cloud platform itself, making other modules of the cloud platform insensitive to the exception warning information of the message middleware, having better versatility, and can be applied to various real production environments, improving the reliability of the exception monitoring system of the entire cloud platform.

[0075] See Figure 5 , Figure 5 shows a schematic structural diagram of an exception monitoring system for a cloud platform according to an embodiment of this specification. As Figure 5 shown, it includes a monitoring unit 502 and a cloud platform 504, and the monitoring unit 502 is deployed independently of the cloud platform 504; The monitoring unit 502 is configured to collect the running information of the message middleware 506 in the cloud platform 504, where the message middleware 506 is a message passing component of the cloud platform 504; determine whether the message middleware 506 in the cloud platform 504 has an abnormality according to the running information and the set alarm policy of the message middleware 506; if the message middleware 506 has an abnormality, then push the abnormal alarm information of the message middleware 506 in the cloud platform 504.

[0076] In an optional implementation manner of this embodiment, the monitoring unit 502 includes an index collection component 5022 and a monitoring service 5024; The index collection component 5022 is configured to collect the running information of the message middleware 506 from the cloud platform 504, where the running information includes at least one of the memory size occupied by the message middleware 506, the process running status, and the available disk size; push the running information to the monitoring service 5024 in a set format, where the set format is a data format recognizable by the monitoring service 5024; The monitoring service 5024 is configured to store the running information and determine whether the message middleware 506 in the cloud platform 504 has an abnormality according to the running information and the set alarm policy of the message middleware.

[0077] In actual implementation, the monitoring unit may include an index collection component and a monitoring service. The index collection component is used to collect the running information of the message middleware from the cloud platform and push it to the monitoring service in a data format recognizable by the monitoring service to implement subsequent data analysis. As an example, the index collection component may be Rabbitmq-Exporter, and the monitoring service is a Prometheus server. The running information of the message middleware in the cloud platform is collected based on Rabbitmq-Exporter and exposed to the Prometheus server in a format recognizable by Prometheus. The Prometheus server can capture this running information through the HTTP protocol, and then store and use it to generate warnings, visualization charts, etc.

[0078] In the embodiments of this specification, the operation information of the message middleware in the cloud platform can be collected through the index collection component (Rabbitmq-Exporter) and the monitoring service (Prometheus server), so as to realize the abnormal monitoring of the message middleware in the cloud platform. Through the collected operation information, potential problems of the message middleware can be discovered and solved in a timely manner, ensuring the reliability of message transmission in the cloud platform and the efficient operation of the system. In addition, the index collection component (Rabbitmq-Exporter) and the monitoring service (Prometheus server) cooperate to monitor the operation information, without changing the native structure of the cloud platform, having less impact on the cloud platform, and improving the reliability of the abnormal monitoring system of the entire cloud platform.

[0079] In an optional implementation manner of this embodiment, the monitoring unit 502 is further configured to push the abnormal alarm information of the message middleware to the cloud platform 504; The cloud platform 504 is configured to store the abnormal alarm information of the message middleware in the database; and feedback the abnormal alarm information of the message middleware to the front end through its own alarm module.

[0080] In an optional implementation manner of this embodiment, the cloud platform 504 is further configured to: In response to the management instruction, display the operation and maintenance interface of the cloud platform, and display the abnormal alarm information of the message middleware in the operation and maintenance interface.

[0081] It should be noted that the monitoring unit can push the abnormal alarm information of the message middleware to the cloud platform, and the cloud platform stores the abnormal alarm information of the message middleware in the database for subsequent query; and the abnormal alarm information of the message middleware can be fed back to the front end through its own alarm module.

[0082] In actual implementation, the push service Alertmanager in the monitoring unit can be used to push the abnormal alarm information of the message middleware to the business layer (Portal) of the cloud platform, that is, the own alarm module of the cloud platform, based on Webhook. Then, this own alarm module feeds back the abnormal alarm information of the message middleware to the front end through websocket. That is, the front end of the cloud platform can display the operation and maintenance interface, and display the abnormal alarm information of the message middleware in the operation and maintenance interface. The operation and maintenance personnel can learn that there is an abnormality in the message middleware in the cloud platform by entering the operation and maintenance interface of the cloud platform, and thus take corresponding handling measures.

[0083] It should be noted that the monitoring unit can display the exception alarm information of the message middleware on the operation and maintenance interface of the cloud platform through Alertmanager and websocket, so as to perform exception alarm information on the message middleware. The operation and maintenance personnel can obtain the alarm details of the message middleware through the operation and maintenance interface of the cloud platform to adopt corresponding processing strategies, which helps the operation and maintenance personnel to detect the exceptions of the message middleware in time, resolve the exceptions before major problems occur in the message middleware, and ensure the operation safety of the entire cloud platform. In addition, the push service Alertmanager pushes the exception alarm information of the message middleware to the cloud platform, and the cloud platform can return the exception alarm information of the message middleware to the front end through its own alarm module. The entire monitoring unit is deployed and runs independently of the cloud platform, and the monitoring process has a low impact on the cloud platform itself, making other modules insensitive to the alarm, and having better generality.

[0084] The embodiment of the present specification provides an exception monitoring system for a cloud platform. By deploying a monitoring unit independently of the cloud platform, it realizes the automatic monitoring and timely alarm of exceptions in the message middleware in the cloud platform, without relying on manual experience, avoiding the omission of exception information, and can timely discover exceptions in the message middleware and resolve the exceptions in a timely manner when exceptions occur, ensuring the stable and safe operation of the entire cloud platform. Moreover, the monitoring unit is deployed and runs independently of the cloud platform, and the exception monitoring of the message middleware has a low impact on the cloud platform itself, making other modules of the cloud platform insensitive to the exception alarm information of the message middleware, having better generality, and can be applied to various real production environments, improving the reliability of the exception monitoring system of the entire cloud platform.

[0085] The above is a schematic solution of an exception monitoring system for a cloud platform in this embodiment. It should be noted that the technical solution of the exception monitoring system for the cloud platform and the technical solution of the above-mentioned exception monitoring method for the cloud platform belong to the same concept. For the details not described in detail in the technical solution of the exception monitoring system for the cloud platform, reference can be made to the description of the technical solution of the above-mentioned exception monitoring method for the cloud platform.

[0086] The following combines the appended Figure 6 , taking the application of the exception monitoring method for the cloud platform provided in this specification in Openstack as an example, to further illustrate the exception monitoring method for the cloud platform. Among them, Figure 6 shows the processing procedure flowchart of an exception monitoring method for a cloud platform provided by an embodiment of this specification, which specifically includes the following steps.

[0087] Step 602: Rabbitmq-Exporter (metric collection component) collects the running information of Rabbitmq (message middleware) from Openstack (cloud platform), and pushes the running information to Prometheus (monitoring service) in a data format recognizable by Prometheus (monitoring service).

[0088] Step 604: Prometheus (monitoring service) determines whether there is an abnormality in Rabbitmq (message middleware) in Openstack (cloud platform) according to the running information and the set alarm policy of Rabbitmq (message middleware).

[0089] Step 606: If there is an abnormality in Rabbitmq (message middleware), an abnormal alarm information email of Rabbitmq (message middleware) is pushed to the operation and maintenance personnel of Openstack (cloud platform) through Alertmanager (push service); and / or, through Alertmanager (push service), the abnormal alarm information of Rabbitmq (message middleware) is pushed to Openstack (cloud platform) using Websocket, and the alarm module of Openstack (cloud platform) itself feeds back the abnormal alarm information of Rabbitmq (message middleware) of Openstack (cloud platform) to the front end based on Websocket.

[0090] The embodiment of this specification provides an abnormal monitoring method for a cloud platform. By deploying monitoring units (Rabbitmq-Exporter (metric collection component), Prometheus (monitoring service), Alertmanager (push service)) independently of Openstack (cloud platform), it realizes automatic monitoring and timely alarming of abnormalities in Rabbitmq (message middleware) in Openstack (cloud platform), without relying on manual experience, avoiding omission of abnormal information, and can timely detect and resolve abnormalities when Rabbitmq (message middleware) has an abnormality, ensuring the stable and secure operation of the entire Openstack (cloud platform). Moreover, the monitoring units are deployed and run independently of Openstack (cloud platform), and the abnormal monitoring of Rabbitmq (message middleware) has a relatively low impact on Openstack (cloud platform) itself, making other modules of Openstack (cloud platform) insensitive to the abnormal alarm information of Rabbitmq (message middleware), having better generality, and can be applied to various real production environments, improving the reliability of the entire abnormal monitoring system of Rabbitmq (message middleware).

[0091] Corresponding to the above method embodiments, this specification also provides an embodiment of a monitoring unit in an exception monitoring system of a platform. Figure 7 FIG. shows a schematic structural diagram of a monitoring unit in an exception monitoring system of a cloud platform provided by an embodiment of this specification. As Figure 7 shown, the monitoring unit is independently deployed from the cloud platform and includes: An index collection component 702, configured to collect the running information of the message middleware in the cloud platform, where the message middleware is a message transfer component of the cloud platform; A monitoring service 704, configured to determine whether an exception occurs in the message middleware in the cloud platform according to the running information and the set alarm policy of the message middleware; A push service 706, configured to push the exception alarm information of the message middleware in the cloud platform if an exception occurs in the message middleware.

[0092] Optionally, the index collection component 702 is further configured to: Collect the running information of the message middleware from the cloud platform, where the running information includes at least one of the memory size occupied by the message middleware, the process running status, and the available disk size; Push the running information to the monitoring service 704 in a set format, where the set format is a data format recognizable by the monitoring service 704.

[0093] Optionally, the set alarm policy includes at least one of the memory size exceeding the memory threshold, the process running status being abnormal, and the available disk size being lower than the disk threshold; The monitoring service 704 is further configured to perform at least one of the following: If the memory size occupied by the message middleware exceeds the memory threshold, determine that an exception occurs in the message middleware in the cloud platform; If the process running status is abnormal, determine that an exception occurs in the message middleware in the cloud platform; If the available disk size is lower than the disk threshold, determine that an exception occurs in the message middleware in the cloud platform.

[0094] Optionally, the push service 706 is further configured to: Push the exception alarm information of the message middleware to the operation and maintenance personnel of the cloud platform through a network communication method; and / or, Push the exception alarm information of the message middleware to the cloud platform, so that the cloud platform feeds back the exception alarm information of the message middleware to the front end through its own alarm module.

[0095] Optionally, the monitoring unit further includes a deployment module, configured to: Deploy a monitoring unit in the exception monitoring system of the cloud platform, and configure the setting alarm policy of the message middleware in the monitoring unit through a configuration file during the deployment process of the monitoring unit.

[0096] Optionally, the monitoring service 704 is further configured to: Aggregate the operation information through the multi-dimensional data model in the monitoring unit to obtain the operation information of at least one dimension; Determine whether there is an exception in the message middleware in the cloud platform according to the operation information of at least one dimension and the setting alarm policy corresponding to each dimension.

[0097] Optionally, the monitoring unit further includes a query module, which is configured to: Receive a monitoring metric query request, where the monitoring metric query request carries a query dimension parameter; Query and return the corresponding target operation information according to the query dimension parameter.

[0098] The embodiment of the present specification provides a monitoring unit in an exception monitoring system of a cloud platform. The monitoring unit is deployed independently of the cloud platform, realizing automatic monitoring and timely alarm of exceptions in the message middleware in the cloud platform, without relying on manual experience, avoiding omission of exception information, and can timely detect exceptions in the message middleware and solve the exceptions in time when an exception occurs, ensuring the stable and safe operation of the entire cloud platform. Moreover, the monitoring unit is deployed and runs independently of the cloud platform, and the exception monitoring of the message middleware has a low impact on the cloud platform itself, making other modules of the cloud platform insensitive to the exception alarm information of the message middleware, having better versatility, and can be applied to various real production environments, improving the reliability of the exception monitoring system of the entire cloud platform.

[0099] The above is a schematic solution of a monitoring unit in an exception monitoring system of a cloud platform in this embodiment. It should be noted that the technical solution of the monitoring unit in the exception monitoring system of the cloud platform belongs to the same concept as the above-mentioned exception monitoring method of the cloud platform and the technical solution of the exception monitoring system of the cloud platform. For the details not described in the technical solution of the monitoring unit in the exception monitoring system of the cloud platform, reference can be made to the descriptions of the above-mentioned exception monitoring method of the cloud platform and the technical solution of the exception monitoring system of the cloud platform.

[0100] Figure 8 The structural block diagram of a computing device provided according to an embodiment of the present specification is shown. The components of the computing device 800 include but are not limited to a memory 810 and a processor 820. The processor 820 is connected to the memory 810 through a bus 830, and a database 850 is used to store data.

[0101] The computing device 800 also includes an access device 840, which enables the computing device 800 to communicate via one or more networks 860. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 840 may include one or more of any type of wired or wireless network interface (e.g., a network interface controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.

[0102] In one embodiment of the present specification, the above components of the computing device 800 and Figure 8 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 8 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art may add or replace other components as needed.

[0103] The computing device 800 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 800 can also be a mobile or stationary server.

[0104] Among them, the processor 820 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above abnormal monitoring method of the cloud platform are implemented.

[0105] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above-mentioned abnormal monitoring method of the cloud platform belong to the same concept. For the detailed content not described in the technical solution of the computing device, reference can be made to the description of the technical solution of the above-mentioned abnormal monitoring method of the cloud platform.

[0106] An embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the above-mentioned abnormal monitoring method of the cloud platform are implemented.

[0107] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above-mentioned abnormal monitoring method of the cloud platform belong to the same concept. For the detailed content not described in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above-mentioned abnormal monitoring method of the cloud platform.

[0108] An embodiment of this specification also provides a computer program. When the computer program is executed on a computer, the computer is made to execute the steps of the above-mentioned abnormal monitoring method of the cloud platform.

[0109] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above-mentioned abnormal monitoring method of the cloud platform belong to the same concept. For the detailed content not described in the technical solution of the computer program, reference can be made to the description of the technical solution of the above-mentioned abnormal monitoring method of the cloud platform.

[0110] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0111] Computer instructions include computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms, etc. A computer-readable medium can include: any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0112] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, some steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential for the embodiments of this specification.

[0113] In the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0114] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not elaborate on all details and do not limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. An abnormal monitoring method for a cloud platform, characterized in that, Applied to a monitoring unit, the monitoring unit is deployed independently of the cloud platform, and the method includes: Collect the operation information of the message middleware in the cloud platform, where the message middleware is the message passing component of the cloud platform; Determine whether there is an abnormality in the message middleware in the cloud platform according to the operation information and the set alarm policy of the message middleware; If the message middleware has an abnormality, push the abnormality alarm information of the message middleware in the cloud platform.

2. The abnormal monitoring method of the cloud platform according to claim 1, wherein The monitoring unit includes a metric collection component and a monitoring service; collecting the operation information of the message middleware in the cloud platform includes: Collect the operation information of the message middleware from the cloud platform through the metric collection component, where the operation information includes at least one of the memory size occupied by the message middleware, the process running status, and the available disk size; Push the operation information to the monitoring service in a set format, where the set format is a data format recognizable by the monitoring service.

3. The abnormal monitoring method of the cloud platform according to claim 2, characterized in that, The set alarm policy includes at least one of the memory size exceeding the memory threshold, the process running status being abnormal, and the available disk size being lower than the disk threshold; Determining whether there is an abnormality in the message middleware in the cloud platform according to the operation information and the set alarm policy of the message middleware includes at least one of the following: If the memory size occupied by the message middleware exceeds the memory threshold, determine that there is an abnormality in the message middleware in the cloud platform; If the process running status is abnormal, determine that there is an abnormality in the message middleware in the cloud platform; If the available disk size is lower than the disk threshold, determine that there is an abnormality in the message middleware in the cloud platform.

4. The abnormal monitoring method of the cloud platform according to claim 1, wherein Pushing the abnormality alarm information of the message middleware in the cloud platform includes: Pushing the abnormality alarm information of the message middleware to the operation and maintenance personnel of the cloud platform through a network communication method; and / or, Pushing the abnormality alarm information of the message middleware to the cloud platform so that the cloud platform feeds back the abnormality alarm information of the message middleware to the front end through its own alarm module.

5. The abnormal monitoring method of the cloud platform according to claim 1, wherein The method further includes: Deploy the monitoring unit in the abnormality monitoring system of the cloud platform, and configure the set alarm policy of the message middleware in the monitoring unit through a configuration file during the deployment process of the monitoring unit.

6. The abnormal monitoring method of the cloud platform according to claim 1, characterized in that Determining whether there is an abnormality in the message middleware in the cloud platform according to the operation information and the set alarm policy of the message middleware includes: Aggregate the operation information through a multi-dimensional data model in the monitoring unit to obtain operation information of at least one dimension; Determine whether there is an abnormality in the message middleware in the cloud platform according to the operation information of at least one dimension and the set alarm policy corresponding to each dimension.

7. The abnormal monitoring method of the cloud platform according to claim 6, characterized in that The method further includes: Receive a monitoring metric query request, where the monitoring metric query request carries a query dimension parameter; Query and return the corresponding target operation information according to the query dimension parameter.

8. An anomaly monitoring system for a cloud platform, characterized in that, Includes a monitoring unit and a cloud platform, and the monitoring unit is deployed independently of the cloud platform; The monitoring unit is configured to collect the operation information of the message middleware in the cloud platform, where the message middleware is the message transmission component of the cloud platform; determine whether there is an abnormality in the message middleware in the cloud platform according to the operation information and the set alarm policy of the message middleware; if the message middleware has an abnormality, push the abnormality alarm information of the message middleware in the cloud platform.

9. The anomaly monitoring system of the cloud platform according to claim 8, characterized in that, The monitoring unit includes a metric collection component and a monitoring service; The metric collection component is configured to collect the operation information of the message middleware from the cloud platform, where the operation information includes at least one of the memory size occupied by the message middleware, the process running status, and the available disk size; push the operation information to the monitoring service in a set format, where the set format is a data format recognizable by the monitoring service; The monitoring service is configured to store the operation information and determine whether there is an abnormality in the message middleware in the cloud platform according to the operation information and the set alarm policy of the message middleware.

10. The abnormal monitoring system of the cloud platform according to claim 8, wherein The monitoring unit is further configured to push the abnormality alarm information of the message middleware to the cloud platform; The cloud platform is configured to store the abnormality alarm information of the message middleware in a database; feedback the abnormality alarm information of the message middleware to the front end through its own alarm module.

11. The anomaly monitoring system of the cloud platform according to claim 8, characterized in that, The cloud platform is further configured to: In response to a management instruction, display the operation and maintenance interface of the cloud platform and display the abnormality alarm information of the message middleware in the operation and maintenance interface.

12. A computing device, characterized in that, including: a memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the abnormal monitoring method of the cloud platform according to any one of claims 1-7 are implemented.

13. A computer-readable storage medium, characterized in that, It stores computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the abnormal monitoring method of the cloud platform according to any one of claims 1-7 are implemented.

14. A computer program product, characterized in that, including a computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the abnormal monitoring method of the cloud platform according to any one of claims 1-7 are implemented.