System monitoring management method and device, electronic equipment and nonvolatile storage medium

By combining the chat operation and maintenance robot and service management framework, using the service level agreement to generate alarm information and automatically create work orders, the problems of slow failure response and insufficient operation and maintenance transparency in the monitoring service system are solved, and efficient operation and maintenance management is achieved.

CN120474892APending Publication Date: 2025-08-12CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510519117.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing monitoring service system lacks unified management and standardized service processes, resulting in slower failure response, low operation and maintenance efficiency, and insufficient transparency in operation and maintenance work.

Method used

By combining the chat operation and maintenance robot and service management framework, the service level agreement is used to generate alarm information, automatically create work orders, and obtain control instructions through the front-end interactive interface, efficient push and processing of alarm information is achieved.

Benefits of technology

It improves the overall efficiency and service quality of IT operations and maintenance, and realizes timely response to failures and transparent operation and maintenance processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120474892A_ABST
    Figure CN120474892A_ABST
Patent Text Reader

Abstract

The invention discloses a system monitoring management method and device, electronic equipment and a nonvolatile storage medium. The method comprises the following steps: acquiring monitoring data in a target system, and generating alarm information under the condition that the monitoring data triggers an alarm condition; in response to the alarm information, a target process instance is created, the target process instance is used for generating a target work order, and the target work order is used for processing an alarm event corresponding to the alarm information; and pushing the alarm information and / or the processing state information of the target work order to a front-end interaction interface of the terminal equipment of the target object by adopting a chat operation and maintenance robot, and obtaining a control instruction triggered in the front-end interaction interface for execution, the control instruction is triggered when the target object inputs a message on the front-end interaction interface in a dialogue mode with the chat operation and maintenance robot. According to the method and the device, the technical problems that a monitoring service operation and maintenance technology in related technologies is relatively slow in fault response and the transparency of operation and maintenance work is insufficient are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of monitoring service technology, and in particular to a system monitoring and management method, device, electronic device and non-volatile storage medium. Background Art

[0002] With the accelerated construction of next-generation cloud network operation systems, the number and scale of IT systems are also increasing, leading to a continuous increase in the complexity of IT operation and maintenance service management. Monitoring service systems in related technologies (e.g., Prometheus) lack unified management and standardized service processes for equipment operation and maintenance, while information technology service management tools in related technologies (e.g., ITIL (Information Technology Infrastructure Library)) still have shortcomings in automated supervision and fault prevention. As a result, monitoring service operation and maintenance technologies in related technologies suffer from technical issues such as slow response to faults, low operation and maintenance efficiency, and insufficient transparency in operation and maintenance work.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] The embodiments of the present application provide a system monitoring and management method, device, electronic device and non-volatile storage medium to at least solve the technical problems in the related art of monitoring service operation and maintenance technology, such as slow response to faults and insufficient transparency of operation and maintenance work.

[0005] According to one aspect of an embodiment of the present application, a system monitoring and management method is provided, including: obtaining monitoring data in a target system, and generating alarm information when the monitoring data triggers an alarm condition, wherein the alarm condition is determined based on a service level agreement corresponding to the target system; creating a target process instance in response to the alarm information, wherein the target process instance is used to generate a target work order, and the target work order is used to process an alarm event corresponding to the alarm information; using a chat operation and maintenance robot to push the alarm information and / or the processing status information of the target work order to the front-end interactive interface of the terminal device of the target object, and obtaining a control instruction triggered in the front-end interactive interface for execution, wherein the control instruction is triggered when the target object inputs a message on the front-end interactive interface in the form of a dialogue with the chat operation and maintenance robot.

[0006] Optionally, obtaining the monitoring data in the target system includes: in the case where there is no network isolation between the monitoring module and the target system, using the monitoring module to pull the monitoring data in the target system through the service interface of the target system, wherein the service interface is implemented by an exporter set in the target system; in the case where there is network isolation between the monitoring module and the target system, using the monitoring module to pull the monitoring data from the push gateway, wherein the monitoring data in the push gateway is pushed to the push gateway by the target system.

[0007] Optionally, when the monitoring data triggers an alarm condition, generating alarm information includes: using the service management framework operation and maintenance module to obtain the service level agreement corresponding to the target system, wherein the service level agreement is used to represent the service standards agreed upon between the customer and the service provider, and the service level agreement includes multiple specified thresholds of monitoring indicators for quantifying the operating status of the target system; using the monitoring module to determine the alarm condition based on the service level agreement, wherein the alarm condition includes the alarm thresholds of multiple monitoring indicators, and the alarm threshold is determined based on the specified thresholds of the monitoring indicators in the service level agreement; when the value of the monitoring indicator corresponding to the monitoring data exceeds the alarm threshold, generating alarm information.

[0008] Optionally, in response to the alarm information, creating a target process instance includes: using the service management framework operation and maintenance module to obtain the alarm information in the web page display module, wherein the alarm information is sent to the web page display module by the monitoring module; determining the alarm type and alarm priority corresponding to the alarm information, wherein the alarm priority is used to characterize the urgency of the alarm; creating a target process instance corresponding to the alarm type and alarm priority in accordance with the alarm processing rules specified in the service level agreement to generate a target work order, wherein the target work order is used to indicate the processing flow of the alarm information and the processing requirement parameters corresponding to each link in the processing flow, wherein the processing requirement parameters include at least one of the following: response time, processing time.

[0009] Optionally, obtaining the control instructions triggered in the front-end interactive interface for execution includes: using a chat operation and maintenance robot to obtain the message content input by the target object in the front-end interactive interface; determining the customer needs corresponding to the message content by performing semantic analysis on the message content, and determining the control instructions corresponding to the customer needs, wherein the control instructions are used to achieve customer needs; judging whether the target object of the input message content has the authority to issue control instructions; if the target object has the authority to issue control instructions, executing the control instructions, and sending the execution results returned after executing the control instructions to the front-end interactive interface in the form of a dialogue message.

[0010] Optionally, the method also includes: when the information planned to be sent to the front-end interactive interface is a pending message that requires approval, determining the approval person in charge corresponding to the pending message, and sending the pending message to the front-end interactive interface of the approval person in charge separately; when receiving the message content of the approval person in charge's reply to the pending message, determining the approval result corresponding to the pending message based on the message content, and sending the approval result in a group to the front-end interactive interfaces of the terminal devices of all target objects related to the pending message.

[0011] Optionally, the method also includes: obtaining monitoring data of the target system within a first time period; using a fault risk prediction model to analyze the monitoring data within the first time period to determine the failure risk probability of the target system failing within a second time period, wherein the second time period is a time period immediately after the first time period, and the fault risk prediction model is trained based on a training data set, and the training data set includes: historical monitoring data corresponding to historical failures; when the failure risk probability exceeds a preset probability threshold, an alarm message is generated.

[0012] According to another aspect of an embodiment of the present application, a system monitoring and management device is also provided, including: a data acquisition module, used to acquire monitoring data in a target system, and generate alarm information when the monitoring data triggers an alarm condition, wherein the alarm condition is determined based on a service level agreement corresponding to the target system; a work order generation module, used to create a target process instance in response to the alarm information, wherein the target process instance is used to generate a target work order, and the target work order is used to process an alarm event corresponding to the alarm information; a chat interaction module, used to use a chat operation and maintenance robot to push the alarm information and / or the processing status information of the target work order to the front-end interactive interface of the terminal device of the target object, and obtain the control instruction triggered in the front-end interactive interface for execution, wherein the control instruction is triggered when the target object inputs a message on the front-end interactive interface in the form of a dialogue with the chat operation and maintenance robot.

[0013] According to another aspect of the embodiments of the present application, an electronic device is provided, including: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the system monitoring and management method is executed when the program is running.

[0014] According to another aspect of the embodiments of the present application, a non-volatile storage medium is provided. The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the system monitoring and management method by running the computer program.

[0015] According to another aspect of the embodiments of the present application, a computer program product is provided, including a computer program, which implements the steps of the system monitoring and management method when executed by a processor.

[0016] In an embodiment of the present application, monitoring data from a target system is acquired, and when the monitoring data triggers an alarm condition, an alarm message is generated, wherein the alarm condition is determined based on a service level agreement corresponding to the target system; in response to the alarm message, a target process instance is created, wherein the target process instance is used to generate a target work order, and the target work order is used to handle the alarm event corresponding to the alarm message; a chat operation and maintenance robot is used to push the alarm message and / or the processing status information of the target work order to a front-end interactive interface of a terminal device of the target object, and obtain a control instruction triggered in the front-end interactive interface for execution, wherein the control instruction is triggered when the target object enters a message on the front-end interactive interface in the form of a dialogue with the chat operation and maintenance robot. By combining technology and services, the operation and maintenance management process and modern collaboration tools are utilized to optimize the alarm mechanism and incident handling process, thereby achieving the purpose of timely generating alarms and automatically generating incident work orders when IT (Information Technology) equipment fails, and using the chat operation and maintenance robot for efficient and accurate push, thereby improving the overall efficiency and service quality of IT operation and maintenance, thereby solving the technical problems of slow response to faults and insufficient transparency of operation and maintenance work in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 This is a hardware structure block diagram of a computer terminal (or electronic device) for implementing a system monitoring and management method provided in an embodiment of the present application;

[0019] Figure 2 This is a schematic diagram of a method flow for system monitoring and management provided according to an embodiment of the present application;

[0020] Figure 3 This is a schematic diagram of a system workflow of a monitoring service system provided according to an embodiment of the present application;

[0021] Figure 4 is a schematic diagram of different methods of obtaining monitoring data provided according to an embodiment of the present application;

[0022] Figure 5 This is a schematic diagram of an event processing flow of a work order provided according to an embodiment of the present application;

[0023] Figure 6 This is a schematic diagram of the overall design architecture of a chat operation and maintenance robot provided according to an embodiment of the present application;

[0024] Figure 7 This is a schematic diagram of an approval process provided according to an embodiment of the present application;

[0025] Figure 8 It is a structural diagram of a system monitoring and management device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0028] To facilitate those skilled in the art to better understand the embodiments of the present application, some technical terms or nouns involved in the embodiments of the present application are explained as follows:

[0029] ITIL (Information Technology Infrastructure Library): is a best practice framework for IT Service Management (ITSM). It provides a comprehensive and systematic knowledge base designed to guide and support organizations in best practices in designing, delivering, operating, and improving IT services.

[0030] Prometheus: is an open-source monitoring service system and time-series database. It provides a common data model and fast data collection, storage, and query interfaces. Its core component, the Prometheus server, regularly pulls data from statically configured monitoring targets or from self-targets automatically configured based on service discovery. When the newly pulled data exceeds the configured memory cache, the data is persisted to the storage device.

[0031] An SLA (Service-Level Agreement) is a key means of ensuring the quality and availability of enterprise services. An SLA is a written contract that defines the service standards and expectations between a service provider and its customers. SLA monitoring is the process of tracking, evaluating, and reporting on these service standards to ensure the service provider is meeting the agreed-upon service levels.

[0032] ChatOps is a real-time chat-driven operations model that allows teams to collaborate and manage infrastructure, code, and data through chat rooms. By integrating bots into chat sessions, it creates an automated and transparent connection between humans, machines, and data, enabling operations teams to efficiently execute tasks and communicate and collaborate.

[0033] With the accelerated construction of next-generation cloud network operating systems, the number and scale of IT systems are also increasing, leading to a continuous increase in the complexity of IT operations and maintenance service management. This has, to a certain extent, increased the complexity of IT operations and maintenance. Operations and maintenance teams are investing excessive time and energy in handling a large number of simple and repetitive tasks, which is not conducive to optimal resource allocation. The traditional reactive "firefighting" operations and maintenance model, with its low efficiency, high recurring failure rate, and difficult-to-quantify workload for operations and maintenance staff, is no longer able to meet the new demands of today's IT operations and maintenance.

[0034] Among related technologies, the open source monitoring system Promethues, relying on a cloud-native architecture, continues to become a mainstream choice for building monitoring system platform solutions. Its powerful functions are suitable for monitoring and alerting distributed systems of various sizes; ITIL, the core content of its 3.0 version covers four core functions and 26 key processes. These elements together constitute an efficient and practical operation and maintenance management system.

[0035] However, ITIL systems within related technologies still lack automated monitoring and fault prevention. Furthermore, the Prometheus monitoring system lacks unified management and standardized service processes for device operations and maintenance. Consequently, related technologies lack a comprehensive and efficient operations and maintenance system to meet the operational and management needs of modern enterprises.

[0036] To address the above issues, the present application provides a solution in an embodiment. By integrating ITIL's service management process with Prometheus's monitoring capabilities through a chat operation and maintenance robot (e.g., Chatops), a comprehensive and efficient IT operation and maintenance system can be constructed to better meet the operational needs of modern enterprises. This is described in detail below.

[0037] According to an embodiment of the present application, an embodiment of a method for system monitoring and management is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0038] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computer terminal (or electronic device) for implementing a system monitoring and management method. Figure 1 As shown, the computer terminal 10 (or electronic device) may include one or more (illustrated as 102a, 102b, ..., 102n in the figure) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0039] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or electronic device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0040] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the system monitoring and management method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned system monitoring and management method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0041] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0042] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or electronic device).

[0043] In the above operating environment, the embodiment of the present application provides a system monitoring and management method. Figure 2 This is a schematic diagram of a method flow for system monitoring and management provided in accordance with an embodiment of the present application. Figure 2 As shown, the method includes the following steps:

[0044] Step S202: Acquire monitoring data from the target system and generate alarm information when the monitoring data triggers an alarm condition, wherein the alarm condition is determined according to the service level agreement corresponding to the target system;

[0045] Step S204: In response to the alarm information, a target process instance is created, wherein the target process instance is used to generate a target work order, and the target work order is used to handle the alarm event corresponding to the alarm information;

[0046] In step S206, a chat operation and maintenance robot is used to push the alarm information and / or the processing status information of the target work order to the front-end interactive interface of the terminal device of the target object, and obtain the control instructions triggered in the front-end interactive interface for execution, wherein the control instructions are triggered when the target object inputs a message on the front-end interactive interface in the form of a dialogue with the chat operation and maintenance robot.

[0047] Through the above steps, by combining technology with services, and utilizing operation and maintenance management processes and modern collaboration tools to optimize the alarm mechanism and accident handling process, we achieved the goal of generating timely alarms and automatically generating accident work tickets when IT equipment failures occur, as well as using chat operation and maintenance robots to efficiently and accurately push them, thereby improving the overall efficiency and service quality of IT operation and maintenance, and thus solving the technical problems of slow fault response of monitoring service operation and maintenance technology in related technologies and insufficient transparency of operation and maintenance work.

[0048] The following further introduces the system monitoring and management method in steps S202 to S208 of the embodiment of the present application.

[0049] In an embodiment of the present application, the above-mentioned system monitoring and management method can be executed in a monitoring service system, which includes: a monitoring module (taking Prometheus as an example), a Web display module, and a service management framework operation and maintenance module (taking ITIL as an example).

[0050] Among them, the monitoring module is responsible for collecting data from the monitored equipment or system (that is, obtaining monitoring data in the target system), and displaying the collected relevant data information in the interface of the Web display module. It can also customize alarm indicators through the service level agreement (SLA) in the service management framework operation and maintenance module (ITIL operation and maintenance module), and transmit alarm information to the Web interface.

[0051] The Web display module is mainly used to receive, store and visualize monitoring data and alarm information, and then establish a channel with the service management framework operation and maintenance module to enable it to obtain alarm information.

[0052] The service management framework's operations module aggregates and organizes IT system service level agreements (SLAs), acquires real-time alert information, and calls the Create Incident Process Instance API to create process instances, thereby processing alert events and automatically generating work orders. Relevant work order and alert information can be pushed to IT system managers or project clients via a chatbot-based operations bot. These managers or project clients can also participate in the work order handling process through the chatbot, for example, querying work order status and approving risk-based actions.

[0053] The following is a further introduction to the system workflow of the monitoring service system in the embodiment of the present application: Figure 3 As shown, the details are as follows.

[0054] First, use the monitoring module (Prometheus) to obtain monitoring data from the monitored IT system (i.e., the target system). The specific steps are as follows.

[0055] In some embodiments of the present application, obtaining monitoring data in the target system includes the following steps: in the case where there is no network isolation between the monitoring module and the target system, using the monitoring module to pull the monitoring data in the target system through the service interface of the target system, wherein the service interface is implemented by an exporter set in the target system; in the case where there is network isolation between the monitoring module and the target system, using the monitoring module to pull the monitoring data from the push gateway, wherein the monitoring data in the push gateway is pushed to the push gateway by the target system.

[0056] In an embodiment of the present application, a corresponding exporter or push gateway may be deployed in the monitored target system so that the monitoring module can obtain monitoring data of the corresponding target system.

[0057] Specifically, in this embodiment, there are two ways to obtain monitoring data, such as Figure 4 As shown, when there is no network isolation between the monitoring module and the target system, the timed pull mode can be used by default to pull the targets data (i.e., monitoring data). In this embodiment, the monitoring module (Prometheus server) periodically grabs monitoring data from the specified target system. Each target system must provide an HTTP service interface for the monitoring module's timed grabbing operation. In this embodiment, the service interface can be implemented by deploying an exporter. This calling method of obtaining monitoring data through a monitoring object is called the "pull" method.

[0058] However, if the monitoring module and the target system are not in the same subnet or firewall, that is, there is network isolation between the monitoring module and the target system, the monitoring module cannot pull the targets data (i.e. monitoring data), or when monitoring business data, different data need to be aggregated. In this case, the targets data (i.e. monitoring data) can be pushed to the push gateway pushgateway, and the pushgateway collects the data uniformly. Then the monitoring module can pull the pull data from the pushgateway at a regular interval.

[0059] Through the combination of the above technical features, full coverage of IT system monitoring is achieved with the help of two acquisition methods. Monitoring data can be effectively collected and processed in various network environments, improving the adaptability and flexibility of the system.

[0060] After acquiring monitoring data, the monitoring module can aggregate the pulled monitoring data and synchronize it in the web display module for visual display. In addition, the service management framework operation and maintenance module can aggregate the service level agreement (SLA) of the target system and present it numerically in the web display module. The monitoring module can configure the conditions for triggering alarms based on the service level agreement (SLA) indicators in the service management framework operation and maintenance module, as follows.

[0061] In some embodiments of the present application, when monitoring data triggers an alarm condition, generating alarm information includes the following steps: using the service management framework operation and maintenance module to obtain the service level agreement corresponding to the target system, wherein the service level agreement is used to characterize the service standards agreed upon between the customer and the service provider, and the service level agreement includes multiple prescribed thresholds of monitoring indicators for quantifying the operating status of the target system; using the monitoring module to determine the alarm condition based on the service level agreement, wherein the alarm condition includes the alarm thresholds of multiple monitoring indicators, and the alarm threshold is determined based on the prescribed thresholds of the monitoring indicators in the service level agreement; generating alarm information when the value of the monitoring indicator corresponding to the monitoring data exceeds the alarm threshold.

[0062] Specifically, the service level agreement (SLA) in the embodiments of this application combines the service standards and expectations between the service provider and the customer with monitoring indicators, allowing the SLA to be quantified and displayed through the monitoring indicators. Conversely, the alarm thresholds defined by the monitoring indicators can also be clearly defined based on the SLA. The quantified SLA is numerically presented in the web display module, which facilitates subsequent data collection and analysis, conducts regular performance evaluations, and thus tracks service execution and establishes clear service direction.

[0063] For example, in an online payment system, the SLA might stipulate a transaction success rate of at least 99.5% and an average transaction response time of no more than 150ms. The service management framework converts these thresholds into alert conditions. When the transaction success rate in the monitored data falls below 99.5% or the average transaction response time exceeds 150ms, the system automatically triggers an alert and generates an alert message. This mechanism ensures that the system automatically identifies and reports performance degradation or failures according to pre-defined SLA standards, addressing the limitations and lags of manual monitoring and improving the automation and responsiveness of operations and maintenance.

[0064] In addition, as an optional implementation, the method also includes the following steps: obtaining monitoring data of the target system within a first time period; using a fault risk prediction model to analyze the monitoring data within the first time period to determine the failure risk probability of the target system failing within a second time period, wherein the second time period is a time period immediately after the first time period, and the fault risk prediction model is trained based on a training data set, and the training data set includes: historical monitoring data corresponding to historical failures; when the failure risk probability exceeds a preset probability threshold, an alarm message is generated.

[0065] Specifically, through a fault risk prediction model, the system can predict the future failure risk probability of a target system based on historical monitoring data and machine learning algorithms, providing early warning of failures. For example, the model might be trained based on historical monitoring data such as CPU usage, memory usage, and network latency from before a system outage over the past year. When the values of these metrics in the current system are similar to patterns seen before historical failures, the model will calculate a higher failure risk probability. When the failure risk probability exceeds a preset threshold, the system generates an alert in advance, notifying operations and maintenance personnel to take preventive measures. This prediction mechanism addresses the delayed failure response problem in traditional operations and maintenance. By providing early warning, operations and maintenance personnel can proactively troubleshoot and resolve issues, effectively preventing system failures, ensuring stable system operation, and improving user experience. In building the fault risk prediction model, in addition to using historical failure data, it can also incorporate multiple data sources such as real-time monitoring data, system configuration information, and external environmental data to further improve the accuracy and comprehensiveness of predictions, providing more precise decision support for operations and maintenance.

[0066] When monitoring data triggers an alarm condition, the detection module generates an alarm message. Furthermore, the alarm message can be visually displayed on the Web display module. In this embodiment, a channel can be established between the Web display module and the service management framework operation and maintenance module, enabling the service management framework operation and maintenance module to obtain the alarm message. The service management framework operation and maintenance module receives the alarm message in real time and can define indicators such as the response time and processing time for different work order types based on information such as the type and priority of the alarm message. The Create Incident Process Instance interface is then called to create a process instance, thereby processing the alarm event and completing the automatic work order generation function. The specific steps are as follows.

[0067] In some embodiments of the present application, in response to alarm information, creating a target process instance includes the following steps: using the service management framework operation and maintenance module to obtain the alarm information in the web page display module, wherein the alarm information is sent to the web page display module by the monitoring module; determining the alarm type and alarm priority corresponding to the alarm information, wherein the alarm priority is used to characterize the urgency of the alarm; creating a target process instance corresponding to the alarm type and alarm priority in accordance with the alarm processing rules specified in the service level agreement to generate a target work order, wherein the target work order is used to indicate the processing flow of the alarm information and the processing requirement parameters corresponding to each link in the processing flow, wherein the processing requirement parameters include at least one of the following: response time, processing time.

[0068] Specifically, after receiving an alert, the service management framework's operations and maintenance module automatically creates a target process instance and a target work order based on the alert type and priority, in accordance with the alert handling rules specified in the SLA. This process includes identifying, recording, and classifying events. High-priority alerts require a rapid response. When an alert triggers an automated rule, the service management tool generates a work order based on the alert information, including necessary details such as an event description, scope of impact, and urgency.

[0069] For example, for high-priority alerts (such as system downtime), the SLA may require a response within 5 minutes of the alert being triggered and a resolution within 1 hour. Based on these requirements, the service management framework's operations module creates corresponding work tickets, assigns them to specific operations teams, and sets response and resolution deadlines to ensure that the issue is promptly addressed and resolved. This SLA-based work ticket processing process resolves issues such as unclear responsibilities and non-standardized processes in troubleshooting, improves the efficiency and quality of troubleshooting, and ensures stable system operation.

[0070] For example, the specific processing flow after the work order is generated can be as follows Figure 5 As shown, in the first step, the system generates a work order and dispatches it for processing, while notifying the second-line operation and maintenance engineer; in the second step, the second-line operation and maintenance engineer accepts the order and immediately assesses the impact scope based on the fault level. General faults are handled by the second-line operation and maintenance engineer, while major faults are immediately escalated to the third-line engineer; in the third step, the third-line operation and maintenance expert activates the emergency plan. If the problem is not resolved within 15 minutes, the operation and maintenance expert intervenes and coordinates with the manufacturer's engineers to jointly resolve the fault; in the fourth step, after the fault is eliminated, the operation and maintenance team is responsible for summarizing the fault handling results and compiling them into a fault replay report, which is then emailed back within 24 hours.

[0071] Furthermore, in this embodiment of the present application, the service management framework operation and maintenance module can also deploy a chatops-based chatbot for interactive use. This bot can push alerts and work order information to IT system managers or project clients. IT system managers or project clients can also participate in the work order handling process by interacting with the chatbot, for example, querying work order status and approving risk-based operations. Details are as follows.

[0072] In some embodiments of the present application, obtaining the control instructions triggered in the front-end interactive interface for execution includes the following steps: using a chat operation and maintenance robot to obtain the message content input by the target object in the front-end interactive interface; determining the customer needs corresponding to the message content by performing semantic analysis on the message content, and determining the control instructions corresponding to the customer needs, wherein the control instructions are used to achieve customer needs; judging whether the target object of the input message content has the authority to issue control instructions; if the target object has the authority to issue control instructions, executing the control instructions, and sending the execution results returned after executing the control instructions to the front-end interactive interface in the form of a dialogue message.

[0073] Specifically, the chat operation and maintenance robot is responsible for interacting with the IT system manager or the client side of the project. The overall architecture is as follows: Figure 6 As shown, the chatbot uses semantic analysis technology to understand the natural language messages entered by operators in the front-end interactive interface and convert them into specific control commands, thus enabling natural language-based operations. For example, operators can trigger the robot to perform the corresponding system operation by entering a command such as "restart the database server." Before executing the command, the robot verifies that the operator has the appropriate permissions to ensure operational safety. This approach reduces operational difficulty for operators and improves response speed and accuracy. The introduction of the chatbot eliminates the complexity and error-proneness of command-line operations in traditional operations, improving the intelligence and efficiency of operations.

[0074] In an embodiment of the present application, during the information collection phase, the Web display module provides a standard http interface. According to the interface specification, the chat operation and maintenance robot can access the interface through Python's requests.get or requests.post method to obtain monitoring indicators and alarms of the IT system in the Web display module; during the data processing phase, it is necessary to fully clarify customer needs and data presentation formats; during the process design phase, based on the new generation of atomic capability registration and process engine capabilities, the approval process based on the chat operation and maintenance robot is orchestrated and designed. The process mainly involves 6 modules, including entering the process, obtaining approval information, pushing customer approval through single chat, group notification of approval results, connecting to the SMS interface to send approval results, and ending the process.

[0075] Among them, the approval process based on the chat operation and maintenance robot is as follows.

[0076] In some embodiments of the present application, the method also includes the following steps: when the information planned to be sent to the front-end interactive interface is a pending message that requires approval, determining the person in charge of approval corresponding to the pending message, and sending the pending message separately to the front-end interactive interface of the terminal device of the person in charge of approval; when receiving the message content of the approval person in charge's reply to the pending message, determining the approval result corresponding to the pending message based on the message content, and sending the approval result in a group to the front-end interactive interfaces of the terminal devices of all target objects related to the pending message.

[0077] In addition, in an embodiment of the present application, the service management framework operation and maintenance module can also collect resource utilization and performance data of the detection module through the asset management and configuration management process, process the collected resource utilization and performance data, and optimize resource management and system.

[0078] The embodiment of the present application realizes the creation of automated alarm and work order processing processes by combining service level agreements with system monitoring data, thereby improving operation and maintenance efficiency. The introduction of the chat operation and maintenance robot enables operation and maintenance personnel to query and control the system status in a natural language manner, reducing the difficulty of operation and improving the response speed. In addition, through the fault risk prediction model, it is possible to issue an early warning before a fault occurs and take measures in advance, effectively preventing system failures, ensuring the stable operation of the system, and improving the user experience. In the approval process, the person in charge of approval is notified separately and the approval results are sent in groups, which simplifies the approval process and improves the approval efficiency. Overall, this solution significantly improves the intelligence level of system monitoring and management, and brings substantial improvements and efficiency improvements to operation and maintenance work.

[0079] In order to make the above method process easier to understand, the following is a further example of the above method process combined with a complete project case.

[0080] In this embodiment, the project involves Gigabit Internet egress ports, circuits, a 10G secondary backbone network, and VPDN (Virtual Private Dial-up Network) services. Therefore, the monitoring service system needs to focus on the operating status of the backbone network, the power supply status of transmission equipment, the stable operation indicators of the Gigabit Internet egress, VPDN certification status, and the stability, performance, and security indicators of conventional IT systems.

[0081] Specifically, monitoring is performed using the monitoring module, with monitoring indicators exposed via HTTP pull or pulled via the deployed exporter. Some of this project's business resources are deployed on the customer's intranet, while the monitoring module deployed by the monitoring service system is on the public network. Since the two networks are disconnected, the pull method using the monitoring module is incapable of acquiring monitoring data. Therefore, data acquisition is performed by deploying a pushgateway. The pushgateway acts as a bridge, connecting the monitoring module service and business resources. Business resources push data to the pushgateway, and the monitoring module then periodically pulls data from the pushgateway.

[0082] The project's service level agreement (SLA) determines the metrics that need to be monitored, such as simulated dialing task indicators, dedicated line traffic, Internet egress packet loss rate, Internet broadband latency, and the availability, throughput, and response time of the IT system itself. The SLA also determines the thresholds for these metrics, thereby determining the conditions for triggering alarms.

[0083] The service management framework operation and maintenance module of the monitoring service system deploys a chatops-based chat operation and maintenance robot that can be used for interaction. The robot obtains data through standard interfaces during the information collection phase, such as the total number of VPDN simulation dial tests configured for the project, the number of successes, the number of failures, etc. From this, the success rate of the VPDN certification of the project according to the time dimension can be calculated; the data processing phase requires full clarification of customer needs and the form of data presentation. For example, the province's backbone network alarms need to be classified by matching identifiers to form statistical icon outputs; the power supply guarantee for the computer room needs to extract the AC three-phase voltage, rectifier output voltage, rectifier output current, and battery pack voltage indicators, and compare them with the normal threshold value, output abnormal indicators, and alert the project leader to pay attention; the approval scenario is designed during the process design phase, specifically Figure 7 shown.

[0084] Step 1: Enter the process. You need to configure the conditions that trigger entering the process. This function can be activated by setting a password for interacting with the bot. Step 2: Obtain approval information. Leverage the encapsulation capabilities of conversation atoms for interactive conversation and information acquisition. Step 3: Push the customer approval via one-on-one chat. After field processing, configure the one-on-one chat bot to push the relevant approval notification to the customer. This design avoids interference from group messages. This node can be implemented using a form and a button. Step 4: Notify the group of the approval result. After the customer's approval is completed, a message will be broadcast within the group to alert the applicant. Step 5: Connect to the SMS interface to distribute the approval result. Configure the SMS interface and call it via POST. The interface will return the call status. After the call, the approval content and results can be distributed to the customer or system project manager. Step 6: End the process. The process ends normally, and a timer is set for each step to prevent the process from being stuck.

[0085] This application solution creates a scalable monitoring system architecture that significantly reduces incident response time, improves team collaboration efficiency and overall transparency through real-time communication and automated operations; automates work order generation and allocation, reduces manual intervention, and uses chat operation and maintenance robots to simplify the incident management process, thereby improving work efficiency; the open API design of the chat operation and maintenance robot allows the system to easily design operation and maintenance function processes to meet diverse needs and enhance system flexibility; the service management framework (ITIL) process ensures the reliability of the system, and permission control and security measures ensure the security of monitoring data.

[0086] According to an embodiment of the present application, an embodiment of a system monitoring and management device is also provided. Figure 8 This is a schematic diagram of the structure of a system monitoring and management device provided according to an embodiment of the present application. Figure 8 As shown, the device includes:

[0087] A data acquisition module 80 is configured to acquire monitoring data from the target system and generate an alarm message when the monitoring data triggers an alarm condition, wherein the alarm condition is determined based on the service level agreement corresponding to the target system;

[0088] A work order generation module 82 is configured to create a target process instance in response to the alarm information, wherein the target process instance is used to generate a target work order, and the target work order is used to process the alarm event corresponding to the alarm information;

[0089] The chat interaction module 84 is used to use the chat operation and maintenance robot to push the alarm information and / or the processing status information of the target work order to the front-end interaction interface of the terminal device of the target object, and obtain the control instructions triggered in the front-end interaction interface for execution, wherein the control instructions are triggered when the target object enters a message on the front-end interaction interface in the form of a dialogue with the chat operation and maintenance robot.

[0090] Optionally, obtaining the monitoring data in the target system includes: in the case where there is no network isolation between the monitoring module and the target system, using the monitoring module to pull the monitoring data in the target system through the service interface of the target system, wherein the service interface is implemented by an exporter set in the target system; in the case where there is network isolation between the monitoring module and the target system, using the monitoring module to pull the monitoring data from the push gateway, wherein the monitoring data in the push gateway is pushed to the push gateway by the target system.

[0091] Optionally, when the monitoring data triggers an alarm condition, generating alarm information includes: using the service management framework operation and maintenance module to obtain the service level agreement corresponding to the target system, wherein the service level agreement is used to represent the service standards agreed upon between the customer and the service provider, and the service level agreement includes multiple specified thresholds of monitoring indicators for quantifying the operating status of the target system; using the monitoring module to determine the alarm condition based on the service level agreement, wherein the alarm condition includes the alarm thresholds of multiple monitoring indicators, and the alarm threshold is determined based on the specified thresholds of the monitoring indicators in the service level agreement; when the value of the monitoring indicator corresponding to the monitoring data exceeds the alarm threshold, generating alarm information.

[0092] Optionally, in response to the alarm information, creating a target process instance includes: using the service management framework operation and maintenance module to obtain the alarm information in the web page display module, wherein the alarm information is sent to the web page display module by the monitoring module; determining the alarm type and alarm priority corresponding to the alarm information, wherein the alarm priority is used to characterize the urgency of the alarm; creating a target process instance corresponding to the alarm type and alarm priority in accordance with the alarm processing rules specified in the service level agreement to generate a target work order, wherein the target work order is used to indicate the processing flow of the alarm information and the processing requirement parameters corresponding to each link in the processing flow, wherein the processing requirement parameters include at least one of the following: response time, processing time.

[0093] Optionally, obtaining the control instructions triggered in the front-end interactive interface for execution includes: using a chat operation and maintenance robot to obtain the message content input by the target object in the front-end interactive interface; determining the customer needs corresponding to the message content by performing semantic analysis on the message content, and determining the control instructions corresponding to the customer needs, wherein the control instructions are used to achieve customer needs; judging whether the target object of the input message content has the authority to issue control instructions; if the target object has the authority to issue control instructions, executing the control instructions, and sending the execution results returned after executing the control instructions to the front-end interactive interface in the form of a dialogue message.

[0094] Optionally, the chat interaction module 84 is also used to: when the information planned to be sent to the front-end interactive interface is a pending message that requires approval, determine the person in charge of approval corresponding to the pending message, and send the pending message to the front-end interactive interface of the terminal device of the person in charge of approval separately; when receiving the message content of the approval person in charge's reply to the pending message, determine the approval result corresponding to the pending message based on the message content, and send the approval result in a group to the front-end interactive interfaces of the terminal devices of all target objects related to the pending message.

[0095] Optionally, the system monitoring and management device is also used to: obtain monitoring data of the target system within a first time period; use a fault risk prediction model to analyze the monitoring data within the first time period to determine the failure risk probability of the target system failing within a second time period, wherein the second time period is a time period immediately after the first time period, and the fault risk prediction model is trained based on a training data set, and the training data set includes: historical monitoring data corresponding to historical failures; when the failure risk probability exceeds a preset probability threshold, an alarm message is generated.

[0096] It should be noted that the various modules in the above-mentioned system monitoring and management device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.

[0097] It should be noted that the system monitoring and management device provided in this embodiment can be used to perform Figure 2 The system monitoring and management method shown, therefore, the relevant explanations and descriptions of the above-mentioned system monitoring and management method are also applicable to the embodiments of this application and will not be repeated here.

[0098] An embodiment of the present application also provides a non-volatile storage medium, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the following system monitoring and management method by running the computer program: obtaining monitoring data in the target system, and generating alarm information when the monitoring data triggers an alarm condition, wherein the alarm condition is determined based on the service level agreement corresponding to the target system; creating a target process instance in response to the alarm information, wherein the target process instance is used to generate a target work order, and the target work order is used to process the alarm event corresponding to the alarm information; using a chat operation and maintenance robot, the alarm information and / or the processing status information of the target work order is pushed to the front-end interactive interface of the terminal device of the target object, and the control instruction triggered in the front-end interactive interface is obtained for execution, wherein the control instruction is triggered when the target object enters a message on the front-end interactive interface in the form of a dialogue with the chat operation and maintenance robot.

[0099] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of the system monitoring and management method described in each embodiment of the present application: obtaining monitoring data in the target system, and generating alarm information when the monitoring data triggers an alarm condition, wherein the alarm condition is determined based on the service level agreement corresponding to the target system; creating a target process instance in response to the alarm information, wherein the target process instance is used to generate a target work order, and the target work order is used to process the alarm event corresponding to the alarm information; using a chat operation and maintenance robot, pushing the alarm information and / or the processing status information of the target work order to the front-end interactive interface of the terminal device of the target object, and obtaining the control instructions triggered in the front-end interactive interface for execution, wherein the control instructions are triggered when the target object enters a message on the front-end interactive interface in the form of a dialogue with the chat operation and maintenance robot.

[0100] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0101] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0102] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0103] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0104] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0105] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0106] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A system monitoring and management method, characterized in that: include: Acquiring monitoring data from a target system and generating an alarm message when the monitoring data triggers an alarm condition, wherein the alarm condition is determined based on a service level agreement corresponding to the target system; In response to the alarm information, creating a target process instance, wherein the target process instance is used to generate a target work order, and the target work order is used to handle the alarm event corresponding to the alarm information; A chat operation and maintenance robot is used to push the alarm information and / or the processing status information of the target work order to the front-end interactive interface of the terminal device of the target object, and obtain the control instructions triggered in the front-end interactive interface for execution, wherein the control instructions are triggered when the target object enters a message on the front-end interactive interface in the form of a dialogue with the chat operation and maintenance robot.

2. The system monitoring and management method according to claim 1, characterized in that: Obtaining monitoring data from the target system includes: In the absence of network isolation between the monitoring module and the target system, using the monitoring module to pull the monitoring data in the target system through a service interface of the target system, wherein the service interface is implemented by an exporter provided in the target system; In the case where there is network isolation between the monitoring module and the target system, the monitoring module is used to pull the monitoring data from the push gateway, wherein the monitoring data in the push gateway is pushed to the push gateway by the target system.

3. The system monitoring and management method according to claim 1, characterized in that: When the monitoring data triggers an alarm condition, generating alarm information includes: Using a service management framework operation and maintenance module, obtaining the service level agreement corresponding to the target system, wherein the service level agreement is used to represent the service standards agreed upon between the customer and the service provider, and the service level agreement includes a plurality of prescribed thresholds for monitoring indicators used to quantify the operating status of the target system; Using a monitoring module, determining the alarm condition according to the service level agreement, wherein the alarm condition includes alarm thresholds of a plurality of the monitoring indicators, and the alarm thresholds are determined according to the prescribed thresholds of the monitoring indicators in the service level agreement; When the value of the monitoring indicator corresponding to the monitoring data exceeds the alarm threshold, the alarm information is generated.

4. The system monitoring and management method according to claim 3, characterized in that: In response to the warning information, creating a target process instance includes: Using the service management framework operation and maintenance module, obtaining the alarm information in the web page display module, wherein the alarm information is sent to the web page display module by the monitoring module; Determining an alarm type and an alarm priority corresponding to the alarm information, wherein the alarm priority is used to characterize the urgency of the alarm; In accordance with the alarm processing rules specified in the service level agreement, the target process instance corresponding to the alarm type and the alarm priority is created to generate the target work order, wherein the target work order is used to indicate the processing flow of the alarm information and the processing requirement parameters corresponding to each link in the processing flow, wherein the processing requirement parameters include at least one of the following: response time and processing time.

5. The system monitoring and management method according to claim 1, characterized in that: Obtaining the control instruction triggered in the front-end interactive interface for execution includes: Using the chat operation and maintenance robot, obtain the message content input by the target object in the front-end interactive interface; Determining customer needs corresponding to the message content by performing semantic analysis on the message content, and determining control instructions corresponding to the customer needs, wherein the control instructions are used to achieve the customer needs; Determining whether the target object that inputs the message content has the authority to issue the control instruction; In the case where the target object has the authority to issue the control instruction, the control instruction is executed, and the execution result returned after executing the control instruction is sent to the front-end interactive interface in the form of a dialogue message.

6. The system monitoring and management method according to claim 1, characterized in that: The method further comprises: In the case where the information planned to be sent to the front-end interactive interface is a pending approval message that requires approval, determining the person in charge of approval corresponding to the pending approval message, and sending the pending approval message separately to the front-end interactive interface of the terminal device of the person in charge of approval; Upon receiving the message content of the reply from the approval person in charge to the message to be approved, the approval result corresponding to the message to be approved is determined based on the message content, and the approval result is sent in a group to the front-end interactive interface of the terminal devices of all target objects related to the message to be approved.

7. The system monitoring and management method according to claim 1, characterized in that: The method further comprises: Acquiring the monitoring data of the target system within a first time period; Using a fault risk prediction model, by analyzing the monitoring data within the first time period, determining a fault risk probability of the target system failing within a second time period, wherein the second time period is a time period immediately after the first time period, and the fault risk prediction model is trained based on a training data set, the training data set including: historical monitoring data corresponding to historical faults; When the failure risk probability exceeds a preset probability threshold, the alarm information is generated.

8. A system monitoring and management device, characterized in that: include: a data acquisition module, configured to acquire monitoring data from a target system and generate an alarm message when the monitoring data triggers an alarm condition, wherein the alarm condition is determined according to a service level agreement corresponding to the target system; a work order generation module, configured to create a target process instance in response to the alarm information, wherein the target process instance is used to generate a target work order, and the target work order is used to process the alarm event corresponding to the alarm information; A chat interaction module is used to use a chat operation and maintenance robot to push the alarm information and / or the processing status information of the target work order to the front-end interaction interface of the terminal device of the target object, and obtain the control instructions triggered in the front-end interaction interface for execution, wherein the control instructions are triggered when the target object enters a message on the front-end interaction interface in the form of a dialogue with the chat operation and maintenance robot.

9. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the system monitoring and management method according to any one of claims 1 to 7 is executed when the program is run.

10. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the system monitoring and management method according to any one of claims 1 to 7 by running the computer program.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the system monitoring and management method according to any one of claims 1 to 7 are implemented.