Operation and maintenance monitoring method and device based on business indexes
By acquiring and separating business indicator data and building alarm rules and topology structures, the problem of traditional operation and maintenance monitoring being unable to integrate monitoring at all levels is solved, and efficient troubleshooting and performance optimization of complex systems are achieved.
Patent Information
- Application Number
- CN202510814204.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-05
AI Technical Summary
Traditional operation and maintenance monitoring methods lack business monitoring and are unable to integrate monitoring at all levels, making it difficult to fully understand the business operation status and quickly locate the root cause of failures in complex systems.
By obtaining business indicator data from multiple nodes, performing hot and cold separation and calculating health indicators, alarm rules and topology structures are constructed, and operation and maintenance monitoring is performed using the topology structure configured with alarm rules.
It achieves efficient troubleshooting and performance optimization of large amounts of business indicator data in complex operation and maintenance environments, and provides a fast and accurate means of fault location.
Smart Images

Figure CN120602373A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of operation and maintenance monitoring technology, and in particular to an operation and maintenance monitoring method and device based on business indicators. Background Art
[0002] In internet business scenarios, system stability and availability are crucial, and any failure can have a serious impact on the business. In such an environment, due to the large scale of business and high system complexity, when failures occur, they often face the challenge of slow troubleshooting efficiency.
[0003] In existing technologies, traditional operation and maintenance monitoring methods usually deploy and monitor the system layer, component layer, application layer, etc. in a layered manner. This layered deployment method will lead to a lack of business monitoring and the inability to integrate the monitoring of each layer, making it difficult for operation and maintenance personnel to fully understand the business and system operation status. As the scale of business grows, the system architecture will also become more complex. The existing production scenario troubleshooting uses a bottom-up approach, starting from the basic components at the bottom of the system and gradually conducting decentralized inspections to locate the root cause of the problem. This cannot meet the needs of comprehensive monitoring and rapid problem location of complex systems. Therefore, there is an urgent need for an operation and maintenance monitoring method based on business indicators to provide efficient troubleshooting and performance optimization means for complex operation and maintenance environments. Summary of the Invention
[0004] Given that traditional operation and maintenance monitoring methods typically deploy and monitor the system layer, component layer, application layer, etc. in a layered deployment approach, this layered deployment approach results in a lack of business monitoring and an inability to integrate and correlate monitoring at each level, making it difficult for operation and maintenance personnel to fully understand the business and system operating status. As the scale of business grows, the system architecture will also become more complex. Existing production scenario troubleshooting uses a bottom-up approach, starting from the basic components at the bottom of the system and gradually conducting decentralized inspections to locate the root cause of the problem. This cannot meet the needs of comprehensive monitoring and rapid problem location for complex systems. The embodiments of this specification provide an operation and maintenance monitoring method and device based on business indicators to overcome the above-mentioned problems or at least partially solve the above-mentioned problems.
[0005] On the one hand, some embodiments of this specification aim to provide an operation and maintenance monitoring method based on business indicators, the method comprising:
[0006] Acquire first business indicator data from multiple nodes and transmit it to a monitoring node;
[0007] Performing hot and cold separation on the first business indicator data to obtain second business indicator data and third business indicator data, and respectively calculating business health indicators corresponding to the second business indicator data and the third business indicator data;
[0008] Acquire historical fault information, and construct an alarm rule based on the historical fault information and a business health indicator corresponding to the second business indicator data;
[0009] A topology structure corresponding to the type of the first business indicator data is generated, and corresponding alarm rules are configured for the topology structure, so as to perform operation and maintenance monitoring on the third business indicator data using the topology structure configured with the alarm rules.
[0010] Furthermore, the first business indicator data is subjected to hot and cold separation to obtain second business indicator data and third business indicator data, and business health indicators corresponding to the second business indicator data and the third business indicator data are calculated respectively, including:
[0011] performing hot and cold separation on the first business indicator data according to the timeliness of the first business indicator data to obtain second business indicator data and third business indicator data;
[0012] The business health index of the second business indicator data and the business health index of the third business indicator data are calculated respectively using a preset health evaluation model.
[0013] Furthermore, historical fault information is obtained, and an alarm rule is constructed based on the historical fault information and the business health indicator corresponding to the second business indicator data, including:
[0014] Get historical fault information;
[0015] Searching the second service indicator data for fourth service indicator data corresponding to the historical fault information, and determining fifth service indicator data in the second service indicator data that has no fault except the fourth service indicator data;
[0016] Analyzing, based on the business health indicator of the fourth business indicator data and the business health indicator of the fifth business indicator data, a correlation between the business health indicator of each business indicator data and the occurrence of a fault;
[0017] An alarm threshold for each type of business indicator data is determined according to the association relationship, and an alarm rule is constructed according to the alarm threshold for each type of business indicator data.
[0018] Furthermore, an alarm threshold for each type of business indicator data is determined based on the association relationship, and an alarm rule is constructed based on the alarm threshold for each type of business indicator data, including:
[0019] Determine the business health indicator range when the fault occurs based on the association relationship;
[0020] For each type of business indicator data, the business health indicator range is mapped using a mapping coefficient corresponding to the type of each type of business indicator data to obtain a corresponding alarm threshold;
[0021] Build alarm rules based on the alarm thresholds of each business indicator data.
[0022] Furthermore, generating a topology structure corresponding to the type of the first business indicator data, and configuring corresponding alarm rules for the topology structure, so as to perform operation and maintenance monitoring on the third business indicator data using the topology structure configured with the alarm rules, including:
[0023] determining a data call relationship between different business indicator data according to the type of the first business indicator data;
[0024] Obtaining topological nodes according to each business indicator data, and generating a topological structure using the topological nodes and data call relationships;
[0025] Configuring corresponding alarm rules based on each topological node in the topological structure;
[0026] The topology structure configured with alarm rules is used to perform operation and maintenance monitoring on the third business indicator data.
[0027] Furthermore, the operation and maintenance monitoring of the third business indicator data is performed using the topology structure configured with the alarm rules, further comprising:
[0028] Real-time monitoring of abnormal nodes that trigger alarm rules;
[0029] Determining an upstream link and a downstream link according to a position of the abnormal node in the topology structure;
[0030] Check each node on the downstream link and check the designated adjacent nodes on the upstream link;
[0031] Match the corresponding processing strategy from the preset processing strategy table based on the troubleshooting situation.
[0032] On the other hand, some embodiments of this specification further provide an operation and maintenance monitoring device based on business indicators, the device comprising:
[0033] A receiving module, configured to obtain first business indicator data from multiple nodes and transmit the data to a monitoring node;
[0034] a calculation module, configured to perform hot and cold separation on the first business indicator data to obtain second business indicator data and third business indicator data, and respectively calculate business health indicators corresponding to the second business indicator data and the third business indicator data;
[0035] An acquisition module, configured to acquire historical fault information and construct an alarm rule based on the historical fault information and a business health indicator corresponding to the second business indicator data;
[0036] The operation and maintenance module is used to generate a topology structure corresponding to the type of the first business indicator data and configure corresponding alarm rules for the topology structure, so as to use the topology structure configured with the alarm rules to perform operation and maintenance monitoring on the third business indicator data.
[0037] On the other hand, some embodiments of this specification further provide a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the computer program executes instructions of the above method when executed by the processor.
[0038] On the other hand, some embodiments of this specification further provide a computer storage medium having a computer program stored thereon, wherein the computer program executes the instructions of the above method when executed by a processor of a computer device.
[0039] On the other hand, some embodiments of this specification further provide a computer program product, which includes a computer program. When the computer program is executed by a processor of a computer device, the computer program executes instructions of the above method.
[0040] Some embodiments of this specification provide one or more technical solutions that have at least the following technical effects:
[0041] The embodiments of the present specification automatically obtain first business indicator data from multiple nodes and transmit it to a monitoring node, then perform hot and cold separation on the first business indicator data to obtain second business indicator data and third business indicator data, and respectively calculate the business health indicators corresponding to the second business indicator data and the third business indicator data. Alarm rules are quickly and accurately constructed based on the intrinsic correlation between historical fault information and the business health indicators corresponding to the second business indicator data. On this basis, a topological structure corresponding to the type of the first business indicator data is generated, thereby clarifying the data call relationship between different types of first business indicator data, and configuring corresponding alarm rules for the topological structure, so as to use the topological structure configured with alarm rules to perform operation and maintenance monitoring on the third business indicator data, thereby providing efficient troubleshooting and performance optimization means for complex operation and maintenance environments with multiple nodes that require calculation and processing of a large amount of business indicator data.
[0042] The above description is only an overview of the technical solutions of some embodiments of this specification. In order to more clearly understand the technical means of some embodiments of this specification, they can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of some embodiments of this specification more obvious and easy to understand, the specific implementation methods of some embodiments of this specification are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate some embodiments of this specification or technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments described in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without inventive work. In the drawings:
[0044] Figure 1 A schematic diagram of an implementation system of an operation and maintenance monitoring method based on business indicators in some embodiments of this specification is shown;
[0045] Figure 2 A flowchart of an operation and maintenance monitoring method based on business indicators in some embodiments of this specification is shown;
[0046] Figure 3 This is a schematic diagram of the steps for calculating the business health indicator corresponding to the second business indicator data and the third business indicator data in some embodiments of this specification;
[0047] Figure 4 This is a schematic diagram of the steps of constructing an alarm rule based on historical fault information and a service health indicator corresponding to the second service indicator data in some embodiments of this specification;
[0048] Figure 5 This is a schematic diagram of the steps of determining the alarm threshold of each type of business indicator data according to the association relationship and constructing an alarm rule according to the alarm threshold of each type of business indicator data in some embodiments of this specification;
[0049] Figure 6 A schematic diagram of steps for generating a topology structure corresponding to the type of first business indicator data and configuring corresponding alarm rules for the topology structure in some embodiments of this specification, so as to perform operation and maintenance monitoring on third business indicator data using the topology structure configured with the alarm rules;
[0050] Figure 7 This is a schematic diagram of the steps of performing operation and maintenance monitoring on third business indicator data using a topology structure configured with alarm rules in some embodiments of this specification;
[0051] Figure 8 This is a structural diagram of an operation and maintenance monitoring device based on business indicators in some embodiments of this specification;
[0052] Figure 9 This is a schematic diagram of the computer device structure provided in some embodiments of this specification.
[0053] [Description of Reference Numerals]
[0054] 101, node;
[0055] 102. Monitoring node;
[0056] 801, receiving module;
[0057] 802, calculation module;
[0058] 803. Get module;
[0059] 804, operation and maintenance module;
[0060] 902. Computer equipment;
[0061] 904, processor;
[0062] 906. Memory;
[0063] 908, driving mechanism;
[0064] 910, input / output interface;
[0065] 912. Input devices;
[0066] 914. Output device;
[0067] 916. Presentation equipment;
[0068] 918. Graphical User Interface;
[0069] 920, network interface;
[0070] 922, communication link;
[0071] 924. Communication bus. DETAILED DESCRIPTION
[0072] To help those skilled in the art better understand the technical solutions in this specification, the following will provide a clear and complete description of the technical solutions in the embodiments of this specification, in conjunction with the accompanying drawings of some embodiments of this specification. Obviously, the embodiments described are only some of the embodiments of this specification, not all of them. All other embodiments obtained by those skilled in the art based on some of the embodiments in this specification without creative work should fall within the scope of protection of this specification.
[0073] It should be noted that the terms "first," "second," and the like in the specification and claims herein and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0074] It should be noted that the acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of relevant laws and regulations.
[0075] It should be noted that in the embodiments of this specification, certain software, components, models and other existing solutions in the industry may be mentioned. They should be regarded as exemplary and their purpose is only to illustrate the feasibility of implementing the technical solution of this application, but it does not mean that the applicant has or will necessarily use the solution.
[0076] like Figure 1 The diagram shows a schematic diagram of an implementation system for an operation and maintenance monitoring method based on business indicators according to an embodiment of the present invention. The diagram may include: a node 101 and a monitoring node 102. Nodes 101 and 102 communicate with each other via a network. The network may include a local area network (LAN), a wide area network (WAN), the Internet, or a combination thereof, and may be connected to a website, a user device (e.g., a computing device), and a back-end system. A staff member may send an operation and maintenance monitoring request to the monitoring node 102. After receiving the operation and maintenance monitoring request, the monitoring node 102 retrieves data from the node 101 for operation and maintenance monitoring, and sends the operation and maintenance monitoring results to the monitoring node 101 so that the staff member can process the business according to the operation and maintenance monitoring results.
[0077] In the embodiments of this specification, the monitoring node 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms.
[0078] In an optional embodiment, node 101 may include, but is not limited to, electronic devices such as self-service terminals, desktop computers, tablet computers, laptop computers, and smart wearable devices. Optionally, the operating system running on the electronic device may include, but is not limited to, Android, iOS, Linux, and Windows. Of course, node 101 is not limited to the aforementioned electronic devices having a certain physical form; it may also be software running on the aforementioned electronic devices.
[0079] In addition, it should be noted that Figure 1 What is shown is only one application environment provided by the present disclosure. In actual application, more nodes 101 may be included, and this specification does not limit this.
[0080] Figure 2 It is a flowchart of an operation and maintenance monitoring method based on business indicators provided by an embodiment of the present invention. This specification provides the method operation steps described in the embodiment or flowchart, but it may include more or fewer operation steps based on conventional or non-creative labor. The order of steps listed in the embodiment is only one way of executing the steps among many steps, and does not represent the only execution order. When the actual system or device product is executed, it can be executed in sequence or in parallel according to the method shown in the embodiment or the accompanying drawings. Specifically, Figure 2 As shown, applying the above-mentioned server side, the method may include:
[0081] S201: Acquire first business indicator data from multiple nodes and transmit it to a monitoring node;
[0082] S202: performing hot and cold separation on the first business indicator data to obtain second business indicator data and third business indicator data, and calculating business health indicators corresponding to the second business indicator data and the third business indicator data respectively;
[0083] S203: Acquire historical fault information, and construct an alarm rule based on the historical fault information and the business health indicator corresponding to the second business indicator data;
[0084] S204: Generate a topology structure corresponding to the type of the first business indicator data, and configure corresponding alarm rules for the topology structure, so as to perform operation and maintenance monitoring on the third business indicator data using the topology structure configured with the alarm rules.
[0085] The embodiments of the present specification automatically obtain first business indicator data from multiple nodes and transmit it to a monitoring node, then perform hot and cold separation on the first business indicator data to obtain second business indicator data and third business indicator data, and respectively calculate the business health indicators corresponding to the second business indicator data and the third business indicator data. Alarm rules are quickly and accurately constructed based on the intrinsic correlation between historical fault information and the business health indicators corresponding to the second business indicator data. On this basis, a topological structure corresponding to the type of the first business indicator data is generated, thereby clarifying the data call relationship between different types of first business indicator data, and configuring corresponding alarm rules for the topological structure, so as to use the topological structure configured with alarm rules to perform operation and maintenance monitoring on the third business indicator data, thereby providing efficient troubleshooting and performance optimization means for complex operation and maintenance environments with multiple nodes that require calculation and processing of a large amount of business indicator data.
[0086] It can be understood that in some embodiments, in a distributed system, first business indicator data is stored in multiple different nodes, which can reflect the indicators of the business objectives. For example, if the business objective is the profitability of a certain product, then the first business indicator data may include net profit, number of active users, customer conversion rate, etc. If the business objective is server performance optimization, then the first business indicator data may include ambient temperature, bandwidth, memory usage, etc. Furthermore, in some embodiments, for certain types of first business indicator data, it can be obtained by aggregating multiple other business indicator data. For example, the user satisfaction index can be obtained by aggregating the user evaluation index, the average user access delay, and the average user access success rate at a ratio of 60%, 20%, and 20%, respectively.
[0087] Refer to the attached Figure 3 In some embodiments, performing hot and cold separation on the first business indicator data to obtain second business indicator data and third business indicator data, and calculating business health indicators corresponding to the second business indicator data and the third business indicator data, respectively, may include:
[0088] S301: performing hot and cold separation on the first business indicator data according to the timeliness of the first business indicator data to obtain second business indicator data and third business indicator data;
[0089] S302: Calculate the business health index of the second business indicator data and the business health index of the third business indicator data respectively using a preset health evaluation model.
[0090] It can be understood that in some embodiments, in a distributed system, data is dispersed and stored on multiple data storage nodes for distributed storage to improve data reliability. Furthermore, based on a combination of multiple data acquisition methods, such as interface requests, database queries, text queries, etc., a monitoring agent can be used to collect time series indicator data at the second or minute level. After the collected data is aggregated, calculated, and pre-processed by adding indicator corresponding tags, it is reported to the data server of the monitoring system, i.e., the monitoring node, via the HTTP network protocol. Hot and cold data are separated at the monitoring node for persistent storage, and high-performance storage resources are allocated to hot data with high access frequency, such as SSD solid-state drives, to ensure the rapid response of these data. At the same time, cold data with low access frequency is stored on lower-cost storage media to improve data access performance and increase data capacity.
[0091] In some embodiments, based on the timeliness of the first business indicator data, the historical data (i.e., the second business indicator data) and the real-time data (the third business indicator data) can be partitioned and stored by hot and cold separation, and then the business health indicators of the second business indicator data and the business health indicators of the third business indicator data are calculated respectively using a preset health assessment model, so that subsequent operation and maintenance monitoring work can be carried out based on the business health indicators of the second business indicator data and the business health indicators of the third business indicator data. Among them, the various outputs of the health assessment model can include business interface health check values, business interface response time, business interface response success rate, etc. In some embodiments, the various outputs of the health assessment model can also be weighted and fused to obtain a total health assessment result. Since the business interface health check value, business interface response time, business interface response success rate, etc. are not the focus of the present invention, they will not be described here.
[0092] Refer to the attached Figure 4 In some embodiments, obtaining historical fault information and constructing an alarm rule based on the historical fault information and the business health indicator corresponding to the second business indicator data may include:
[0093] S401: Obtain historical fault information;
[0094] S402: searching the second service indicator data for fourth service indicator data corresponding to the historical fault information, and determining fifth service indicator data in the second service indicator data that has no fault except the fourth service indicator data;
[0095] S403: Analyze the correlation between the business health indicator of each business indicator data and the occurrence of a fault based on the business health indicator of the fourth business indicator data and the business health indicator of the fifth business indicator data;
[0096] S404: Determine an alarm threshold for each type of business indicator data according to the association relationship, and construct an alarm rule according to the alarm threshold for each type of business indicator data.
[0097] It can be understood that in some embodiments, historical fault information can be used to reflect faults that occurred during historical monitoring and the corresponding business indicator data and business health indicators. In the second business indicator data, only a small portion of the business indicator data is related to the occurrence of faults. Therefore, it is necessary to search the second business indicator data for the fourth business indicator data corresponding to the historical fault information. It should be noted that all fault information corresponding to the second business indicator data is not equivalent to the historical fault information. That is, there is an intersection between all fault information corresponding to the second business indicator data and the historical fault information. The business indicator data corresponding to this intersection is the fourth business indicator data. Subsequently, it is necessary to determine the fifth business indicator data in the second business indicator data that has not experienced a fault, excluding the fourth business indicator data, in order to compare and analyze the business health indicators of the fourth business indicator data that has experienced a fault and the business health indicators of the fifth business indicator data that has not experienced a fault, and to obtain the correlation between the business health indicator of each type of business indicator data and the occurrence of a fault. The method for analyzing the correlation can be implemented through nonlinear regression analysis, neural network model, etc., which is not limited herein. The alarm threshold of each type of business indicator data is determined based on the correlation, and the alarm rule is constructed based on the alarm threshold of each type of business indicator data.
[0098] Refer to the attached Figure 5 In some embodiments, determining an alarm threshold for each type of business indicator data based on the association relationship and constructing an alarm rule based on the alarm threshold for each type of business indicator data may include:
[0099] S501: Determine the service health indicator range when the fault occurs based on the association relationship;
[0100] S502: For each type of business indicator data, mapping the business health indicator range using a mapping coefficient corresponding to the type of each type of business indicator data to obtain a corresponding alarm threshold;
[0101] S503: Constructing alarm rules according to the alarm threshold of each business indicator data.
[0102] It can be understood that in some embodiments, the correlation between the business health indicator of each business indicator data and the occurrence of a fault can be obtained through nonlinear regression analysis and a neural network model. According to the correlation, the range of the business health indicator when the fault occurs can be quickly obtained. Afterwards, for each business indicator data, the business health indicator range is mapped using a mapping coefficient corresponding to the type of each business indicator data. The mapping coefficient is actually a gain coefficient, which can expand the range of the business health indicator to a certain extent and obtain the corresponding alarm threshold. In some embodiments, the alarm threshold can also be a numerical threshold or an interval range. This article does not limit this. According to the alarm threshold of each business indicator data, an alarm rule can be quickly constructed. The alarm rule is used to quickly and accurately determine whether each business indicator data exceeds the safety range specified by its corresponding alarm threshold.
[0103] Refer to the attached Figure 6 In some embodiments, generating a topology structure corresponding to the type of the first business indicator data and configuring corresponding alarm rules for the topology structure, so as to perform operation and maintenance monitoring on the third business indicator data using the topology structure configured with the alarm rules, may include:
[0104] S601: Determine a data call relationship between different business indicator data according to the type of the first business indicator data;
[0105] S602: Obtaining topological nodes according to each business indicator data, and generating a topological structure using the topological nodes and data call relationships;
[0106] S603: configuring corresponding alarm rules based on each topological node in the topological structure;
[0107] S604: Perform operation and maintenance monitoring on the third business indicator data using the topology structure configured with alarm rules.
[0108] It can be understood that, in some embodiments, since there is a data call relationship between different types of first business indicator data, there is a corresponding topological structure between different types of first business indicator data. Specifically, a topological node can be obtained according to each type of business indicator data, and the type of the first business indicator data is regarded as a node. The node has the first business indicator data of the corresponding type. The connection path between the topological nodes is obtained according to the data call relationship, so as to quickly generate a topological structure corresponding to the type of the first business indicator data. Then, the topological structure is configured with corresponding alarm rules based on each topological node, including configuring corresponding alarm rules at each topological node to evaluate the business health indicator status of each topological node in real time, and configuring corresponding alarm rules at the connection path between each topological node and the topological node to monitor in real time whether the communication path corresponding to each connection path can communicate normally, thereby utilizing the topological structure configured with alarm rules to perform efficient and accurate operation and maintenance monitoring of real-time third business indicator data.
[0109] Refer to the attached Figure 7 In some embodiments, performing operation and maintenance monitoring on the third business indicator data using a topology structure configured with alarm rules may further include:
[0110] S701: Real-time monitoring of abnormal nodes that trigger alarm rules;
[0111] S702: Determine an upstream link and a downstream link according to the position of the abnormal node in the topology structure;
[0112] S703: Check the downstream link node by node and check the designated adjacent nodes of the upstream link;
[0113] S704: Match a corresponding processing strategy from a preset processing strategy table according to the troubleshooting situation.
[0114] It can be understood that in some embodiments, the abnormal node that triggers the alarm rule is monitored in real time. The abnormal node can be a single node, a pair of nodes (for example, a pair of nodes corresponding to a communication path with an abnormality), or multiple nodes, for example, multiple abnormal nodes with a business association relationship (usually clustered in the topology structure). Then, the upstream link and downstream link are determined based on the position of the abnormal node in the topology structure. For the downstream link of the abnormal node, since the topological nodes in the downstream link may need to use the data of the abnormal node, it is often necessary to conduct a node-by-node investigation. For the upstream link of the abnormal node, the designated neighboring nodes of the upstream link can be investigated. That is, the abnormality of the abnormal node does not necessarily mean that the upstream link has an abnormality. When investigating, it is possible to consider only investigating the designated neighboring nodes of the upstream link to avoid excessive analysis of the upstream link, which may cause the real fault event to be drowned out by the noise caused by the non-faulty topological nodes. Common investigation methods include basic inspections (for example, node reachability tests, system resource status checks, etc.), log analysis, and in-depth network investigations (for example, connection tracking, packet capture and analysis, etc.), so as to quickly and accurately match the corresponding processing strategy from the preset processing strategy table based on the investigation results.
[0115] It should be noted that although the operations of the method of the present invention are described in a specific order in the above embodiments and drawings, this does not require or imply that these operations must be performed in this specific order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0116] Corresponding to the above-mentioned operation and maintenance monitoring method based on business indicators, some embodiments of this specification also provide an operation and maintenance monitoring device based on business indicators, referring to Figure 8 As shown, in some embodiments, the apparatus may include:
[0117] The receiving module 801 is configured to obtain first business indicator data from multiple nodes and transmit the data to the monitoring node;
[0118] A calculation module 802 is configured to perform hot and cold separation on the first business indicator data to obtain second business indicator data and third business indicator data, and calculate business health indicators corresponding to the second business indicator data and the third business indicator data, respectively;
[0119] An acquisition module 803 is configured to acquire historical fault information and construct an alarm rule based on the historical fault information and a business health indicator corresponding to the second business indicator data;
[0120] The operation and maintenance module 804 is configured to generate a topology structure corresponding to the type of the first business indicator data and configure corresponding alarm rules for the topology structure, so as to perform operation and maintenance monitoring on the third business indicator data using the topology structure configured with the alarm rules.
[0121] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0122] It should be noted that in the embodiments of this specification, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved are all information and data authorized by the user and fully authorized by all parties.
[0123] It should be noted that the computer program product described in this specification is a software product that mainly implements the method described in this specification through a computer program.
[0124] The embodiment of this specification also provides a computer device. Figure 9 As shown, in some embodiments of this specification, the computer device 902 may include one or more processors 904, such as one or more central processing units (CPUs) or graphics processing units (GPUs), each of which may implement one or more hardware threads. The computer device 902 may also include any memory 906 for storing any type of information, such as code, settings, data, etc. In a specific embodiment, the memory 906 may contain a computer program that can be executed on the processor 904. When executed by the processor 904, the computer program can execute instructions of the method described in any of the above embodiments. For example, without limitation, the memory 906 may include any one or more combinations of the following: any type of RAM, any type of ROM, a flash memory device, a hard disk, an optical disk, etc. More generally, any memory may use any technology to store information. Furthermore, any memory may provide volatile or non-volatile retention of information. Furthermore, any memory may represent a fixed or removable component of the computer device 902. In one embodiment, when the processor 904 executes the associated instructions stored in any memory or combination of memories, the computer device 902 may perform any operation of the associated instructions. The computer device 902 also includes one or more drive mechanisms 908 for interacting with any storage, such as a hard disk drive mechanism, an optical disk drive mechanism, and the like.
[0125] The computer device 902 may also include an input / output interface 910 (I / O) for receiving various inputs (via input devices 912) and for providing various outputs (via output devices 914). A specific output mechanism may include a presentation device 916 and an associated graphical user interface 918 (GUI). In other embodiments, the input / output interface 910 (I / O), input devices 912, and output devices 914 may not be included, and the computer device 902 may simply be a computer device in a network. The computer device 902 may also include one or more network interfaces 920 for exchanging data with other devices via one or more communication links 922. One or more communication buses 924 couple the components described above together.
[0126] The communication link 922 may be implemented in any manner, for example, via a local area network, a wide area network (e.g., the Internet), a point-to-point connection, etc., or any combination thereof. The communication link 922 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.
[0127] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), computer-readable storage media, and computer program products of some embodiments of the present specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processor to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processor generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0128] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processor to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, the instruction device being implemented in the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0129] These computer program instructions can also be loaded onto a computer or other programmable data processor so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0130] In a typical configuration, a computer device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0131] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0132] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computer device. As defined in this specification, computer-readable media does not include temporary computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0133] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the embodiments of this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0134] Embodiments of this specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. Embodiments of this specification may also be practiced in distributed computing environments where tasks are performed by remote processors connected via a communications network. In distributed computing environments, program modules may be located in local and remote computer storage media, including storage devices.
[0135] It should also be understood that in the embodiments of this specification, the term "and / or" is merely a description of the association relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.
[0136] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.
[0137] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the embodiments of this specification. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0138] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. An operation and maintenance monitoring method based on business indicators, characterized in that: The method comprises: Acquire first business indicator data from multiple nodes and transmit it to a monitoring node; Performing hot and cold separation on the first business indicator data to obtain second business indicator data and third business indicator data, and respectively calculating business health indicators corresponding to the second business indicator data and the third business indicator data; Acquire historical fault information, and construct an alarm rule based on the historical fault information and a business health indicator corresponding to the second business indicator data; A topology structure corresponding to the type of the first business indicator data is generated, and corresponding alarm rules are configured for the topology structure, so as to perform operation and maintenance monitoring on the third business indicator data using the topology structure configured with the alarm rules.
2. The method according to claim 1, characterized in that Performing hot and cold separation on the first business indicator data to obtain second business indicator data and third business indicator data, and calculating business health indicators corresponding to the second business indicator data and the third business indicator data, respectively, including: performing hot and cold separation on the first business indicator data according to the timeliness of the first business indicator data to obtain second business indicator data and third business indicator data; The business health index of the second business indicator data and the business health index of the third business indicator data are calculated respectively using a preset health evaluation model.
3. The method according to claim 2, characterized in that Obtain historical fault information, and construct an alarm rule based on the historical fault information and the business health indicator corresponding to the second business indicator data, including: Get historical fault information; Searching the second service indicator data for fourth service indicator data corresponding to the historical fault information, and determining fifth service indicator data in the second service indicator data that has no fault except the fourth service indicator data; Analyzing, based on the business health indicator of the fourth business indicator data and the business health indicator of the fifth business indicator data, a correlation between the business health indicator of each business indicator data and the occurrence of a fault; An alarm threshold for each type of business indicator data is determined according to the association relationship, and an alarm rule is constructed according to the alarm threshold for each type of business indicator data.
4. The method according to claim 3, characterized in that Determine the alarm threshold of each business indicator data according to the association relationship, and construct an alarm rule according to the alarm threshold of each business indicator data, including: Determine the business health indicator range when the fault occurs based on the association relationship; For each type of business indicator data, the business health indicator range is mapped using a mapping coefficient corresponding to the type of each type of business indicator data to obtain a corresponding alarm threshold; Build alarm rules based on the alarm thresholds of each business indicator data.
5. The method according to claim 1, wherein Generating a topology structure corresponding to the type of the first business indicator data, and configuring corresponding alarm rules for the topology structure, so as to perform operation and maintenance monitoring on the third business indicator data using the topology structure configured with the alarm rules, including: determining a data call relationship between different business indicator data according to the type of the first business indicator data; Obtaining topological nodes according to each business indicator data, and generating a topological structure using the topological nodes and data call relationships; Configuring corresponding alarm rules based on each topological node in the topological structure; The topology structure configured with alarm rules is used to perform operation and maintenance monitoring on the third business indicator data.
6. The method according to claim 1, characterized in that The operation and maintenance monitoring of the third business indicator data is performed using a topology structure configured with alarm rules, further comprising: Real-time monitoring of abnormal nodes that trigger alarm rules; Determining an upstream link and a downstream link according to a position of the abnormal node in the topology structure; Check each node on the downstream link and check the designated adjacent nodes on the upstream link; Match the corresponding processing strategy from the preset processing strategy table based on the troubleshooting situation.
7. An operation and maintenance monitoring device based on business indicators, characterized in that: The device comprises: A receiving module, configured to obtain first business indicator data from multiple nodes and transmit the data to a monitoring node; a calculation module, configured to perform hot and cold separation on the first business indicator data to obtain second business indicator data and third business indicator data, and respectively calculate business health indicators corresponding to the second business indicator data and the third business indicator data; An acquisition module, configured to acquire historical fault information and construct an alarm rule based on the historical fault information and a business health indicator corresponding to the second business indicator data; The operation and maintenance module is used to generate a topology structure corresponding to the type of the first business indicator data and configure corresponding alarm rules for the topology structure, so as to use the topology structure configured with the alarm rules to perform operation and maintenance monitoring on the third business indicator data.
8. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that: When the computer program is executed by the processor, the computer program executes the instructions of the method according to any one of claims 1 to 6.
9. A computer storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor of a computer device, the computer program executes the instructions of the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the computer program executes instructions of the method according to any one of claims 1 to 6.