Fault node detection and alarm method, device and equipment of business system and medium

By generating a virtual node call relationship graph and detecting parameter values ​​in real time, the problem of low efficiency in manual detection in existing technologies is solved, and efficient fault node detection and alarm are achieved.

CN116800635BActive Publication Date: 2026-05-19CHINA UNITED NETWORK COMM GRP CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA UNITED NETWORK COMM GRP CO LTD
Filing Date
2023-06-20
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing methods for detecting fault nodes in business systems rely on manual inspection, resulting in low detection efficiency.

Method used

By obtaining the function type of each virtual node in the business system, a call relationship diagram of the target virtual node is generated. The parameter values ​​of the parameters to be detected are obtained at preset intervals. Fault detection is performed according to preset detection rules, alarm information is generated and marked in the call relationship diagram, and the call relationship diagram of the fault node is output.

Benefits of technology

It improves the efficiency of fault node detection, allowing users to intuitively view fault nodes and enhancing the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116800635B_ABST
    Figure CN116800635B_ABST
Patent Text Reader

Abstract

The application provides a fault node detection and alarm method, device and equipment of a business system and a medium, which can be used in the field of computers. In the method, after obtaining the function type corresponding to each virtual node in the business system, a first target virtual node call relationship diagram of the business system is generated. Then, the preset detection rules and parameter values corresponding to the detection parameters of each virtual node are obtained, and then the virtual node is detected for failure every interval of a preset second detection time; if the virtual node fails, the generated alarm information and the virtual node are marked in the first target virtual node call relationship diagram to obtain and output a second target virtual node call relationship diagram. According to the parameter values, the present scheme determines whether the virtual node in the business system is faulty, and outputs the faulty virtual node and the alarm information in the virtual node call relationship diagram, so that the user can intuitively view the fault node, and the detection efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and in particular to a method, apparatus, device and medium for fault node detection and alarm in a business system. Background Technology

[0002] With the development of technology, cloud computing has played an important role in the digital transformation of enterprises. More and more enterprises are migrating their business systems to the cloud, thereby achieving advantages such as resource sharing, elastic scaling, and rapid deployment.

[0003] In existing technologies, users can migrate their business systems to the cloud, which involves building multiple virtual nodes. These virtual nodes constitute the business system, enabling the processing of business data. These virtual nodes can be virtual servers, virtual hard drives, virtual gateways, etc. During the operation of the business system, failures are inevitable. Typically, troubleshooting faulty nodes involves staff reviewing historical data to identify them.

[0004] In summary, the existing methods for detecting fault nodes in business systems rely on manual inspection, resulting in low detection efficiency. Summary of the Invention

[0005] This application provides a method, apparatus, device, and medium for fault node detection and alarm in a business system, which addresses the problem that existing fault node detection methods in business systems rely on manual detection, resulting in low detection efficiency.

[0006] Firstly, this application provides a method for fault node detection and alarm in a business system, including:

[0007] Obtain the function type corresponding to each virtual node in the business system;

[0008] Based on the functional type of each virtual node, generate the first target virtual node call relationship diagram of the business system;

[0009] For each parameter to be detected corresponding to each virtual node, obtain the preset detection rule corresponding to the parameter to be detected, and obtain the parameter value corresponding to the parameter to be detected once every preset first detection time interval;

[0010] For each interval of the parameter to be detected, a preset second detection duration is defined. Based on the preset detection rules and the parameter values ​​corresponding to the parameter to be detected obtained within the preset second detection duration, the virtual node is subjected to fault detection to determine whether the virtual node is faulty. The preset second detection duration is an integer multiple of the preset first detection duration.

[0011] If the virtual node fails, an alarm message is generated, and the alarm message and the virtual node are marked in the first target virtual node call relationship diagram to obtain the second target virtual node call relationship diagram;

[0012] Output the call relationship diagram of the second target virtual node.

[0013] In one specific implementation, generating the first target virtual node call relationship diagram of the business system based on the functional type of each virtual node includes:

[0014] Based on the functional type of each virtual node, the initial call relationship between each virtual node is determined from the pre-generated virtual node baseline call relationship graph;

[0015] Generate an initial virtual node call relationship graph based on the initial call relationships between each virtual node;

[0016] In response to the user's operation to adjust the relationship graph, a relationship graph for the first target virtual node is generated.

[0017] In one specific implementation, if the preset detection rule is a preset first detection rule, then the preset second detection duration is equal to the preset first detection duration. The step of performing fault detection on the virtual node based on the preset detection rule and the parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration, to determine whether the virtual node is faulty, includes:

[0018] Determine whether the parameter to be detected is marked with an alarm persistence status identifier;

[0019] If the parameter to be detected is not marked with the alarm continuity status identifier, then it is determined whether the parameter value is greater than the preset first alarm parameter threshold.

[0020] If the parameter value is greater than the preset first alarm parameter threshold, then update the cumulative number of consecutive anomalies;

[0021] Determine whether the updated cumulative number of consecutive anomalies is equal to the preset first alarm cumulative threshold;

[0022] If the updated cumulative number of consecutive anomalies is equal to the preset first alarm cumulative threshold, then the virtual node is determined to be faulty, and the alarm persistence status identifier is marked on the parameter to be detected.

[0023] If the updated cumulative number of consecutive anomalies is less than the preset first alarm cumulative threshold, then the virtual node is determined to be normal.

[0024] If the parameter value is less than or equal to the preset first alarm parameter threshold, then the virtual node is determined to be normal, and the cumulative number of consecutive anomalies is updated.

[0025] If the parameter to be detected marks the alarm persistence status identifier, then determine whether the parameter value is less than or equal to the preset first alarm parameter threshold.

[0026] If the parameter value is less than or equal to the preset first alarm parameter threshold, then update the cumulative number of consecutive normal alarms;

[0027] Determine whether the updated cumulative number of consecutive normal occurrences is equal to the preset first warning elimination threshold;

[0028] If the updated cumulative number of consecutive normal occurrences is less than the preset first warning elimination threshold, then the virtual node is determined to be faulty.

[0029] If the updated cumulative number of consecutive normal occurrences is equal to the preset first warning elimination threshold, then the virtual node is determined to be normal, and the alarm persistence status identifier marked on the parameter to be detected is removed.

[0030] If the parameter value is greater than the preset first alarm parameter threshold, then the virtual node is determined to be faulty, and the cumulative number of consecutive normal operations is updated.

[0031] In one specific implementation, if the preset detection rule is a preset second detection rule, then the number of parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration is the target detection quantity. The step of performing fault detection on the virtual node based on the preset detection rule and the parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration to determine whether the virtual node is faulty includes:

[0032] Determine whether the parameter to be detected is marked with an alarm persistence status identifier;

[0033] If the parameter to be detected is not marked with the alarm continuity status identifier, then determine whether the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the target detection quantity is greater than the preset third alarm parameter threshold.

[0034] If the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is greater than the preset third alarm parameter threshold, then the cumulative number of consecutive anomalies is updated.

[0035] Determine whether the updated cumulative number of consecutive anomalies is equal to the preset second alarm cumulative threshold;

[0036] If the updated cumulative number of consecutive anomalies is equal to the preset second alarm cumulative threshold, then the virtual node is determined to be faulty, and the alarm persistence status identifier is marked on the parameter to be detected.

[0037] If the updated cumulative number of consecutive anomalies is less than the preset second alarm cumulative threshold, then the virtual node is determined to be normal.

[0038] If the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is less than or equal to the preset third alarm parameter threshold, then the virtual node is determined to be normal, and the cumulative number of consecutive anomalies is updated.

[0039] If the parameter to be detected is not marked with the alarm continuity status identifier, then it is determined whether the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the target detection quantity is less than the preset fourth alarm parameter threshold, and the preset fourth alarm parameter threshold is less than the preset third alarm parameter threshold.

[0040] If the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is less than the preset fourth alarm parameter threshold, then the cumulative number of consecutive normal occurrences is updated.

[0041] Determine whether the updated consecutive normal cumulative count is equal to the preset second warning elimination threshold;

[0042] If the updated cumulative number of consecutive normal occurrences is less than the preset second warning elimination threshold, then the virtual node is determined to be faulty.

[0043] If the updated cumulative number of consecutive normal occurrences is equal to the preset second warning elimination threshold, then the virtual node is determined to be normal, and the alarm persistence status identifier marked on the parameter to be detected is removed.

[0044] If the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is greater than or equal to the preset fourth alarm parameter threshold, then the virtual node is determined to be faulty, and the cumulative number of consecutive normal operations is updated.

[0045] In one specific implementation, if the preset detection rule is a preset third detection rule, then the preset second detection duration is equal to the preset first detection duration. The step of performing fault detection on the virtual node based on the preset detection rule and the parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration, to determine whether the virtual node is faulty, includes:

[0046] Get the current month and the current date within the current month;

[0047] Based on the current month and the current date within the current month, as well as the correspondence between the month / date and the alarm threshold, a preset fifth alarm threshold corresponding to the current month and the current date within the current month is determined;

[0048] If the parameter value is greater than the preset fifth alarm threshold, then the virtual node is determined to be faulty;

[0049] If the parameter value is less than or equal to the preset fifth alarm threshold, then the virtual node is determined to be normal.

[0050] In one specific implementation, if the preset detection rule is a preset fourth detection rule, then the preset second detection duration is equal to the preset first detection duration. The step of performing fault detection on the virtual node based on the preset detection rule and the parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration, to determine whether the virtual node is faulty, includes:

[0051] Get the current month and the current date within the current month;

[0052] The current month and the current date under the current month are input into the preset alarm threshold prediction model to determine the preset sixth alarm threshold; the preset alarm threshold prediction model is a pre-trained calculation model for determining the alarm threshold based on the month and date;

[0053] If the parameter value is greater than the preset sixth alarm threshold, then the virtual node is determined to be faulty;

[0054] If the parameter value is less than or equal to the preset sixth alarm threshold, then the virtual node is determined to be normal.

[0055] In one specific embodiment, the method further includes:

[0056] For each faulty virtual node, the affected virtual node corresponding to each faulty virtual node is determined by the second target virtual node call relationship graph, and / or the preset affected node prediction model, and / or the correspondence between virtual nodes and affected nodes; the preset affected node prediction model is a pre-trained computational model used to determine the affected node based on the virtual node.

[0057] The affected virtual nodes are marked in the second target virtual node call relationship graph to obtain the third target virtual node call relationship graph;

[0058] Output the call relationship diagram of the third target virtual node.

[0059] In one specific implementation, after generating the first target virtual node call relationship graph of the business system, the method further includes:

[0060] Output the call relationship graph of the first target virtual node.

[0061] In one specific implementation, before obtaining the function type corresponding to each virtual node in the business system, the method further includes:

[0062] Retrieve all baseline function types from the virtual node function type library;

[0063] Retrieve all baseline call relationships from the call relationship database, as well as the baseline function type corresponding to each baseline call relationship. All types of baseline call relationships include one or more of the following: containment relationship, access relationship, access relationship, binding relationship, and mount relationship.

[0064] The virtual node baseline call relationship diagram is generated based on all baseline function types, all baseline call relationships, and the baseline function type corresponding to each baseline call relationship.

[0065] Secondly, this application provides a fault node detection and alarm device for a business system, comprising:

[0066] The acquisition module is used to acquire the function type corresponding to each virtual node in the business system.

[0067] The generation module is used to generate the first target virtual node call relationship diagram of the business system based on the functional type of each virtual node;

[0068] The acquisition module is also used to acquire the preset detection rule corresponding to each parameter to be detected for each virtual node, and acquire the parameter value corresponding to the parameter to be detected once every preset first detection time interval;

[0069] The processing module is used to perform fault detection on the virtual node at intervals corresponding to a preset second detection duration for the parameter to be detected, based on the preset detection rules and the parameter values ​​corresponding to the parameter to be detected obtained within the preset second detection duration, to determine whether the virtual node is faulty; the preset second detection duration is an integer multiple of the preset first detection duration;

[0070] The generation module is further configured to generate alarm information if the virtual node fails, and mark the alarm information and the virtual node in the first target virtual node call relationship diagram to obtain the second target virtual node call relationship diagram;

[0071] The output module is used to output the call relationship diagram of the second target virtual node.

[0072] Thirdly, this application provides an electronic device, comprising:

[0073] Processor, memory, communication interface;

[0074] The memory is used to store the executable instructions of the processor;

[0075] The processor is configured to execute the fault node detection and alarm method of the business system according to any one of the first aspects by executing the executable instructions.

[0076] Fourthly, this application provides a readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the fault node detection and alarm method of the business system described in any one of the first aspects.

[0077] The method, apparatus, equipment, and medium for fault node detection and alarm in a business system provided in this application obtain the function type corresponding to each virtual node in the business system, and then generate a first target virtual node call relationship diagram based on the function type of each virtual node. Next, for each parameter to be detected corresponding to each virtual node, a corresponding preset detection rule is obtained, and the parameter value corresponding to the parameter to be detected is obtained once at a preset first detection interval. Then, at a preset second detection interval corresponding to the parameter to be detected, the virtual node is fault-detected based on the preset detection rule and the parameter value obtained within the preset second detection interval to determine whether the virtual node is faulty. If the virtual node is faulty, alarm information is generated, and the alarm information and the virtual node are marked in the first target virtual node call relationship diagram to obtain a second target virtual node call relationship diagram; finally, the second target virtual node call relationship diagram is output. This solution determines whether a virtual node in the business system is faulty based on the parameter value of each virtual node, and outputs the faulty virtual node and alarm information after marking them in the virtual node call relationship diagram, allowing users to intuitively view faulty nodes and effectively improving detection efficiency. Attached Figure Description

[0078] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0079] Figure 1a A flowchart illustrating an embodiment of the fault node detection and alarm method for the business system provided in this application;

[0080] Figure 1b The virtual node baseline call relationship diagram provided in this application;

[0081] Figure 1c This application provides an initial virtual node call relationship diagram;

[0082] Figure 1d This application provides a first target virtual node call relationship diagram;

[0083] Figure 1e This is the second target virtual node call relationship diagram provided in this application;

[0084] Figure 2a A flowchart illustrating Embodiment 2 of the fault node detection and alarm method for the business system provided in this application;

[0085] Figure 2b A schematic diagram illustrating the application of the preset first detection rule provided in this application;

[0086] Figure 3a A flowchart illustrating Embodiment 3 of the fault node detection and alarm method for the business system provided in this application;

[0087] Figure 3b A schematic diagram illustrating the application of the preset second detection rule provided in this application;

[0088] Figure 4 A flowchart illustrating Embodiment 4 of the fault node detection and alarm method for the business system provided in this application;

[0089] Figure 5 A flowchart illustrating Embodiment 5 of the fault node detection and alarm method for the business system provided in this application;

[0090] Figure 6a A flowchart illustrating Embodiment Six of the fault node detection and alarm method for the business system provided in this application;

[0091] Figure 6b The third target virtual node call relationship diagram provided in this application;

[0092] Figure 7 A flowchart illustrating Embodiment Seven of the Fault Node Detection and Alarm Method for the Business System Provided in this Application;

[0093] Figure 8 A schematic diagram of the structure of an embodiment of the fault node detection and alarm device for the business system provided in this application;

[0094] Figure 9 This is a schematic diagram of the structure of an electronic device provided in this application. Detailed Implementation

[0095] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments made by those skilled in the art under the guidance of these embodiments are within the scope of protection of this application.

[0096] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0097] With the development of technology, cloud computing has played an important role in the digital transformation of enterprises. More and more enterprises are migrating their business systems to the cloud, thereby achieving advantages such as resource sharing, elastic scaling, and rapid deployment.

[0098] In existing technologies, users can migrate their business systems to the cloud, which involves building multiple virtual nodes. These virtual nodes constitute the business system, enabling the processing of business data. These virtual nodes can be virtual servers, virtual hard disks, virtual gateways, etc. However, failures are inevitable during the operation of the business system. Typically, troubleshooting these failures involves staff reviewing historical data to identify the faulty nodes, leading to low detection efficiency.

[0099] To address the problems existing in the prior art, the inventors, during their research on fault node detection and alarm methods for business systems, discovered that to improve detection efficiency, the business system can be monitored in real time, and each virtual node can be periodically checked for faults. When a virtual node fault is detected, an alarm is triggered. First, a first target virtual node call relationship diagram of the business system is generated. Then, for each parameter to be detected corresponding to each virtual node, a preset detection rule corresponding to the parameter is obtained, and the parameter value corresponding to the parameter to be detected is obtained once at a preset first detection interval. At a preset second detection interval corresponding to the parameter to be detected, based on the preset detection rule and the parameter value corresponding to the parameter to be detected obtained within the preset second detection interval, fault detection is performed on the virtual node. If a virtual node is faulty, an alarm message is generated, and the alarm message and the virtual node are marked in the first target virtual node call relationship diagram to obtain a second target virtual node call relationship diagram; the second target virtual node call relationship diagram is then output. This effectively improves detection efficiency and makes the alarm effect more intuitive. Based on the above inventive concept, the fault node detection and alarm scheme for business systems in this application was designed.

[0100] The execution subject of the fault node detection and alarm method of the business system in this application can be a server, or a computer, terminal equipment, or other devices. This application does not limit it. The following description uses a server as an example.

[0101] The following provides an example illustrating the application scenarios of the fault node detection and alarm method for the business system provided in this application.

[0102] For example, in this application scenario, the server can monitor and alert on virtual nodes in the business system. The server first obtains the function type corresponding to each virtual node in the business system; then, based on the function type of each virtual node, it generates the first target virtual node call relationship diagram of the business system.

[0103] Then, for each parameter to be detected corresponding to each virtual node, the preset detection rule corresponding to the parameter to be detected is obtained, and the parameter value corresponding to the parameter to be detected is obtained once every preset first detection time interval.

[0104] Each interval corresponds to a preset second detection duration for the parameter to be detected. Based on preset detection rules and the parameter values ​​obtained within the preset second detection duration, fault detection is performed on the virtual node to determine whether it is faulty. If the virtual node is faulty, an alarm message is generated, and the alarm message and the virtual node are marked in the first target virtual node call relationship diagram to obtain the second target virtual node call relationship diagram. The second target virtual node call relationship diagram is then output. In other words, the second target virtual node call relationship diagram is sent to the user's terminal device or large-screen display device, so the user can intuitively see the faulty node.

[0105] It should be noted that the above scenario is only an example of an application scenario provided by the embodiments of this application. The embodiments of this application do not limit the actual form of the various devices included in the scenario, nor do they limit the interaction method between devices. In the specific application of the solution, it can be set according to actual needs.

[0106] The technical solution of this application will now be described in detail through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0107] Figure 1a This is a flowchart illustrating an embodiment of the fault node detection and alarm method for a business system provided in this application. In this embodiment, the server generates a first target virtual node call relationship diagram for the business system, then obtains preset detection rules and parameter values ​​for the parameters to be detected. Fault detection is performed on the virtual node at intervals corresponding to a preset second detection duration for each parameter to be detected. When a virtual node fails, the generated alarm information and the output of the faulty node after being marked in the first target virtual node call relationship diagram are explained. The method in this embodiment can be implemented through software, hardware, or a combination of both. Figure 1a As shown, the fault node detection and alarm method of this business system specifically includes the following steps:

[0108] S101: Obtain the function type corresponding to each virtual node in the business system.

[0109] In this step, in order to enable users to intuitively see which virtual nodes in the business system are faulty, it is necessary to generate the first target virtual node call relationship diagram of the business system. This requires the server to first obtain the function type corresponding to each virtual node in the business system.

[0110] It should be noted that the function type can be obtained through SNMP, Agent, JDBC, HTTP, etc. The function type can be a server function type, gateway function type, hard disk function type, database function type, etc., and the corresponding virtual nodes are virtual servers, virtual network administrators, virtual hard disks, virtual databases, etc. This application does not limit the method of obtaining the function type, the function type, or the virtual node; these can be determined according to the actual situation.

[0111] S102: Generate the first target virtual node call relationship diagram of the business system based on the function type of each virtual node.

[0112] In this step, after the server obtains the function type corresponding to each virtual node in the business system, it can generate the first target virtual node call relationship diagram of the business system based on the function type of each virtual node.

[0113] Specifically, based on the functional type of each virtual node, the initial call relationships between each virtual node are determined from a pre-generated virtual node baseline call relationship graph. Since there are corresponding initial call relationships between each virtual node in the virtual node baseline call relationship graph, the initial call relationships between each virtual node in the business system can be determined.

[0114] It should be noted that the type of initial call relationship can be an inclusion relationship, access relationship, binding relationship, mounting relationship, etc. This application embodiment does not limit the type of initial call relationship, and it can be set according to the actual situation.

[0115] For example, Figure 1b The virtual node baseline call relationship diagram provided in this application is as follows: Figure 1b As shown in the virtual node baseline call relationship diagram, virtual hard disk nodes are mounted to virtual server nodes and virtual bare metal server nodes. Virtual private network (VPC) nodes are contained to virtual server nodes, virtual database nodes, and virtual bare metal server nodes. Virtual database nodes are deployed to virtual server nodes. Virtual elastic public IP (PEP) nodes are bound to virtual gateway nodes. Virtual PEP nodes are accessed by virtual server nodes and virtual bare metal server nodes. Virtual server load balancer nodes and virtual gateway nodes are bound to VPC nodes. VPC nodes are connected to virtual cloud high-speed nodes.

[0116] Then, based on the initial call relationship between each virtual node, an initial virtual node call relationship graph is generated.

[0117] For example, Figure 1c The initial virtual node call relationship diagram provided for this application is as follows: Figure 1cAs shown, the business system has 16 virtual nodes. Virtual Elastic Public IP node A has access relationships with virtual server nodes D and E. Virtual Private Network node G has containment relationships with virtual server nodes D, E, M, and F. Virtual server load balancer node H is bound to virtual private network node G. Virtual Elastic Public IP node B is bound to virtual gateway node I. Virtual gateway node I is bound to virtual private network node J. Virtual private network node J has containment relationships with virtual server nodes K and L. Virtual Elastic Public IP node B has access relationships with virtual server node N. Virtual private network node O has containment relationships with virtual server node N. Virtual private network nodes G, J, and O have access relationships with virtual cloud high-speed node P.

[0118] In response to the user's adjustment of the virtual node call graph, a first target virtual node call graph is generated. Since the call relationships in the initial virtual node call graph may be inaccurate, the user needs to adjust them again. The initial virtual node call graph is sent to the user's terminal device, and the user adjusts the initial virtual node call graph, sending the adjusted initial virtual node call graph to the server. The server then receives the adjusted initial virtual node call graph, thus generating the first target virtual node call graph in response to the user's adjustment operation. In other words, the adjusted initial virtual node call graph is used as the first target virtual node call graph.

[0119] For example, in Figure 1c On this basis, Figure 1d The first target virtual node call relationship diagram provided in this application is as follows: Figure 1d As shown, the user discovered that virtual private network node J and virtual server node M have an inclusion relationship, but it is not virtual private network node G and virtual server node M that have an inclusion relationship. After adjustment, the following can be obtained: Figure 1d .like Figure 1d As shown, the forwarding strategy of virtual server load balancer node H handles the traffic load balancing between virtual server nodes D and E, connecting access requests to the relatively idle virtual server node E, and then accessing the virtual database node F through virtual server node E. Virtual gateway node I enables public network bandwidth sharing within virtual private network node J, allowing access to virtual server node L within virtual private network node J. Virtual server node N is accessed through virtual elastic public IP node C.

[0120] S103: For each parameter to be detected corresponding to each virtual node, obtain the preset detection rule corresponding to the parameter to be detected, and obtain the parameter value corresponding to the parameter to be detected once every preset first detection time interval.

[0121] In this step, after the server obtains the call relationship graph of the first target virtual node, in order to perform fault detection on the virtual node, it is necessary to obtain the preset detection rule corresponding to each parameter to be detected for each virtual node, and obtain the parameter value corresponding to the parameter to be detected once every preset first detection time interval.

[0122] Since the server has already set up a correspondence between the parameter to be detected and the preset detection rules for each parameter to be detected, the preset detection rules corresponding to each parameter to be detected can be determined.

[0123] It should be noted that the number of parameters to be detected for each virtual node can be one or more. These parameters can include peak bandwidth, real-time inbound bandwidth, real-time outbound bandwidth, number of real-time inbound data packets, number of real-time outbound data packets, Central Processing Unit (CPU) utilization, memory utilization, disk utilization, memory usage, storage resource utilization, storage resource usage, read / write operations per second, read / write bandwidth, read / write latency, network inflow / outflow rate, network inflow / outflow data packets, packet loss rate, number of lost packets, and number of Transmission Control Protocol (TCP) connections. This application does not limit the number or types of parameters to be detected for each virtual node; these can be determined based on actual conditions.

[0124] It should be noted that the preset first detection duration can be 0.01 seconds, 0.1 seconds, 1 second, etc. This application embodiment does not limit the preset first detection duration, and it can be set according to the actual situation.

[0125] It should be noted that parameter values ​​can be obtained through application programming interfaces (APIs), agents, JDBC, HTTP, etc. This application does not limit the method of obtaining parameter values; it can be determined according to the actual situation.

[0126] It should be noted that after the server generates the first target virtual node call relationship graph, it can also output the first target virtual node call relationship graph for users to view, allowing them to intuitively see the call relationships between each virtual node. Furthermore, the parameter values ​​of each virtual node can be plotted on the first target virtual node call relationship graph, thus outputting a first target virtual node call relationship graph with labeled parameter values, allowing users to intuitively view the operational status of each virtual node.

[0127] S104: For each interval of the parameter to be detected, there is a preset second detection time. Based on the preset detection rules and the parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection time, the virtual node is fault detected to determine whether the virtual node is faulty.

[0128] In this step, after the server obtains the parameter values ​​and preset detection rules, it performs fault detection on the virtual node according to the preset detection rules and the parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection time interval, based on the preset detection rules and the parameter values ​​to be detected obtained within the preset second detection time interval, to determine whether the virtual node is faulty.

[0129] It should be noted that the preset second detection duration is an integer multiple of the preset first detection duration. The preset second detection duration can be 0.01 seconds, 0.5 seconds, 3 seconds, etc. This application embodiment does not limit the preset second detection duration, and it can be set according to the actual situation.

[0130] S105: If a virtual node fails, an alarm message is generated, and the alarm message and the virtual node are marked in the first target virtual node call relationship diagram to obtain the second target virtual node call relationship diagram.

[0131] In this step, if the server determines that a virtual node is faulty, it generates an alarm message and marks the alarm message and the virtual node in the first target virtual node call relationship graph to obtain the second target virtual node call relationship graph.

[0132] For example, Figure 1e This is the second target virtual node call relationship diagram provided in this application, such as... Figure 1e As shown, Virtual Elastic Public IP Node A is a faulty node. The fault is caused by the incoming bandwidth exceeding the threshold. Virtual Elastic Public IP Node A is marked, and an alarm message is displayed next to it.

[0133] S106: Output the second target virtual node call relationship diagram.

[0134] In this step, after the server obtains the call relationship diagram of the second target virtual node, in order for the user to intuitively view the faulty node and the cause of the fault, the server needs to output the call relationship diagram of the second target virtual node, that is, send the call relationship diagram of the second target virtual node to the user's terminal device or large screen device for display.

[0135] It should be noted that if the executing entity is a computer or other device with a display, outputting the second target virtual node call relationship diagram means displaying the second target virtual node call relationship diagram on the display.

[0136] It should be noted that the server can also obtain static information for each faulty node, including function type, node name, business tag, resource ID, and inter-node call relationship. The static information is then marked in the second target virtual node call relationship graph, and the second target virtual node call relationship graph marked with static information is then output.

[0137] The fault node detection and alarm method for a business system provided in this embodiment generates a first target virtual node call relationship diagram of the business system. Then, for each parameter to be detected corresponding to each virtual node, a corresponding preset detection rule is obtained, and the parameter value corresponding to the parameter to be detected is obtained once at a preset first detection interval. Then, at a preset second detection interval corresponding to the parameter to be detected, fault detection is performed on the virtual node according to the preset detection rule and the parameter value obtained within the preset second detection interval to determine whether the virtual node is faulty. If the virtual node is faulty, alarm information is generated, and the alarm information and the virtual node are marked in the first target virtual node call relationship diagram to obtain a second target virtual node call relationship diagram. Finally, the second target virtual node call relationship diagram is output. Compared with the prior art that uses manual methods to determine faulty nodes, this solution determines whether the virtual nodes in the business system are faulty based on the parameter values ​​of each virtual node, and outputs the faulty virtual nodes and alarm information after marking them in the virtual node call relationship diagram, allowing users to intuitively view the faulty nodes and effectively improving detection efficiency.

[0138] Figure 2a This is a flowchart illustrating a second embodiment of the fault node detection and alarm method for a business system provided in this application. Based on the above embodiments, this application describes the detection of virtual nodes according to the preset first detection rule when the preset detection rule is the first preset detection rule. For example... Figure 2a As shown, the fault node detection and alarm method of this business system specifically includes the following steps:

[0139] S201: Determine whether the parameter to be detected is marked with an alarm continuity status indicator; if the parameter to be detected is not marked with an alarm continuity status indicator, proceed to step S202; if the parameter to be detected is marked with an alarm continuity status indicator, proceed to step S207.

[0140] In this step, when the preset detection rule is the preset first detection rule, the preset second detection duration is equal to the preset first detection duration. The number of parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration is one. Some parameters to be detected will be marked with an alarm persistence status indicator. There are different processing methods for detection parameters marked with an alarm persistence status indicator and those not marked with an alarm persistence status indicator. Therefore, it is necessary to first determine whether the parameters to be detected are marked with an alarm persistence status indicator.

[0141] S202: Determine whether the parameter value is greater than the preset first alarm parameter threshold; if the parameter value is greater than the preset first alarm parameter threshold, then execute step S203; if the parameter value is less than or equal to the preset first alarm parameter threshold, then execute step S206.

[0142] In this step, if the parameter to be detected is not marked with an alarm persistence status indicator, it is necessary to determine whether the parameter value is greater than the preset first alarm parameter threshold in order to determine whether the parameter value is abnormal.

[0143] It should be noted that the preset first alarm parameter threshold can be 80%, 90%, 95%, 1024KB / S, 2048KB / S, 3072KB / S, etc. This application embodiment does not limit the preset first alarm parameter threshold, and it can be set according to the actual situation.

[0144] S203: Update the cumulative number of consecutive anomalies and determine whether the updated cumulative number of consecutive anomalies is equal to the preset first alarm cumulative threshold; if the updated cumulative number of consecutive anomalies is equal to the preset first alarm cumulative threshold, then execute step S204; if the updated cumulative number of consecutive anomalies is less than the preset first alarm cumulative threshold, then execute step S205.

[0145] S204: Determine if a virtual node is faulty, and mark the parameter to be detected with an alarm persistence status indicator.

[0146] S205: Virtual node confirmed to be normal.

[0147] In the above steps, if the parameter value is greater than the preset first alarm parameter threshold, it indicates that the parameter value is abnormal. Then, the cumulative number of consecutive abnormalities is updated, that is, the cumulative number of consecutive abnormalities is incremented by one, and the virtual node failure is determined based on the updated cumulative number of consecutive abnormalities.

[0148] The system checks whether the updated cumulative number of consecutive anomalies equals the preset first alarm cumulative threshold. If the updated cumulative number of consecutive anomalies equals the preset first alarm cumulative threshold, it indicates a virtual node failure, thus confirming the virtual node failure. Furthermore, since the parameter to be detected is not currently marked with an alarm persistence status indicator, and subsequent determinations of virtual node failure based on the parameter value of this parameter should be implemented using the method used when an alarm persistence status indicator is marked. Therefore, it is necessary to mark the parameter to be detected with an alarm persistence status indicator.

[0149] If the updated cumulative number of consecutive anomalies is less than the preset first alarm cumulative threshold, it indicates that the virtual node is normal, and the virtual node can be determined to be normal.

[0150] It should be noted that the preset first alarm cumulative threshold can be 3, 5, 7, etc. This application embodiment does not limit the preset first alarm cumulative threshold, and it can be set according to the actual situation.

[0151] It should be noted that after marking the alarm persistence status indicator for the parameter to be detected, the system checks whether the parameter is marked with the alarm persistence status indicator at preset reminder intervals. If it is marked, an alarm message is output. If it is not marked, the process stops. The preset reminder interval can be 12 hours, 18 hours, 24 hours, etc. This embodiment does not limit the preset reminder interval and can be set according to actual conditions.

[0152] S206: Determine that the virtual node is normal and update the cumulative number of consecutive anomalies.

[0153] In this step, if the parameter value is less than or equal to the preset first alarm parameter threshold, it indicates that the virtual node is normal, and the virtual node can be confirmed to be normal. It is also necessary to update the cumulative number of consecutive anomalies, that is, to update the cumulative number of consecutive anomalies to 0.

[0154] S207: Determine whether the parameter value is less than or equal to the preset first alarm parameter threshold; if the parameter value is less than or equal to the preset first alarm parameter threshold, then execute step S208; if the parameter value is greater than the preset first alarm parameter threshold, then execute step S211.

[0155] In this step, if the parameter to be detected is marked with an alarm persistence status indicator, it is necessary to determine whether the parameter value is greater than the preset first alarm parameter threshold in order to determine whether the parameter value is abnormal.

[0156] S208: Update the cumulative number of consecutive normal events and determine whether the updated cumulative number of consecutive normal events is equal to the preset first warning elimination threshold; if the updated cumulative number of consecutive normal events is less than the preset first warning elimination threshold, then proceed to step S209; if the updated cumulative number of consecutive normal events is equal to the preset first warning elimination threshold, then proceed to step S210.

[0157] S209: Identify a virtual node failure.

[0158] S210: Determine that the virtual node is normal and remove the alarm persistence status flag from the parameter to be tested.

[0159] In the above steps, if the parameter value is less than or equal to the preset first alarm parameter threshold, it means that the parameter value is normal. Then, the number of consecutive normal cumulative counts is updated, that is, the number of consecutive normal cumulative counts is incremented by one, and the virtual node fault is determined based on the updated number of consecutive normal cumulative counts.

[0160] If the updated cumulative number of consecutive normal occurrences is less than the preset first warning elimination threshold, it indicates that the virtual node is faulty, and the virtual node fault can be determined.

[0161] If the updated cumulative number of consecutive normal occurrences equals the preset first warning elimination threshold, it indicates that the virtual node is normal, and the virtual node can be confirmed to be normal. Additionally, since the parameter to be detected is currently marked with an alarm persistence status indicator, and is no longer in an alarm persistence status indicator, and considering the subsequent determination of whether the virtual node is faulty based on the parameter value of this parameter, the processing method used when the alarm persistence status indicator is not marked should be used. Therefore, it is necessary to remove the alarm persistence status indicator from the parameter to be detected.

[0162] It should be noted that the preset first warning elimination threshold can be 3, 5, 7, etc. This application embodiment does not limit the preset first warning elimination threshold, and it can be set according to the actual situation.

[0163] It should be noted that when the alarm persistence status marker of the parameter to be detected needs to be removed, it means that a second target virtual node call relationship diagram has already been generated, and the virtual node is marked in the second target virtual node call relationship diagram. After removing the alarm persistence status marker of the parameter to be detected, the mark of the virtual node in the second target virtual node call relationship diagram needs to be removed before it is output for user viewing.

[0164] S211: Determine if a virtual node is faulty and update the cumulative number of consecutive normal occurrences.

[0165] In this step, if the parameter value is greater than the preset first alarm parameter threshold, it indicates that the parameter value is abnormal and the alarm is currently in a continuous state, thus confirming a virtual node failure. Simultaneously, it is also necessary to update the cumulative number of consecutive normal occurrences, that is, to update the cumulative number of consecutive normal occurrences to 0.

[0166] For example, Figure 2b An application diagram illustrating the preset first detection rule provided in this application, such as... Figure 2bAs shown, each cell represents a parameter value acquired within a preset second time period. The preset first alarm parameter threshold is 90%, the preset first alarm cumulative threshold is 3, and the preset first alarm cancellation threshold is 3. Initially, the parameter to be detected is not marked with an alarm persistence status indicator. When the first parameter value is acquired, since the parameter is not marked with an alarm persistence status indicator, 85% < 90%, so the virtual node is normal, and the cumulative number of consecutive anomalies is updated to 0. When the second parameter value is acquired, since the parameter is not marked with an alarm persistence status indicator, 91% > 90%, and the cumulative number of consecutive anomalies is updated to 1. Since 1 < 3, the virtual node is normal. When the third parameter value is acquired, since the parameter is not marked with an alarm persistence status indicator, 92% > 90%, and the cumulative number of consecutive anomalies is updated to 2. Since 2 < 3, the virtual node is normal. When the fourth parameter value is acquired, since the parameter is not marked with an alarm persistence status indicator, 95% > 90%, and the cumulative number of consecutive anomalies is updated to 3. Since 3 = 3, the virtual node is faulty, and the parameter to be detected is marked with an alarm persistence status indicator.

[0167] When the fifth parameter value is obtained, because the parameter to be tested is marked with an alarm persistence status indicator, 89% < 90%, and the cumulative number of consecutive normal operations is updated to 1, 1 < 3, therefore the virtual node is faulty. When the sixth parameter value is obtained, because the parameter to be tested is marked with an alarm persistence status indicator, 91% > 90%, therefore the virtual node is faulty, and the cumulative number of consecutive normal operations is updated to 0. When the seventh parameter value is obtained, because the parameter to be tested is marked with an alarm persistence status indicator, 86% < 90%, and the cumulative number of consecutive normal operations is updated to 1, 1 < 3, therefore the virtual node is faulty. When the eighth parameter value is obtained, because the parameter to be tested is marked with an alarm persistence status indicator, 87% < 90%, and the cumulative number of consecutive normal operations is updated to 2, 2 < 3, therefore the virtual node is faulty. When the ninth parameter value is obtained, because the parameter to be tested is marked with an alarm persistence status indicator, 84% < 90%, and the cumulative number of consecutive normal operations is updated to 3, 3 = 3, therefore the virtual node is normal, and the alarm persistence status indicator of the parameter to be tested is removed.

[0168] The fault node detection and alarm method for the business system provided in this embodiment determines whether the parameter to be detected is marked with an alarm persistence status identifier. When the alarm persistence status identifier is not marked, it determines whether the parameter value is greater than a preset first alarm parameter threshold and whether the updated cumulative number of consecutive abnormalities is equal to a preset first alarm cumulative threshold. When the alarm persistence status identifier is marked, it determines whether the parameter value is less than or equal to the preset first alarm parameter threshold and whether the updated cumulative number of consecutive normalities is equal to a preset first warning elimination threshold. This method effectively improves the accuracy of fault detection of virtual nodes.

[0169] Figure 3aThis is a flowchart illustrating a third embodiment of the fault node detection and alarm method for a business system provided in this application. Based on the above embodiments, this application describes the detection of virtual nodes according to the preset second detection rule when the preset detection rule is a preset second detection rule. Figure 3a As shown, the fault node detection and alarm method of this business system specifically includes the following steps:

[0170] S301: Determine whether the parameter to be detected is marked with an alarm persistence status indicator. If the parameter to be detected is not marked with an alarm persistence status indicator, proceed to step S302; if the parameter to be detected is marked with an alarm persistence status indicator, proceed to step S307.

[0171] In this step, when the preset detection rule is the preset second detection rule, the number of parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection time is the target detection number. Some parameters to be detected will be marked with an alarm persistence status indicator. There are different processing methods for detection parameters marked with an alarm persistence status indicator and those not marked with an alarm persistence status indicator. Therefore, it is necessary to first determine whether the parameters to be detected are marked with an alarm persistence status indicator.

[0172] S302: Determine whether the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is greater than the preset third alarm parameter threshold; if the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is greater than the preset third alarm parameter threshold, then proceed to step S303; if the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is less than or equal to the preset third alarm parameter threshold, then proceed to step S306.

[0173] In this step, if the parameter to be detected is not marked with an alarm persistence status indicator, it is necessary to determine whether the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the target detection quantity is greater than the preset third alarm parameter threshold in order to determine whether the parameter value is abnormal.

[0174] It should be noted that the preset second alarm parameter threshold can be 80%, 90%, 95%, 1024KB / S, 2048KB / S, 3072KB / S, etc.; the preset third alarm parameter threshold can be 65%, 70%, 75%, etc. This application embodiment does not limit the preset second alarm parameter threshold and the preset third alarm parameter threshold, and can be set according to the actual situation.

[0175] S303: Update the cumulative number of consecutive anomalies and determine whether the updated cumulative number of consecutive anomalies is equal to the preset second alarm cumulative threshold; if the updated cumulative number of consecutive anomalies is equal to the preset second alarm cumulative threshold, then execute step S304; if the updated cumulative number of consecutive anomalies is less than the preset second alarm cumulative threshold, then execute step S305.

[0176] S304: Determine if a virtual node is faulty and mark the parameter to be detected with an alarm persistence status indicator.

[0177] S305: Virtual node confirmed to be normal.

[0178] In the above steps, if the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is greater than the preset third alarm parameter threshold, it indicates that the parameter value is abnormal. Then, the cumulative number of consecutive abnormalities is updated, that is, the cumulative number of consecutive abnormalities is incremented by one, and the virtual node fault is determined based on the updated cumulative number of consecutive abnormalities.

[0179] The system checks whether the updated cumulative number of consecutive anomalies equals the preset second alarm cumulative threshold. If the updated cumulative number of consecutive anomalies equals the preset second alarm cumulative threshold, it indicates a virtual node failure, thus confirming the virtual node failure. Furthermore, since the parameter to be detected is not currently marked with an alarm persistence status indicator, and subsequent determinations of virtual node failure based on the parameter value should use the processing method for marking alarm persistence status indicators, it is necessary to mark the parameter to be detected with an alarm persistence status indicator.

[0180] If the updated cumulative number of consecutive anomalies is less than the preset second alarm cumulative threshold, it indicates that the virtual node is normal, and the virtual node can be determined to be normal.

[0181] It should be noted that the preset second alarm cumulative threshold can be 2, 3, 5, etc. This application embodiment does not limit the preset second alarm cumulative threshold, and it can be set according to the actual situation.

[0182] It should be noted that after marking the alarm persistence status indicator for the parameter to be detected, the system checks whether the parameter is marked with the alarm persistence status indicator at preset reminder intervals. If it is marked, an alarm message is output. If it is not marked, the process stops. The preset reminder interval can be 12 hours, 18 hours, 24 hours, etc. This embodiment does not limit the preset reminder interval and can be set according to actual conditions.

[0183] S306: Determine that the virtual node is normal and update the cumulative number of consecutive anomalies.

[0184] In this step, if the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is less than or equal to the preset third alarm parameter threshold, it indicates that the virtual node is normal, and the virtual node can be determined to be normal. It is also necessary to update the cumulative number of consecutive anomalies, that is, to update the cumulative number of consecutive anomalies to 0.

[0185] S307: Determine whether the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is less than the preset fourth alarm parameter threshold; if the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is less than the preset fourth alarm parameter threshold, then proceed to step S308; if the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is greater than or equal to the preset fourth alarm parameter threshold, then proceed to step S311.

[0186] In this step, if the parameter to be detected is marked with an alarm continuous status indicator, it is necessary to determine whether the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the target detection quantity is less than the preset fourth alarm parameter threshold in order to determine whether the parameter value is abnormal.

[0187] It should be noted that the preset fourth alarm parameter threshold is less than the preset third alarm parameter threshold; the preset fourth alarm parameter threshold can be 55%, 60%, 65%, etc. This application embodiment does not limit the preset fourth alarm parameter threshold, and it can be set according to the actual situation.

[0188] S308: Update the consecutive normal cumulative count and determine whether the updated consecutive normal cumulative count is equal to the preset second warning elimination threshold; if the updated consecutive normal cumulative count is less than the preset second warning elimination threshold, then execute step S309; ​​if the updated consecutive normal cumulative count is equal to the preset second warning elimination threshold, then execute step S310.

[0189] S309: Identify a virtual node failure.

[0190] S310: Determine that the virtual node is normal and remove the alarm persistence status flag from the parameter to be tested.

[0191] In the above steps, if the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is less than the preset fourth alarm parameter threshold, it indicates that the parameter value is normal. Then, the number of consecutive normal cumulative counts is updated, that is, the number of consecutive normal cumulative counts is incremented by one, and the virtual node fault is determined based on the updated number of consecutive normal cumulative counts.

[0192] If the updated cumulative number of consecutive normal occurrences is less than the preset second warning elimination threshold, it indicates that the virtual node is faulty, and the virtual node fault can be determined.

[0193] If the updated cumulative number of consecutive normal occurrences equals the preset second warning elimination threshold, it indicates that the virtual node is normal, and the virtual node can be confirmed to be normal. Additionally, since the parameter to be detected is currently marked with an alarm persistence status indicator, and is no longer in an alarm persistence status indicator, and considering the subsequent determination of whether the virtual node is faulty based on the parameter value of this parameter, the processing method used when the alarm persistence status indicator is not marked should be used. Therefore, it is necessary to remove the alarm persistence status indicator from the parameter to be detected.

[0194] It should be noted that the preset second warning elimination threshold can be 2, 3, 5, etc. This application embodiment does not limit the preset second warning elimination threshold, and it can be set according to the actual situation.

[0195] It should be noted that when the alarm persistence status marker of the parameter to be detected needs to be removed, it means that a second target virtual node call relationship diagram has already been generated, and the virtual node is marked in the second target virtual node call relationship diagram. After removing the alarm persistence status marker of the parameter to be detected, the mark of the virtual node in the second target virtual node call relationship diagram needs to be removed before it is output for user viewing.

[0196] S311: Determine if a virtual node is faulty and update the cumulative number of consecutive normal occurrences.

[0197] In this step, if the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is greater than or equal to the preset fourth alarm parameter threshold, it indicates that the parameter value is abnormal and the system is currently in an alarm-continuous state, thus confirming a virtual node failure. Simultaneously, the cumulative number of consecutive normal occurrences needs to be updated, specifically to 0.

[0198] For example, Figure 3b An application diagram illustrating the preset second detection rule provided in this application, such as... Figure 3bAs shown, the target detection quantity is 5, and each five cells represent the parameter values ​​acquired within a preset second time period. The preset second alarm parameter threshold is 90%, the preset third alarm parameter threshold is 70%, the preset fourth alarm parameter threshold is 60%, the preset second cumulative alarm threshold is 2, and the preset second warning elimination threshold is 2. Initially, the parameters to be detected are not marked with an alarm persistence status indicator. When the parameter values ​​acquired within the first preset second time period are obtained, since the parameters to be detected are not marked with an alarm persistence status indicator, the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is 4 / 5. 4 / 5 > 70%, so the cumulative number of consecutive anomalies is updated to 1, 1 < 2, confirming the virtual node is normal. When the parameter values ​​acquired within the second preset second time period are obtained, since the parameters to be detected are not marked with an alarm persistence status indicator, the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is 4 / 5, 4 / 5 > 70%, so the cumulative number of consecutive anomalies is updated to 2, 2 = 2, confirming the virtual node is faulty, and the parameter to be detected is marked with an alarm persistence status indicator.

[0199] When the parameter values ​​obtained within the third preset second time period are obtained, since the parameter to be detected is marked with an alarm persistence status identifier, the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is 1 / 5, 1 / 5 < 60%, the cumulative number of consecutive normal operations is updated to 1, 1 < 2, and the virtual node is determined to be faulty. When the parameter values ​​obtained within the fourth preset second time period are obtained, since the parameter to be detected is marked with an alarm persistence status identifier, the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is 1 / 5, 1 / 5 < 60%, the cumulative number of consecutive normal operations is updated to 2, 2 = 2, the virtual node is determined to be normal, and the alarm persistence status identifier marked on the parameter to be detected is removed.

[0200] The fault node detection and alarm method for a business system provided in this embodiment determines whether a virtual node is faulty by determining whether the parameter to be detected is marked with an alarm persistence status identifier. When the alarm persistence status identifier is not marked, it determines whether the ratio of the number of parameter values ​​greater than a preset second alarm parameter threshold to the target detection number is greater than a preset third alarm parameter threshold, and whether the updated cumulative number of consecutive anomalies is equal to a preset second cumulative alarm threshold. When the alarm persistence status identifier is marked, it determines whether the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the target detection number is less than a preset fourth alarm parameter threshold, and whether the updated cumulative number of consecutive normal events is equal to a preset second warning elimination threshold. This method effectively improves the accuracy of fault detection for virtual nodes.

[0201] Figure 4This is a flowchart illustrating Embodiment 4 of the fault node detection and alarm method for the business system provided in this application. Based on the above embodiments, this application embodiment explains the detection of virtual nodes according to the preset third detection rule when the preset detection rule is the preset third detection rule. For example... Figure 4 As shown, the fault node detection and alarm method of this business system specifically includes the following steps:

[0202] S401: Get the current month and the current date within the current month.

[0203] In this step, when the preset detection rule is the third preset detection rule, the preset second detection duration is equal to the preset first detection duration, and the number of parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration is one. To determine whether a virtual node is faulty, it is necessary to obtain the current month and the current date within that month.

[0204] For example, the current month and the current date under the current month can be in the form of June 5, June 18, November 11, etc.

[0205] S402: Based on the current month and the current date within the current month, as well as the correspondence between the month / date and the alarm threshold, determine the preset fifth alarm threshold corresponding to the current month and the current date within the current month.

[0206] In this step, after the server obtains the current month and the current date within the current month, since the staff has evaluated a large amount of historical monitoring data and alarm data, determined the correspondence between the month and date and the alarm threshold, and set it in the server, the preset fifth alarm threshold corresponding to the current month and the current date within the current month can be determined based on the correspondence between the month and date and the alarm threshold.

[0207] For example, the preset fifth alarm threshold for June 5th is 70%, the preset fifth alarm threshold for June 18th is 90%, and the preset fifth alarm threshold for November 11th is 95%, etc. This application does not limit the preset fifth alarm threshold; it can be determined according to actual circumstances.

[0208] S403: Determine whether the parameter value is greater than the preset fifth alarm threshold; if the parameter value is greater than the preset fifth alarm threshold, proceed to step S404; if the parameter value is less than or equal to the preset fifth alarm threshold, proceed to step S405.

[0209] S404: Virtual node failure detected.

[0210] S405: Virtual node confirmed to be normal.

[0211] In the above steps, after the server obtains the preset fifth alarm threshold, it determines whether the parameter value is abnormal by checking if it exceeds the preset fifth alarm threshold. If the parameter value is greater than the preset fifth alarm threshold, it indicates that the parameter value is abnormal, and a virtual node failure can be identified.

[0212] If the parameter value is less than or equal to the preset fifth alarm threshold, it indicates that the parameter value is normal, and the virtual node is confirmed to be normal.

[0213] It should be noted that if a virtual node is determined to be normal, but a virtual node is determined to be faulty based on the parameter values ​​within the previous preset second time period, it means that a second target virtual node call relationship diagram has already been generated, and the virtual node is marked in the second target virtual node call relationship diagram. After determining that the virtual node is normal, the mark of the virtual node in the second target virtual node call relationship diagram needs to be removed before outputting it for user viewing.

[0214] The fault node detection and alarm method for the business system provided in this embodiment determines the corresponding preset fifth alarm threshold based on the current month and the current date under the current month, and determines whether the virtual node is faulty based on the relationship between the parameter value and the preset fifth alarm threshold, which effectively improves the accuracy of fault detection of virtual nodes.

[0215] Figure 5 This is a flowchart illustrating Embodiment 5 of the fault node detection and alarm method for the business system provided in this application. Based on the above embodiments, this application embodiment describes the detection of virtual nodes according to the preset fourth detection rule when the preset detection rule is the preset fourth detection rule. Figure 5 As shown, the fault node detection and alarm method of this business system specifically includes the following steps:

[0216] S501: Get the current month and the current date within the current month.

[0217] In this step, when the preset detection rule is the third preset detection rule, the preset second detection duration is equal to the preset first detection duration, and the number of parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration is one. To determine whether a virtual node is faulty, it is necessary to obtain the current month and the current date within that month.

[0218] For example, the current month and the current date under the current month can be in the form of June 5, June 18, November 11, etc.

[0219] S502: Input the current month and the current date of the current month into the preset alarm threshold prediction model to determine the preset sixth alarm threshold.

[0220] In this step, after the server obtains the current month and the current date within the current month, a preset alarm threshold prediction model is obtained by training based on a large amount of monthly and date data and alarm thresholds. Therefore, the current month and the current date within the current month can be input into the preset alarm threshold prediction model to determine the preset sixth alarm threshold. The preset alarm threshold prediction model is a pre-trained calculation model used to determine the alarm threshold based on the month and date.

[0221] S503: Determine whether the parameter value is greater than the preset sixth alarm threshold; if the parameter value is greater than the preset sixth alarm threshold, proceed to step S504; if the parameter value is less than or equal to the preset sixth alarm threshold, proceed to step S505.

[0222] S504: Virtual node failure detected.

[0223] S505: Virtual node confirmed to be normal.

[0224] In the above steps, after the server obtains the preset sixth alarm threshold, it determines whether the parameter value is abnormal by checking if it exceeds the preset sixth alarm threshold. If the parameter value is greater than the preset sixth alarm threshold, it indicates that the parameter value is abnormal, and a virtual node failure can be identified.

[0225] If the parameter value is less than or equal to the preset sixth alarm threshold, it indicates that the parameter value is normal, and the virtual node is confirmed to be normal.

[0226] It should be noted that if a virtual node is determined to be normal, but a virtual node is determined to be faulty based on the parameter values ​​within the previous preset second time period, it means that a second target virtual node call relationship diagram has already been generated, and the virtual node is marked in the second target virtual node call relationship diagram. After determining that the virtual node is normal, the mark of the virtual node in the second target virtual node call relationship diagram needs to be removed before outputting it for user viewing.

[0227] The fault node detection and alarm method for the business system provided in this embodiment determines the corresponding preset sixth alarm threshold based on the current month, the current date under the current month, and the preset alarm threshold prediction model. Based on the relationship between the parameter value and the preset sixth alarm threshold, it determines whether the virtual node is faulty, which effectively improves the accuracy of fault detection of virtual nodes.

[0228] Figure 6a This is a flowchart illustrating Embodiment Six of the Fault Node Detection and Alarm Method for the Business System Provided in this Application. Based on the above embodiments, this embodiment further explains the situation where, after the server identifies a faulty node, some virtual nodes will be affected due to the virtual node failure, thus identifying the affected virtual nodes. For example... Figure 6aAs shown, the fault node detection and alarm method of this business system specifically includes the following steps:

[0229] S601: For each faulty virtual node, based on the second target virtual node call relationship diagram, and / or, the preset affected node prediction model, and / or, the correspondence between virtual nodes and affected nodes, determine the affected virtual node corresponding to each faulty virtual node.

[0230] In this step, after the server identifies the faulty node, it's crucial to identify other nodes as well, since their failure will affect them. For each faulty virtual node, the affected virtual nodes are determined based on the relationship graph of the second target virtual node, and / or the pre-set affected node prediction model, and / or the correspondence between virtual nodes and affected nodes. The pre-set affected node prediction model is a pre-trained computational model used to determine affected nodes based on virtual nodes.

[0231] The implementation method for determining the affected virtual nodes based on the call relationship diagram of the second target virtual node can be as follows: for each faulty virtual node, determine the first process virtual node that has a call relationship with the faulty virtual node, determine the second process node that the call relationship points to from these first process virtual nodes, and determine the determined second process node as the affected virtual node.

[0232] For example, in Figure 1e Based on this, virtual elastic public IP node A is the faulty node, and the virtual nodes that have a calling relationship with it are virtual server D and virtual server E. Both of these are nodes that the calling relationship points to, so virtual server D and virtual server E are the affected virtual nodes.

[0233] The implementation method for determining the affected virtual nodes based on the preset affected node prediction model can be as follows: since the preset affected node prediction model is a model trained by staff based on a large number of faulty nodes and affected virtual nodes, the faulty nodes can be input into the preset affected node prediction model to obtain the affected virtual nodes.

[0234] The implementation method for determining the affected virtual nodes based on the correspondence between virtual nodes and affected nodes can be as follows: since the correspondence between virtual nodes and affected nodes is obtained by staff through evaluation based on a large number of faulty nodes and affected virtual nodes, the corresponding affected virtual nodes can be obtained based on the faulty nodes.

[0235] S602: Mark the affected virtual nodes in the second target virtual node call relationship diagram to obtain the third target virtual node call relationship diagram.

[0236] S603: Output the call relationship diagram of the third target virtual node.

[0237] In the above steps, after the server identifies the affected virtual nodes, in order for users to view the affected virtual nodes intuitively, the affected virtual nodes will be marked in the second target virtual node call relationship graph to obtain the third target virtual node call relationship graph, and then the third target virtual node call relationship graph will be output.

[0238] For example, Figure 6b The virtual node call relationship diagram provided for the third target in this application is as follows: Figure 6b As shown, virtual server D and virtual server E are the affected virtual nodes, and they are marked.

[0239] It should be noted that when removing the mark from a faulty node, the mark from the corresponding affected virtual node is also removed.

[0240] The fault node detection and alarm method for business systems provided in this embodiment identifies the affected virtual nodes of the fault node, marks them in the virtual node call relationship graph, and outputs the results. This allows users to intuitively see which virtual nodes are affected and take timely protective measures, effectively improving the security of the business system.

[0241] Figure 7 This is a flowchart illustrating Embodiment Seven of the fault node detection and alarm method for the business system provided in this application. Based on the above embodiments, this application embodiment explains the generation of a virtual node baseline call relationship diagram. Figure 7 As shown, the fault node detection and alarm method of this business system specifically includes the following steps:

[0242] S701: Retrieve all baseline function types from the virtual node function type library.

[0243] In this step, since a baseline virtual node call relationship diagram is needed during the generation of the first target virtual node call relationship diagram, it is necessary to generate the baseline virtual node call relationship diagram before the server obtains the function type corresponding to each virtual node in the business system. First, all baseline function types are retrieved from the virtual node function type library.

[0244] Since staff will add the baseline function types to the virtual node function type library for storage, all baseline function types can be obtained. Baseline function types can be server function types, gateway function types, hard disk function types, database function types, etc. This application embodiment does not limit the baseline function types; they can be determined according to the actual situation.

[0245] S702: Retrieve all baseline call relationships from the call relationship database, as well as the baseline function type corresponding to each baseline call relationship.

[0246] In this step, after the server obtains the baseline function type, it retrieves all baseline call relationships from the call relationship database, as well as the baseline function type corresponding to each baseline call relationship.

[0247] Because staff will add the baseline call relationships and the corresponding baseline function types to the call relationship database for storage, all baseline call relationships and the corresponding baseline function types for each baseline call relationship can be obtained. All types of baseline call relationships include one or more of the following: inclusion relationship, access relationship, connection relationship, binding relationship, and mounting relationship. This application embodiment does not limit the types of baseline call relationships and can be set according to actual circumstances.

[0248] S703: Generate a virtual node baseline call relationship diagram based on all baseline function types, all baseline call relationships, and the baseline function type corresponding to each baseline call relationship.

[0249] In this step, after the server obtains the baseline function type, the baseline call relationship, and the baseline function type corresponding to each baseline call relationship, it can generate a virtual node baseline call relationship graph based on all baseline function types, all baseline call relationships, and the baseline function type corresponding to each baseline call relationship.

[0250] It should be noted that the virtual node baseline call relationship graph can be generated using WebGL, Echarts, or other methods. This application embodiment does not limit the method of generating the virtual node baseline call relationship graph, and the method can be selected according to the actual situation.

[0251] It should be noted that after generating the virtual node baseline call relationship diagram, it can be output for users to view. Users can also directly drag, pull, and tug on the virtual node baseline call relationship diagram.

[0252] The fault node detection and alarm method for the business system provided in this embodiment generates a virtual node baseline call relationship diagram based on all baseline function types, all baseline call relationships, and the baseline function type corresponding to each baseline call relationship. This makes the virtual node baseline call relationship diagram more accurate and also makes the generation of the virtual node baseline call relationship diagram more efficient.

[0253] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.

[0254] Figure 8This is a schematic diagram of the structure of an embodiment of the fault node detection and alarm device for the business system provided in this application. Figure 8 As shown, the fault node detection and alarm device 80 of this business system includes:

[0255] Module 81 is used to obtain the function type corresponding to each virtual node in the business system;

[0256] The generation module 82 is used to generate the first target virtual node call relationship diagram of the business system according to the functional type of each virtual node;

[0257] The acquisition module 81 is also used to acquire the preset detection rule corresponding to each parameter to be detected for each virtual node, and acquire the parameter value corresponding to the parameter to be detected once every preset first detection time interval;

[0258] Processing module 83 is used to perform fault detection on the virtual node at intervals corresponding to a preset second detection duration for each parameter to be detected, based on the preset detection rules and the parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration, to determine whether the virtual node is faulty; the preset second detection duration is an integer multiple of the preset first detection duration;

[0259] The generation module 82 is further configured to generate alarm information if the virtual node fails, and mark the alarm information and the virtual node in the first target virtual node call relationship diagram to obtain the second target virtual node call relationship diagram;

[0260] Output module 84 is used to output the call relationship diagram of the second target virtual node.

[0261] Furthermore, the generation module 82 is specifically used for:

[0262] Based on the functional type of each virtual node, the initial call relationship between each virtual node is determined from the pre-generated virtual node baseline call relationship graph;

[0263] Generate an initial virtual node call relationship graph based on the initial call relationships between each virtual node;

[0264] In response to the user's operation to adjust the relationship graph, a relationship graph for the first target virtual node is generated.

[0265] Further, if the preset detection rule is a preset first detection rule, then the preset second detection duration is equal to the preset first detection duration, and the processing module 83 is specifically used for:

[0266] Determine whether the parameter to be detected is marked with an alarm persistence status identifier;

[0267] If the parameter to be detected is not marked with the alarm continuity status identifier, then it is determined whether the parameter value is greater than the preset first alarm parameter threshold.

[0268] If the parameter value is greater than the preset first alarm parameter threshold, then update the cumulative number of consecutive anomalies;

[0269] Determine whether the updated cumulative number of consecutive anomalies is equal to the preset first alarm cumulative threshold;

[0270] If the updated cumulative number of consecutive anomalies is equal to the preset first alarm cumulative threshold, then the virtual node is determined to be faulty, and the alarm persistence status identifier is marked on the parameter to be detected.

[0271] If the updated cumulative number of consecutive anomalies is less than the preset first alarm cumulative threshold, then the virtual node is determined to be normal.

[0272] If the parameter value is less than or equal to the preset first alarm parameter threshold, then the virtual node is determined to be normal, and the cumulative number of consecutive anomalies is updated.

[0273] If the parameter to be detected marks the alarm persistence status identifier, then determine whether the parameter value is less than or equal to the preset first alarm parameter threshold.

[0274] If the parameter value is less than or equal to the preset first alarm parameter threshold, then update the cumulative number of consecutive normal alarms;

[0275] Determine whether the updated cumulative number of consecutive normal occurrences is equal to the preset first warning elimination threshold;

[0276] If the updated cumulative number of consecutive normal occurrences is less than the preset first warning elimination threshold, then the virtual node is determined to be faulty.

[0277] If the updated cumulative number of consecutive normal occurrences is equal to the preset first warning elimination threshold, then the virtual node is determined to be normal, and the alarm persistence status identifier marked on the parameter to be detected is removed.

[0278] If the parameter value is greater than the preset first alarm parameter threshold, then the virtual node is determined to be faulty, and the cumulative number of consecutive normal operations is updated.

[0279] Further, if the preset detection rule is a preset second detection rule, then the number of parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection time is the target detection number. The processing module 83 is specifically used for:

[0280] Determine whether the parameter to be detected is marked with an alarm persistence status identifier;

[0281] If the parameter to be detected is not marked with the alarm continuity status identifier, then determine whether the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the target detection quantity is greater than the preset third alarm parameter threshold.

[0282] If the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is greater than the preset third alarm parameter threshold, then the cumulative number of consecutive anomalies is updated.

[0283] Determine whether the updated cumulative number of consecutive anomalies is equal to the preset second alarm cumulative threshold;

[0284] If the updated cumulative number of consecutive anomalies is equal to the preset second alarm cumulative threshold, then the virtual node is determined to be faulty, and the alarm persistence status identifier is marked on the parameter to be detected.

[0285] If the updated cumulative number of consecutive anomalies is less than the preset second alarm cumulative threshold, then the virtual node is determined to be normal.

[0286] If the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is less than or equal to the preset third alarm parameter threshold, then the virtual node is determined to be normal, and the cumulative number of consecutive anomalies is updated.

[0287] If the parameter to be detected is not marked with the alarm continuity status identifier, then it is determined whether the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the target detection quantity is less than the preset fourth alarm parameter threshold, and the preset fourth alarm parameter threshold is less than the preset third alarm parameter threshold.

[0288] If the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is less than the preset fourth alarm parameter threshold, then the cumulative number of consecutive normal occurrences is updated.

[0289] Determine whether the updated consecutive normal cumulative count is equal to the preset second warning elimination threshold;

[0290] If the updated cumulative number of consecutive normal occurrences is less than the preset second warning elimination threshold, then the virtual node is determined to be faulty.

[0291] If the updated cumulative number of consecutive normal occurrences is equal to the preset second warning elimination threshold, then the virtual node is determined to be normal, and the alarm persistence status identifier marked on the parameter to be detected is removed.

[0292] If the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is greater than or equal to the preset fourth alarm parameter threshold, then the virtual node is determined to be faulty, and the cumulative number of consecutive normal operations is updated.

[0293] Furthermore, if the preset detection rule is a preset third detection rule, then the preset second detection duration is equal to the preset first detection duration, and the processing module 83 is specifically used for:

[0294] Get the current month and the current date within the current month;

[0295] Based on the current month and the current date within the current month, as well as the correspondence between the month / date and the alarm threshold, a preset fifth alarm threshold corresponding to the current month and the current date within the current month is determined;

[0296] If the parameter value is greater than the preset fifth alarm threshold, then the virtual node is determined to be faulty;

[0297] If the parameter value is less than or equal to the preset fifth alarm threshold, then the virtual node is determined to be normal.

[0298] Furthermore, if the preset detection rule is a preset third detection rule, then the preset second detection duration is equal to the preset first detection duration, and the processing module 83 is specifically used for:

[0299] Get the current month and the current date within the current month;

[0300] The current month and the current date under the current month are input into the preset alarm threshold prediction model to determine the preset sixth alarm threshold; the preset alarm threshold prediction model is a pre-trained calculation model for determining the alarm threshold based on the month and date;

[0301] If the parameter value is greater than the preset sixth alarm threshold, then the virtual node is determined to be faulty;

[0302] If the parameter value is less than or equal to the preset sixth alarm threshold, then the virtual node is determined to be normal.

[0303] Furthermore, the processing module 83 is also used for:

[0304] For each faulty virtual node, the affected virtual node corresponding to each faulty virtual node is determined by the second target virtual node call relationship graph, and / or the preset affected node prediction model, and / or the correspondence between virtual nodes and affected nodes; the preset affected node prediction model is a pre-trained computational model used to determine the affected node based on the virtual node.

[0305] The affected virtual nodes are marked in the second target virtual node call relationship graph to obtain the third target virtual node call relationship graph;

[0306] Furthermore, the output module 84 is also used to output the call relationship diagram of the third target virtual node.

[0307] Furthermore, after generating the first target virtual node call relationship diagram of the business system, the output module 84 is also used to output the first target virtual node call relationship diagram.

[0308] Furthermore, before obtaining the function type corresponding to each virtual node in the business system, the acquisition module 81 is also used for:

[0309] Retrieve all baseline function types from the virtual node function type library;

[0310] Retrieve all baseline call relationships from the call relationship database, as well as the baseline function type corresponding to each baseline call relationship. All types of baseline call relationships include one or more of the following: containment relationship, access relationship, access relationship, binding relationship, and mount relationship.

[0311] Furthermore, the generation module 82 is also used to generate the virtual node baseline call relationship diagram based on all baseline function types, all baseline call relationships, and the baseline function type corresponding to each baseline call relationship.

[0312] The fault node detection and alarm device for the business system provided in this embodiment is used to execute the technical solution in any of the aforementioned method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0313] Figure 9 This is a schematic diagram of the structure of an electronic device provided in this application. Figure 9 As shown, the electronic device 90 includes:

[0314] Processor 91, memory 92, and communication interface 93;

[0315] The memory 92 is used to store the executable instructions of the processor 91;

[0316] The processor 91 is configured to execute the technical solutions in any of the foregoing method embodiments by executing the executable instructions.

[0317] Optionally, the memory 92 can be either standalone or integrated with the processor 91.

[0318] Optionally, when the memory 92 is a device independent of the processor 91, the electronic device 90 may further include:

[0319] Bus 94, memory 92 and communication interface 93 are connected to processor 91 through bus 94 and complete communication with each other. Communication interface 93 is used to communicate with other devices.

[0320] Optionally, the communication interface 93 can be implemented using a transceiver. The communication interface is used to enable communication between the database access device and other devices (e.g., clients, read-write databases, and read-only databases). The memory may include random access memory (RAM) and may also include non-volatile memory, such as at least one disk drive.

[0321] Bus 94 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one thick line is used in the diagram, but this does not indicate that there is only one bus or one type of bus.

[0322] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0323] The electronic device is used to execute the technical solutions in any of the foregoing method embodiments. Its implementation principle and technical effect are similar, and will not be described again here.

[0324] This application also provides a readable storage medium storing a computer program thereon, which, when executed by a processor, implements the technical solutions provided in any of the foregoing embodiments.

[0325] This application also provides a computer program product, including a computer program, which, when executed by a processor, is used to implement the technical solutions provided in any of the foregoing method embodiments.

[0326] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0327] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A method for fault node detection and alarm in a business system, characterized in that, include: Obtain the function type corresponding to each virtual node in the business system; Retrieve all baseline function types from the virtual node function type library; Retrieve all baseline call relationships from the call relationship database, as well as the baseline function type corresponding to each baseline call relationship. All types of baseline call relationships include one or more of the following: containment relationship, access relationship, access relationship, binding relationship, and mount relationship. Based on all baseline function types, all baseline call relationships, and the baseline function type corresponding to each baseline call relationship, the virtual node baseline call relationship graph is generated; Based on the functional type of each virtual node, the initial call relationship between each virtual node is determined from the pre-generated virtual node baseline call relationship graph; Generate an initial virtual node call relationship graph based on the initial call relationships between each virtual node; In response to the user's operation to adjust the relationship graph, the first target virtual node is generated to call the relationship graph; For each parameter to be detected corresponding to each virtual node, obtain the preset detection rule corresponding to the parameter to be detected, and obtain the parameter value corresponding to the parameter to be detected once every preset first detection time interval; For each interval of the parameter to be detected, a preset second detection time is defined. Based on the preset detection rules and the parameter values ​​corresponding to the parameter to be detected obtained within the preset second detection time, the virtual node is subjected to fault detection to determine whether the virtual node is faulty. The preset second detection duration is an integer multiple of the preset first detection duration; If the virtual node fails, an alarm message is generated, and the alarm message and the virtual node are marked in the first target virtual node call relationship diagram to obtain the second target virtual node call relationship diagram; Output the call relationship diagram of the second target virtual node.

2. The method according to claim 1, characterized in that, If the preset detection rule is a preset first detection rule, then the preset second detection duration is equal to the preset first detection duration. The step of performing fault detection on the virtual node based on the preset detection rule and the parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration, to determine whether the virtual node is faulty, includes: Determine whether the parameter to be detected is marked with an alarm persistence status identifier; If the parameter to be detected is not marked with the alarm continuity status identifier, then it is determined whether the parameter value is greater than the preset first alarm parameter threshold. If the parameter value is greater than the preset first alarm parameter threshold, then update the cumulative number of consecutive anomalies; Determine whether the updated cumulative number of consecutive anomalies is equal to the preset first alarm cumulative threshold; If the updated cumulative number of consecutive anomalies is equal to the preset first alarm cumulative threshold, then the virtual node is determined to be faulty, and the alarm persistence status identifier is marked on the parameter to be detected. If the updated cumulative number of consecutive anomalies is less than the preset first alarm cumulative threshold, then the virtual node is determined to be normal. If the parameter value is less than or equal to the preset first alarm parameter threshold, then the virtual node is determined to be normal, and the cumulative number of consecutive anomalies is updated. If the parameter to be detected marks the alarm persistence status identifier, then determine whether the parameter value is less than or equal to the preset first alarm parameter threshold. If the parameter value is less than or equal to the preset first alarm parameter threshold, then update the cumulative number of consecutive normal alarms; Determine whether the updated cumulative number of consecutive normal occurrences is equal to the preset first warning elimination threshold; If the updated cumulative number of consecutive normal occurrences is less than the preset first warning elimination threshold, then the virtual node is determined to be faulty. If the updated cumulative number of consecutive normal occurrences is equal to the preset first warning elimination threshold, then the virtual node is determined to be normal, and the alarm persistence status identifier marked on the parameter to be detected is removed. If the parameter value is greater than the preset first alarm parameter threshold, then the virtual node is determined to be faulty, and the cumulative number of consecutive normal operations is updated.

3. The method according to claim 1, characterized in that, If the preset detection rule is a preset second detection rule, then the number of parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration is the target detection quantity. The step of performing fault detection on the virtual node based on the preset detection rule and the parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration to determine whether the virtual node is faulty includes: Determine whether the parameter to be detected is marked with an alarm persistence status identifier; If the parameter to be detected is not marked with the alarm continuity status identifier, then determine whether the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the target detection quantity is greater than the preset third alarm parameter threshold. If the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is greater than the preset third alarm parameter threshold, then the cumulative number of consecutive anomalies is updated. Determine whether the updated cumulative number of consecutive anomalies is equal to the preset second alarm cumulative threshold; If the updated cumulative number of consecutive anomalies is equal to the preset second alarm cumulative threshold, then the virtual node is determined to be faulty, and the alarm persistence status identifier is marked on the parameter to be detected. If the updated cumulative number of consecutive anomalies is less than the preset second alarm cumulative threshold, then the virtual node is determined to be normal. If the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is less than or equal to the preset third alarm parameter threshold, then the virtual node is determined to be normal, and the cumulative number of consecutive anomalies is updated. If the parameter to be detected is not marked with the alarm continuity status identifier, then it is determined whether the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the target detection quantity is less than the preset fourth alarm parameter threshold, and the preset fourth alarm parameter threshold is less than the preset third alarm parameter threshold. If the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is less than the preset fourth alarm parameter threshold, then the cumulative number of consecutive normal occurrences is updated. Determine whether the updated consecutive normal cumulative count is equal to the preset second warning elimination threshold; If the updated cumulative number of consecutive normal occurrences is less than the preset second warning elimination threshold, then the virtual node is determined to be faulty. If the updated cumulative number of consecutive normal occurrences is equal to the preset second warning elimination threshold, then the virtual node is determined to be normal, and the alarm persistence status identifier marked on the parameter to be detected is removed. If the ratio of the number of parameter values ​​greater than the preset second alarm parameter threshold to the number of target detections is greater than or equal to the preset fourth alarm parameter threshold, then the virtual node is determined to be faulty, and the cumulative number of consecutive normal operations is updated.

4. The method according to claim 1, characterized in that, If the preset detection rule is a preset third detection rule, then the preset second detection duration is equal to the preset first detection duration. The step of performing fault detection on the virtual node based on the preset detection rule and the parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration, to determine whether the virtual node is faulty, includes: Get the current month and the current date within the current month; Based on the current month and the current date within the current month, as well as the correspondence between the month / date and the alarm threshold, a preset fifth alarm threshold corresponding to the current month and the current date within the current month is determined; If the parameter value is greater than the preset fifth alarm threshold, then the virtual node is determined to be faulty; If the parameter value is less than or equal to the preset fifth alarm threshold, then the virtual node is determined to be normal.

5. The method according to claim 1, characterized in that, If the preset detection rule is a preset fourth detection rule, then the preset second detection duration is equal to the preset first detection duration. The step of performing fault detection on the virtual node based on the preset detection rule and the parameter values ​​corresponding to the parameters to be detected obtained within the preset second detection duration, to determine whether the virtual node is faulty, includes: Get the current month and the current date within the current month; The current month and the current date under the current month are input into the preset alarm threshold prediction model to determine the preset sixth alarm threshold; the preset alarm threshold prediction model is a pre-trained calculation model for determining the alarm threshold based on the month and date; If the parameter value is greater than the preset sixth alarm threshold, then the virtual node is determined to be faulty; If the parameter value is less than or equal to the preset sixth alarm threshold, then the virtual node is determined to be normal.

6. The method according to claim 1, characterized in that, The method further includes: For each faulty virtual node, the affected virtual node corresponding to each faulty virtual node is determined by the second target virtual node call relationship graph, and / or the preset affected node prediction model, and / or the correspondence between virtual nodes and affected nodes; the preset affected node prediction model is a pre-trained computational model used to determine the affected node based on the virtual node. The affected virtual nodes are marked in the second target virtual node call relationship graph to obtain the third target virtual node call relationship graph; Output the call relationship diagram of the third target virtual node.

7. The method according to claim 1, characterized in that, After generating the first target virtual node call relationship graph of the business system, the method further includes: Output the call relationship graph of the first target virtual node.

8. A fault node detection and alarm device for a business system, characterized in that, include: The acquisition module is used to acquire the function type corresponding to each virtual node in the business system. The acquisition module is also used to acquire all baseline function types from the virtual node function type library; acquire all baseline call relationships from the call relationship library, and the baseline function type corresponding to each baseline call relationship. The types of all baseline call relationships include one or more of the following: inclusion relationship, access relationship, access relationship, binding relationship, and mounting relationship. The generation module is used to generate the virtual node baseline call relationship diagram based on all baseline function types, all baseline call relationships, and the baseline function type corresponding to each baseline call relationship. Based on the functional type of each virtual node, the initial call relationship between each virtual node is determined from the pre-generated virtual node baseline call relationship graph; based on the initial call relationship between each virtual node, an initial virtual node call relationship graph is generated; in response to the user's operation to adjust the relationship graph, a first target virtual node call relationship graph is generated. The acquisition module is also used to acquire the preset detection rule corresponding to each parameter to be detected for each virtual node, and acquire the parameter value corresponding to the parameter to be detected once every preset first detection time interval; The processing module is used to perform fault detection on the virtual node at intervals corresponding to a preset second detection duration for the parameter to be detected, based on the preset detection rules and the parameter values ​​corresponding to the parameter to be detected obtained within the preset second detection duration, to determine whether the virtual node is faulty; the preset second detection duration is an integer multiple of the preset first detection duration; The generation module is further configured to generate alarm information if the virtual node fails, and mark the alarm information and the virtual node in the first target virtual node call relationship diagram to obtain the second target virtual node call relationship diagram; The output module is used to output the call relationship diagram of the second target virtual node.

9. An electronic device, characterized in that, include: Processor, memory, communication interface; The memory is used to store the executable instructions of the processor; The processor is configured to execute the fault node detection and alarm method of the business system according to any one of claims 1 to 7 by executing the executable instructions.

10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the fault node detection and alarm method of the business system according to any one of claims 1 to 7.