Fault alarm processing method and device, electronic equipment and storage medium

By aggregating fault alarm information based on network topology in the application system, the problem of alarm storms is solved, enabling rapid and accurate identification of fault causes and reducing maintenance workload.

CN116545835BActive Publication Date: 2026-05-15INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2023-06-07
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

During the fault monitoring and alarm process of application systems, alarm storms can easily occur, leading to chaotic fault alarm information and making it impossible to quickly and accurately determine the root cause of the fault.

Method used

By acquiring fault alarm information from the application system and determining whether there are network topology branches based on the network topology structure, if so, the fault alarm information is aggregated with the fault alarm information in the network topology branches to obtain aggregated fault alarm information, thereby determining the root cause of the fault.

Benefits of technology

It reduces the likelihood of alarm storms, simplifies fault alarm information, helps operations and maintenance personnel quickly and accurately identify the causes of application system failures, and reduces workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116545835B_ABST
    Figure CN116545835B_ABST
Patent Text Reader

Abstract

The application provides a fault alarm processing method and device, electronic equipment and storage medium, relates to the field of financial technology and other related fields, and the method comprises the following steps: acquiring first fault alarm information of a first object of an application system; determining whether at least one network topology branch of the first object exists according to the first fault alarm information and a network topology structure of the application system; each network topology branch comprises a second object located at a level above the first object and having a network topology relationship with the first object; if the network topology branch exists, determining whether second fault alarm information of the second object in the network topology branch already exists; if the second fault alarm information already exists, aggregating the first fault alarm information and the second fault alarm information of the second object in the network topology branch to obtain aggregated fault alarm information. The method provided in the application reduces the possibility of an alarm storm caused by fault alarm of the application system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of financial technology and other related fields, and in particular to a fault alarm processing method, device, electronic device and storage medium. Background Technology

[0002] During the operation of an application system, the supporting hardware and software, such as servers, switches, operating systems deployed on servers, and application software, may all experience failures. To ensure the normal operation of the application system, current practices typically involve monitoring and issuing alerts for potential failures in the hardware and software during application system operation. This allows for control over the application system's operational status and timely handling of any failures that occur, ensuring the normal operation of the application system.

[0003] However, in the current process of fault monitoring and alarming in application systems, there is a problem of alarm storms, that is, a large number of fault alarm messages can easily appear in a short period of time, resulting in chaotic fault alarm information and making it impossible to quickly and accurately determine the root cause of the fault. Summary of the Invention

[0004] This application provides a fault alarm processing method, apparatus, electronic device, and storage medium to solve the problem of alarm storms that can easily occur during the fault monitoring and alarm process of an application system.

[0005] Firstly, this application provides a task processing method, the method comprising:

[0006] Obtain the first fault alarm information of the first object in the application system;

[0007] Based on the first fault alarm information and the network topology of the application system, it is determined whether there is at least one network topology branch of the first object; each network topology branch includes: a second object located at a level above the level of the first object and having a network topology relationship with the first object;

[0008] If it exists, determine whether the second fault alarm information of the second object in the network topology branch already exists;

[0009] If it already exists, the first fault alarm information is aggregated with the second fault alarm information of the second object in the network topology branch to obtain the aggregated fault alarm information.

[0010] Optionally, the aggregation of the first fault alarm information with the second fault alarm information of the second object in the network topology branch includes:

[0011] If there are second objects at at least two levels in the network topology branch, then determine the third device at the highest level from the second objects at at least two levels;

[0012] The existing fault alarm information of the remaining second objects, as well as the first fault alarm information, are aggregated into the third fault alarm information of the third device to obtain the aggregated fault alarm information.

[0013] Optionally, the step of aggregating the existing fault alarm information of the remaining second objects and the first fault alarm information into the fault alarm information of the third device to obtain the aggregated fault alarm information includes:

[0014] The existing fault alarm information of the remaining second objects, as well as the fault alarm information in the first fault alarm information whose generation time difference with the third fault alarm information is less than or equal to a preset duration, are aggregated into the third fault alarm information to obtain the aggregated fault alarm information.

[0015] Optionally, after determining whether at least one network topology branch of the first object exists, the method further includes:

[0016] If it does not exist, determine whether the historical first fault alarm information of the first object already exists;

[0017] If it exists, the historical first fault alarm information whose time difference with the generation time of the first fault alarm information is less than or equal to a preset duration will be aggregated with the first fault alarm information to obtain the aggregated fault alarm information.

[0018] Optionally, the method further includes:

[0019] Based on the aggregated fault alarm information, the root cause of the fault alarm is determined.

[0020] Optionally, the method further includes:

[0021] Based on the first fault alarm information and the network topology of the application system, the layer where the first object is located is determined;

[0022] If the level of the first object cannot be determined, an abnormal fault alarm message will be output.

[0023] Optionally, the method further includes:

[0024] Obtain the configuration information of the devices in the application system;

[0025] Based on the configuration information of the device, the objects in the application system are associated to obtain the network topology.

[0026] Secondly, this application provides a fault alarm processing device, the device comprising:

[0027] The acquisition module is used to acquire the first fault alarm information of the first object in the application system;

[0028] The first determining module is configured to determine, based on the first fault alarm information and the network topology of the application system, whether at least one network topology branch of the first object exists; each network topology branch includes: a second object located at a level above the level of the first object and having a network topology relationship with the first object;

[0029] The second determining module is used to determine, if present, whether the second fault alarm information of the second object in the network topology branch already exists;

[0030] An aggregation module is used to aggregate the first fault alarm information with the second fault alarm information of the second object in the network topology branch, if the fault alarm information already exists, to obtain aggregated fault alarm information.

[0031] Thirdly, this application provides an electronic device, the electronic device comprising: a processor, and a memory communicatively connected to the processor;

[0032] The memory stores computer-executed instructions;

[0033] The processor executes computer execution instructions stored in the memory to implement the method as described in any one of the first aspects.

[0034] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the task processing method as described in any one of the first aspects.

[0035] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the first aspects.

[0036] The fault alarm processing method, apparatus, electronic device, and storage medium provided in this application determine, based on first fault alarm information and the network topology of the application system, whether there exists a second object at a level higher than the first object and with a network topology relationship to the first object; if so, whether there is already second fault alarm information for the second object in the network topology branch; and after confirming its existence, aggregating the first fault alarm information with the second fault alarm information of the second object in the network topology branch to obtain aggregated fault alarm information. Through this method, the electronic device can determine the cause of the first object's fault based on the application system's network topology, and accordingly aggregate the first fault alarm information and the second fault alarm information of the second object in the network topology branch, thereby reducing the possibility of alarm storms, simplifying fault alarm information, and enabling maintenance personnel to quickly and accurately understand the cause of application system faults, reducing the workload of staff. Attached Figure Description

[0037] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0038] Figure 1 This is a schematic diagram of the architecture of an application system;

[0039] Figure 2 A flowchart illustrating the first fault alarm handling method provided in this application;

[0040] Figure 3 A schematic diagram of the network topology of an application system provided in this application;

[0041] Figure 4 A flowchart illustrating the second fault alarm handling method provided in this application;

[0042] Figure 5 A schematic diagram of the third type of fault alarm processing device provided in this application;

[0043] Figure 6 A schematic diagram of the structure of a fault alarm processing device provided in this application;

[0044] Figure 7 This is a schematic diagram of the structure of an electronic device 700 provided in this application.

[0045] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0046] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0047] It should be noted that the fault alarm processing method, device, electronic device and storage medium of this application can be used in the field of financial technology, or in any field other than financial technology. This application does not limit the application field of the fault alarm processing method, device, electronic device and storage medium.

[0048] The normal operation of existing application systems (such as banking information systems) typically requires the joint support of numerous hardware and software components. For example, application systems usually include network devices such as routers and switches; physical devices such as computers and servers; software deployed on the physical devices, such as operating systems; and application system software modules, such as databases and middleware. If virtual machines are deployed on the physical devices, these virtual machines will contain at least two operating systems.

[0049] However, currently, both the hardware and software components of application systems are prone to failure during operation. Therefore, to ensure the normal operation of application systems, it is common practice to monitor and issue alerts for operational failures in various hardware and software components during application system operation. When any hardware or software component in the application system malfunctions, a corresponding fault alarm is triggered, providing accurate and timely alerts on the application system's operational status and prompting maintenance personnel to repair application system faults as early as possible.

[0050] As mentioned above, existing application systems typically include numerous hardware and software components. These different hardware devices often have complex interrelationships. For example, Figure 1 This is a schematic diagram of the architecture of an application system, such as... Figure 1As shown, network devices typically connect to at least one physical device via network ports to support normal network communication for the physical devices of the application system. At least one piece of software is usually deployed on the physical device to provide the hardware foundation for its normal operation. Furthermore, software deployed on different physical devices may have dependencies on each other. For example, the normal operation of one piece of software may require calling another piece of software, such as by calling the Internet Protocol Address (IP address) of the other software. In this case, the normal operation of the first piece of software depends on the normal operation of the second piece of software.

[0051] Based on the above, when any software or hardware failure occurs in the application system, triggering a corresponding fault alarm, it is highly likely that connected and dependent software and hardware will also malfunction, triggering fault alarms for those connected software and hardware. However, this will lead to an alarm storm, where a large number of fault alarms are triggered in a short period, causing confusion in the alarm information. This makes it difficult for maintenance personnel to quickly and accurately determine the root cause of the failure. Manually investigating the root cause is labor-intensive, and it is difficult to determine the scope of the failure's impact, reducing fault handling efficiency and affecting the normal operation of the application system.

[0052] The inventors considered that when an application system failure occurs, if it is possible to automatically distinguish which software and hardware are truly malfunctioning and which are not, and to aggregate the fault alarm information in a targeted manner, it can both provide prompts for application system failures and prevent the occurrence of alarm storms.

[0053] In view of this, this application provides a fault alarm processing method, which can aggregate fault alarm information based on the fault alarm information and the correlation between various software and hardware within the application system, thereby avoiding alarm storms and preventing fault alarm information from becoming chaotic.

[0054] The subject of this application is an electronic device, such as a mobile phone, computer, or server. This electronic device can be a physical device within an application system or an electronic device independent of the application system; this application does not limit its scope.

[0055] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0056] Figure 2A flowchart illustrating the first fault alarm handling method provided in this application is shown below. Figure 2 As shown, the method includes:

[0057] S101. Obtain the first fault alarm information of the first object of the application system.

[0058] This application does not limit the architecture or business type of the aforementioned application system; for example, it could be a bank information system.

[0059] This application defines the specific types of the first object mentioned above and the second object mentioned below. For example, it may be a network device, a physical device, or software deployed on a physical device.

[0060] The first fault alarm information of the aforementioned first object may include, for example, the identifier of the first object. For instance, when the first object is a network device or a physical device, the identifier of the first object may be, for example, the serial number (SN) of the first object or a preset number of the first object; if the first object is software deployed on a physical device, the identifier of the first object may be, for example, a preset number of the first object; if the first object is an operating system deployed on a physical device, the identifier of the first object may be, for example, the IP address of the first object. This application does not limit the specific value of the preset number of the first object. This application does not limit whether the first fault alarm information includes other content besides the identifier of the first object.

[0061] One possible implementation is that the electronic device can monitor the operation of a first object in the application system. If the first object malfunctions, the electronic device acquires a first fault alarm message from the first object in the application system. The electronic device can acquire the first fault alarm message sent by the first object, or it can generate the first fault alarm message based on monitoring the first object; this application does not limit this. This application does not limit the specific implementation method of the electronic device monitoring the operation of the first object in the application system; specific implementations can be found in existing technologies.

[0062] Another possible implementation is that the electronic device can obtain first fault alarm information sent by other electronic devices (such as servers, computers, etc.). For example, other electronic devices can monitor the fault alarm status of the application system, and after obtaining the first fault alarm information sent by the first object, the other electronic devices can send the first fault alarm information to the electronic device.

[0063] S102. Based on the first fault alarm information and the network topology of the application system, determine whether there is at least one network topology branch of the first object.

[0064] This application does not limit the layers included in the network topology of the aforementioned application system, nor the relationships between objects at each layer. For example, Figure 3 A network topology diagram of an application system provided in this application is shown below. Figure 3 As shown, the network topology of this application system may include, for example, a network layer, a device layer, and a software layer. The network layer includes network devices, the device layer includes physical devices, and the software layer includes software deployed on the physical devices. Any network device in the network layer is connected to at least one physical device in the device layer; at least one type of software is deployed on any physical device in the device layer. Optionally, objects within a layer may also have associated relationships. For example, continue referring to... Figure 3 Software 7 needs to call software 4 to function properly, so there is a relationship between software 7 and software 4.

[0065] Each of the above network topology branches includes: a second object located at a level above the level of the first object and having a network topology relationship with the first object.

[0066] The existence of network topology relationships mentioned here refers to the existence of associations. These associations can be direct or indirect. For example, continue to refer to... Figure 3 Network device A in the network layer has a network topology relationship with physical device A and physical device B in the device layer; physical device A has a network topology relationship with software 1 and software 2; network device A has a network topology relationship with software 1 and 2 through physical device A, and with software 3, 4 and 5 through physical device B.

[0067] It should be understood that there are hierarchical relationships within the network topology of an application system. The normal operation of a lower level depends on the upper level. For example, continue to refer to... Figure 3 The application system's network topology, from top to bottom, consists of the network layer, the device layer, and the software layer. The normal operation of physical devices in the device layer depends on the normal operation of network devices in the network layer; similarly, the normal operation of software in the software layer depends on the normal operation of the physical devices deployed in the device layer.

[0068] In this step, the electronic device can identify the first object that has malfunctioned based on the first fault alarm information. Furthermore, the electronic device can determine whether there is at least one network topology branch of the first object based on the network topology of the application system and the known first object that has malfunctioned.

[0069] If it exists, it indicates that the first object triggering the first fault alarm information may be due to a fault in the second object in a layer above the first object, then proceed to step S103.

[0070] S103. Determine whether a second fault alarm message already exists for the second object in the network topology branch.

[0071] The content included in the second fault alarm information mentioned above can be referred to the first fault alarm information, and will not be repeated here.

[0072] For example, the electronic device can record all fault alarm information; that is, if the second object fails, the electronic device will acquire and record the second fault alarm information of the second object. In this step, the electronic device determines whether the second fault alarm information of the second object already exists in the network topology branch, and then determines whether the first object triggered the first fault alarm information due to the failure of the second object.

[0073] If it already exists, it indicates that the first object triggered the first fault alarm information because the second object was faulty, that is, the first object itself did not actually fail, so step S104 is executed.

[0074] S104. Aggregate the first fault alarm information with the second fault alarm information of the second object in the network topology branch to obtain the aggregated fault alarm information.

[0075] The aggregation of the first fault alarm information with the second fault alarm information of the second object in the network topology branch to obtain the aggregated fault alarm information can be achieved by, for example, by having the electronic device retain only the second fault alarm information and not the first fault alarm information in the aggregated fault alarm information.

[0076] Alternatively, the electronic device adds identifiers to both the first and second fault alarm messages, and then obtains aggregated fault alarm messages based on the first and second fault alarm messages with added identifiers. The identifiers mentioned here can be, for example, identifiers used to characterize whether a fault is the root cause. For instance, if a second object in a network topology branch is located at the highest level in the network topology branch, an identifier indicating that the second object is the root cause is added to the second fault alarm message corresponding to that second object; the remaining second objects and the first object are identified as not being the root cause. This application does not limit the specific representation of the aforementioned identifiers.

[0077] The method by which the electronic device aggregates the first fault alarm information with the second fault alarm information of the second object in the network topology branch is related to the number of levels existing in the network topology branch. For example, if there is a second object at one level in the network topology branch, the electronic device aggregates the first fault alarm information into the second fault alarm information of the second device to obtain aggregated fault alarm information. If there are second objects at multiple levels in the network topology branch, the electronic device can, for example, aggregate the first fault alarm information into the fault alarm information of the highest-level second object to obtain aggregated fault alarm information.

[0078] In this embodiment, the electronic device determines, based on the first fault alarm information and the network topology of the application system, whether there exists a second object at a level higher than the first object and with a network topology relationship to the first object. If such an object exists, it determines whether a second fault alarm information for the second object already exists in the network topology branch. Upon confirmation of existence, the first fault alarm information is aggregated with the second fault alarm information of the second object in the network topology branch to obtain aggregated fault alarm information. Through this method, the electronic device can determine the cause of the first object's fault based on the application system's network topology, and accordingly aggregate the first fault alarm information and the second fault alarm information of the second object in the network topology branch. This reduces the possibility of alarm storms, simplifies fault alarm information, and allows maintenance personnel to quickly and accurately identify the cause of application system faults, reducing the workload of staff.

[0079] Optionally, in the above embodiments, after determining whether at least one network topology branch of the first object exists, if the electronic device does not exist, it indicates that the first electronic device triggering the first fault alarm information is unlikely to be caused by a fault in a second object at a level higher than the first object. In this case, one possible implementation is that the electronic device can directly generate aggregated fault alarm information based solely on the first fault alarm information. For example, the aggregated fault alarm information may only include the first fault alarm information.

[0080] Another possible implementation is that the electronic device can determine whether a historical first fault alarm message for the first object already exists. If it does, it indicates that a first fault alarm message has already been recorded for that first object, and the subsequent processing method of the electronic device depends on the actual application scenario.

[0081] For example, if the electronic device only records the first fault alarm information of the first object whose fault has not been eliminated in the past, then the electronic device can directly aggregate the historical first fault alarm information with the first fault alarm information to obtain the aggregated fault alarm information.

[0082] If, regardless of whether the fault has been eliminated, the electronic device records all historically acquired first fault alarm information for the first object, indicating that the electronic device currently does not determine whether the first fault alarm information for the first object was triggered by this fault, then subsequently, the electronic device can aggregate historical first fault alarm information with the first fault alarm information whose generation time is less than or equal to a preset duration, to obtain aggregated fault alarm information. This aggregation can, for example, simply involve retaining any one of the first fault alarm information within the aggregated fault alarm information. This application does not limit the specific value of the aforementioned preset duration; those skilled in the art can determine it according to the actual situation.

[0083] For example, the fault alarm information may also include the generation time of the fault alarm information. The electronic device can determine whether it is necessary to aggregate the first fault alarm information and the historical first fault alarm information to obtain the aggregated fault alarm information based on the generation time.

[0084] Since a fault typically lasts only a short time, this method of determining fault alarm information based on its generation time allows electronic devices to distinguish between first fault alarms of a first object that have already been cleared in the past and first fault alarms of a first object that have not yet been cleared. This reduces the possibility of erroneous aggregation of fault alarm information and improves aggregation accuracy.

[0085] The following describes how an electronic device aggregates the first fault alarm information with the second fault alarm information of the second object in the network topology branch when there are at least two levels of second objects in the network topology branch, i.e., step S104 in the above embodiment. Figure 4 A flowchart illustrating the second fault alarm handling method provided in this application is shown below. Figure 4 As shown, step S104 may include:

[0086] S201. Determine the third device located at the highest level from the second objects at at least two levels.

[0087] In this step, if there are at least two levels of second objects in the network topology branch, the electronic device first determines the third device at the highest level from the second objects at at least two levels, in order to aggregate alarm information. For example, the electronic device can determine the third device at the highest level from the second objects at at least two levels based on the network topology of the application system.

[0088] S202. The existing fault alarm information of the remaining second objects and the first fault alarm information are aggregated into the third fault alarm information of the third device to obtain the aggregated fault alarm information.

[0089] Since the third object, the second object, and the first object are all topologically connected, and the third object is located at the highest level in the network topology branch in this step, the first object triggering the first fault alarm information, and the second object triggering the second fault alarm information, are both caused by a fault in the third object. Therefore, in this step, the electronic device aggregates the existing fault alarm information of the remaining second objects and the first fault alarm information into the third fault alarm information of the third device to obtain the aggregated fault alarm information. For example, the aggregated fault alarm information may only include the third fault alarm information, excluding the first and second fault alarm information.

[0090] The method by which the electronic device aggregates the existing fault alarm information of the remaining second objects, as well as the first fault alarm information, into the third fault alarm information of the third device to obtain the aggregated fault alarm information is related to the actual application scenario. For example, if the electronic device only retains the fault alarm information of the objects whose faults have not been eliminated, it indicates that the fault of the third object indicated by the currently recorded third fault alarm information has not been eliminated, that is, the reason why the fault alarm information of the first object and the remaining second objects is triggered is due to the fault of the third object. In this case, the electronic device can aggregate the existing fault alarm information of the remaining second objects, as well as the first fault alarm information, into the third fault alarm information without distinction to obtain the aggregated fault alarm information.

[0091] For example, continue to refer to Figure 3 Taking software 1 as the first object as an example, the network topology of the application system may contain two layers of second objects: a device layer and a network layer, namely physical device A and network device A, respectively. Thus, the electronic device can determine network device A as the third device at the highest level from these two layers of second objects based on the network topology of the application system. Subsequently, the electronic device aggregates the first fault alarm information of software 1 and the second fault alarm information of physical device A into the third fault alarm information of network device A, obtaining aggregated fault alarm information. This aggregated fault alarm information only includes the third fault alarm information of network device A and does not include the aforementioned first and second fault alarm information.

[0092] If, regardless of whether the fault has been eliminated, the electronic device records fault alarm information for all historically acquired objects, indicating that the electronic device currently cannot determine whether the third fault alarm information for the third object was triggered by this fault, then subsequently, the electronic device can aggregate the existing fault alarm information for the remaining second objects, as well as the fault alarm information in the first fault alarm information whose generation time difference with the third fault alarm information is less than or equal to a preset duration, into the third fault alarm information to obtain the aggregated fault alarm information. Its technical effect is similar to that of determining whether the first fault alarm information was triggered by the current fault based on the fault alarm information generation time in the above embodiment, and will not be elaborated further here.

[0093] This implementation allows electronic devices to distinguish between current fault alarms and previous fault alarms in historical records, preventing the aggregation of first and second fault alarm information into a third fault alarm in historical records that was not caused by the current fault. For example, in some scenarios, electronic devices may record and retain all historical fault alarm information. The above method can reduce the possibility of incorrect aggregation of fault alarm information in such scenarios and improve aggregation accuracy.

[0094] Optionally, after acquiring the aggregated fault alarm information, the electronic device can also determine the root cause of the fault alarm based on the aggregated fault alarm information.

[0095] The method by which electronic devices determine the root cause of fault alarms based on aggregated fault alarm information is related to the content included in the fault alarm information. For example, if the aggregated fault alarm information only includes fault alarm information of the highest-level object in the network topology branch, then the electronic device determines that the highest-level object is the root cause of the fault alarm.

[0096] If the aggregated fault alarm information includes a first object, as well as objects in all network topology branches of the first object, and the highest-level object in each network topology branch is marked, then the electronic device can determine that the highest-level object is the root cause of the fault alarm based on the marking.

[0097] Optionally, the electronic device can also determine the level of the first object based on the first fault alarm information and the network topology of the application system; if the level of the first object cannot be determined, it can output a prompt message indicating that the fault alarm information is abnormal.

[0098] This application does not limit the method by which an electronic device outputs fault alarm information or abnormal prompts. For example, an electronic device may output fault alarm information or abnormal prompts via voice output and / or text output.

[0099] If the electronic device cannot determine the level of the first object, it indicates an anomaly in the first fault alarm message or the network topology of the application system. For example, the network topology of the application system may have missed the first object corresponding to the first fault alarm message, or the identifier of the first object included in the first fault alarm message may not match the identifier of the first object in the network topology. Therefore, in this way, the electronic device can promptly alert to anomalies in fault alarm messages and network topology, enabling maintenance personnel to be aware of the anomalies in a timely manner and take appropriate measures.

[0100] Regarding the acquisition of the aforementioned network topology, one possible implementation is that the electronic device can acquire the network topology input by the user. For example, the electronic device has a user interface, and the electronic device acquires the network topology input by the user through this user interface.

[0101] Another possible implementation is that the electronic device can obtain the configuration information of the device in the application system; then, based on the device configuration information, it can associate the objects in the application system to obtain the above network topology.

[0102] The devices in the aforementioned application system can be, for example, physical devices and / or network devices within the application system. The device configuration information characterizes the network topology of the object, and its specific content is related to the device type. For example, when the device is a physical device, the configuration information may include, for example, the identifier of the physical device; and / or, the identifier of the software deployed on the physical device, such as the identifier of the operating system deployed on the device, and the identifier of the software modules of the application system deployed under that operating system; and / or, the identifier of the network device connected to the physical device; and / or, the identifier of the software called by the software deployed on the physical device. When the device is a network device, the configuration information may, for example, include the identifier of the physical device connected to the network device.

[0103] For example, an electronic device can obtain the configuration information of all physical devices in the application system, and then associate objects in the application system based on the configuration information of all physical devices to obtain the network topology of the application system.

[0104] Alternatively, electronic devices can obtain the configuration information of all physical devices in the application system, as well as the configuration information of all network devices, and then associate the objects in the application system accordingly to obtain the aforementioned network topology.

[0105] Using the above method, electronic devices can automatically generate the network topology based on the configuration information of devices in the application system, eliminating the need for manual input and improving the efficiency of network topology acquisition. Furthermore, the method can generate the network topology based on the actual configuration information of the devices included in the application system, thus improving the accuracy of the generated network topology.

[0106] Optionally, if the first object is software in the software layer of the application system, the electronic device may perform the following steps after obtaining the first fault alarm information of the first object in the application system:

[0107] S301. Determine whether the first object has an associated node in the software layer.

[0108] The associated node mentioned here can be any software in the software layer other than the first object. The normal operation of the first object depends on the normal operation of the associated node.

[0109] If it exists, it indicates that the operation of the first object will be affected by the associated node, then proceed to step S302.

[0110] If it does not exist, it indicates that the operation of the first object will not be affected by the associated node, and then step S102 in the above embodiment is executed.

[0111] S302. Determine whether a fourth fault alarm message already exists for the associated node.

[0112] If it exists, it indicates that the first object triggering the first fault alarm information may be caused by the fault of its associated node. Therefore, it is necessary to determine the root cause of the fault based on its associated node. Then, the associated node is determined as the first object, and step S102 in the above embodiment is executed.

[0113] Through the above implementation, when the first object is software in the software layer of the application system, the electronic device can further aggregate fault alarm information in the application system fault monitoring and alarm process by combining its association with other software, thereby expanding the application scenarios of the fault alarm processing method provided in this application.

[0114] The application system adopts Figure 3 Using the network topology shown as an example, the fault alarm handling method provided in this application will be explained. Figure 5 A flowchart illustrating the third fault alarm handling method provided in this application is shown below. Figure 5 As shown, the fault alarm handling method may include the following steps:

[0115] S401. Obtain the first fault alarm information of the first object of the application system.

[0116] S402. Based on the first fault alarm information and the network topology of the application system, determine the level of the first object to determine whether there is at least one network topology branch of the first object.

[0117] If the first object is located at the network layer, then the first object is a network device. If the first object is located at the highest level of the network topology, then the first object does not have a network topology branch, and step S403 can be executed.

[0118] If the first object is located at the device layer, then the first object is a physical device. If the first object is located at the second layer of the network topology, then the first object has at least one network topology branch, and step S406 can be executed.

[0119] If the first object is located at a software layer, then the first object is software. If the first object is located at the third layer of the network topology, then the first object has at least one network topology branch, and step S408 can be executed.

[0120] If the level of the first object cannot be determined, proceed to step S415.

[0121] S403. Determine whether the first historical fault alarm information of the first object already exists.

[0122] If it exists, proceed to step S404.

[0123] If it does not exist, proceed to step S405.

[0124] S404. Aggregate the historical first fault alarm information with the first fault alarm information whose time difference with the generation time of the first fault alarm information is less than or equal to a preset duration with the first fault alarm information to obtain the aggregated fault alarm information.

[0125] S405, End the processing of this fault alarm information.

[0126] S406. Determine whether a second fault alarm message already exists for the second object in the network topology branch;

[0127] If it already exists, proceed to step S407.

[0128] If it does not exist, proceed to step S403.

[0129] S407. Aggregate the first fault alarm information with the second fault alarm information of the second object in the network topology branch to obtain the aggregated fault alarm information.

[0130] S408. Determine whether the first object has an associated node.

[0131] If it exists, proceed to step S409.

[0132] If it does not exist, proceed to step S411.

[0133] S409. Determine whether a fourth fault alarm message already exists for the associated node.

[0134] If it exists, then the associated node is taken as the first object, and step S410 is executed.

[0135] If it does not exist, proceed to step S411.

[0136] S410. Based on the fourth fault alarm information and the network topology of the application system, obtain the network topology branch of the first object.

[0137] S411. Based on the network topology branch of the first object, determine whether there is already a third fault alarm message for a third object located in the network layer within the network topology branch.

[0138] If it exists, proceed to step S412.

[0139] If it does not exist, proceed to step S413.

[0140] S412. The existing fault alarm information of the remaining second objects and the first fault alarm information are aggregated into the third fault alarm information of the third object to obtain the aggregated fault alarm information.

[0141] S413. Determine whether there is a second fault alarm message for the remaining second object located at the device layer in the network topology branch.

[0142] If it exists, proceed to step S414.

[0143] If it does not exist, proceed to step S403.

[0144] S414. The first fault alarm information is aggregated into the second fault alarm information of the second object to obtain the aggregated fault alarm information.

[0145] S415, Output fault alarm information and abnormal prompts.

[0146] This embodiment uses an application system as an example. Figure 3Taking the network topology shown as an example, in this embodiment, the electronic device first determines the layer of the first object based on the first fault alarm information of the first object and the network topology, thereby determining its network topology branch. Then, for the first object located at different layers, the electronic device uses corresponding methods to determine whether the fault alarm information of the first object needs to be aggregated, and performs aggregation processing after determining that aggregation is necessary. Through the above method, the electronic device can aggregate fault alarm information when a fault alarm occurs in the application system, reducing the possibility of alarm storms and thus avoiding the occurrence of chaotic fault alarm information. This method can reduce labor costs and improve the efficiency of determining the root cause of faults.

[0147] Figure 6 This is a schematic diagram of the structure of a fault alarm processing device provided in this application, as shown below. Figure 6 As shown, the task processing device includes: an acquisition module 11, a first determination module 12, a second determination module 13, and an aggregation module 14. Optionally, the device may further include the following module: a third determination module 15.

[0148] Module 11 is used to acquire the first fault alarm information of the first object of the application system;

[0149] The first determining module 12 is used to determine, based on the first fault alarm information and the network topology of the application system, whether there is at least one network topology branch of the first object; each network topology branch includes: a second object located at a level above the level of the first object and having a network topology relationship with the first object;

[0150] The second determining module 13 is used to determine, if present, whether the second fault alarm information of the second object in the network topology branch already exists;

[0151] The aggregation module 14 is used to aggregate the first fault alarm information with the second fault alarm information of the second object in the network topology branch if the fault alarm information already exists, so as to obtain the aggregated fault alarm information.

[0152] Optionally, the aggregation module 14 is specifically used to determine the third device located at the highest level from the second objects at the at least two levels if there are second objects at the at least two levels in the network topology branch; and to aggregate the fault alarm information of the remaining second objects and the first fault alarm information into the third fault alarm information of the third device to obtain aggregated fault alarm information.

[0153] For example, the aggregation module 14 is specifically used to aggregate the existing fault alarm information of the other second objects, as well as the fault alarm information in the first fault alarm information whose generation time difference with the third fault alarm information is less than or equal to a preset duration, into the third fault alarm information to obtain the aggregated fault alarm information.

[0154] Optionally, after the first determining module 12 determines whether at least one network topology branch of the first object exists, the second determining module 13 is used to determine whether historical first fault alarm information of the first object already exists if it does not exist; if it does exist, the historical first fault alarm information with a time difference less than or equal to the generation time of the first fault alarm information is aggregated with the first fault alarm information to obtain aggregated fault alarm information.

[0155] Optionally, the third determining module 15 is used to determine the root cause of the fault alarm based on the aggregated fault alarm information.

[0156] Optionally, the first determining module 12 is further configured to determine the level of the first object based on the first fault alarm information and the network topology of the application system; if the level of the first object cannot be determined, an abnormal fault alarm information prompt message is output.

[0157] Optionally, the first acquisition module 11 is further configured to acquire configuration information of devices in the application system; and associate objects in the application system according to the configuration information of the devices to obtain the network topology.

[0158] The task processing device provided in this application embodiment can execute the fault alarm processing method in the above method embodiment. Its implementation principle and technical effects are similar, and will not be repeated here. It should be noted that the above... Figure 6 The division of modules shown is merely illustrative. This application does not limit the division of modules or the naming of modules.

[0159] Figure 7 This is a schematic diagram of the structure of an electronic device 700 provided in this application. Figure 7 As shown, the electronic device 700 may include at least one processor 701 and a memory 702.

[0160] The memory 702 is used to store programs. Specifically, the program may include program code, which includes computer operation instructions.

[0161] The memory 702 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0162] The processor 701 is used to execute computer execution instructions stored in the memory 702 to implement the fault alarm processing method described in the foregoing method embodiments. The processor 701 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0163] The electronic device 700 may also include a communication interface 703, through which it can communicate and interact with external devices, such as user terminals (e.g., computers, tablets). In specific implementations, if the communication interface 703, memory 702, and processor 701 are implemented independently, they can be interconnected via a bus to complete communication. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc., but this does not imply that there is only one bus or one type of bus.

[0164] Optionally, in a specific implementation, if the communication interface 703, memory 702, and processor 701 are integrated on a single chip, then the communication interface 703, memory 702, and processor 701 can communicate through an internal interface.

[0165] This application also provides a computer-readable storage medium, which may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. Specifically, the computer-readable storage medium stores program instructions, which are used for the fault alarm processing method in the above embodiments.

[0166] This application also provides a program product including executable instructions stored in a readable storage medium. At least one processor of an electronic device can read the executable instructions from the readable storage medium, and the at least one processor executes the executable instructions to cause the electronic device to implement the fault alarm handling methods provided in the various embodiments described above.

[0167] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0168] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A fault alarm handling method, characterized in that, The method includes: Obtain the first fault alarm information of the first object in the application system; Based on the first fault alarm information and the network topology of the application system, the layer where the first object is located is determined to determine whether there is at least one network topology branch of the first object; the network topology includes multiple layers, including a software layer, and the software layer includes objects with related relationships; each network topology branch includes: a second object located at a layer above the layer where the first object is located and having a network topology relationship with the first object; If the first object is located at a software layer, and the first object has at least one network topology branch, determine whether the first object has an associated node in the software layer. If the first object has associated nodes, then determine whether the fourth fault alarm information of the associated nodes already exists; If the fourth fault alarm information of the associated node exists, the associated node is taken as a new first object, and it is determined in the network topology branch of the first object whether the second fault alarm information of the second object already exists. If it already exists, the first fault alarm information is aggregated with the second fault alarm information of the second object in the network topology branch to obtain the aggregated fault alarm information.

2. The method according to claim 1, characterized in that, The aggregation of the first fault alarm information with the second fault alarm information of the second object in the network topology branch includes: If there are second objects at at least two levels in the network topology branch, then determine the third object at the highest level from the second objects at at least two levels; The existing fault alarm information of the remaining second objects, as well as the first fault alarm information, are aggregated into the third fault alarm information of the third object to obtain the aggregated fault alarm information.

3. The method according to claim 2, characterized in that, The step of aggregating the existing fault alarm information of the remaining second objects, and the first fault alarm information, into the fault alarm information of the third object to obtain the aggregated fault alarm information includes: The existing fault alarm information of the remaining second objects, as well as the fault alarm information in the first fault alarm information whose generation time difference with the third fault alarm information is less than or equal to a preset duration, are aggregated into the third fault alarm information to obtain the aggregated fault alarm information.

4. The method according to claim 3, characterized in that, After determining whether at least one network topology branch of the first object exists, the method further includes: If at least one network topology branch of the first object does not exist, determine whether the first object has a historical first fault alarm message. If at least one network topology branch of the first object exists, then the historical first fault alarm information whose generation time difference with the first fault alarm information is less than or equal to a preset duration is aggregated with the first fault alarm information to obtain aggregated fault alarm information.

5. The method according to any one of claims 1-4, characterized in that, The method further includes: Based on the aggregated fault alarm information, the root cause of the fault alarm is determined.

6. The method according to any one of claims 1-4, characterized in that, The method further includes: If the level of the first object cannot be determined, an abnormal fault alarm message will be output.

7. The method according to any one of claims 1-4, characterized in that, The method further includes: Obtain the configuration information of the devices in the application system; Based on the configuration information of the device, the objects in the application system are associated to obtain the network topology.

8. A fault alarm processing device, characterized in that, The device includes: The acquisition module is used to acquire the first fault alarm information of the first object in the application system; The first determining module is configured to determine the level of the first object based on the first fault alarm information and the network topology of the application system, so as to determine whether there is at least one network topology branch of the first object; the network topology includes multiple levels, the multiple levels include a software layer, and the software layer includes objects with related relationships; each network topology branch includes: a second object located at a level above the level of the first object and having a network topology relationship with the first object; The second determining module is used to determine whether the first object has an associated node in the software layer if the first object is located at a software layer and the first object has at least one network topology branch. If the first object has associated nodes, then determine whether the fourth fault alarm information of the associated nodes already exists; If the fourth fault alarm information of the associated node exists, the associated node is taken as a new first object, and it is determined in the network topology branch of the first object whether the second fault alarm information of the second object already exists. An aggregation module is used to aggregate the first fault alarm information with the second fault alarm information of the second object in the network topology branch, if the fault alarm information already exists, to obtain aggregated fault alarm information.

9. An electronic device, characterized in that, The electronic device includes: a processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the fault alarm processing method as described in any one of claims 1 to 7.

11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, is used to implement the fault alarm processing method as described in any one of claims 1 to 7.