Alarm record analysis method and device, equipment and storage medium

By obtaining alarm records and determining the target business system and host identifiers, and obtaining the alarm root event and handling status, the problems of long alarm processing time and difficulty in merging multiple alarms in the existing technology are solved, and the effect of rapid fault location and reduced operation and maintenance costs are achieved.

CN119988081APending Publication Date: 2025-05-13BEIJING YOUTEJIE INFORMATION TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510148418.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, it is difficult for users to quickly lock the business system, source data, root logs, and host performance when processing alarms, resulting in too long troubleshooting time, and the lack of effective merge query capabilities when multiple alarms are triggered at the same time, which increases operation and maintenance costs.

Method used

By obtaining alarm records, the target service system name and target host identification are determined, the alarm root event is obtained based on this information, and the handling status of the alarm record is determined based on the root event, so as to realize the automatic analysis and processing of alarm records.

Benefits of technology

It significantly reduces the fault processing time, improves the correlation analysis ability of alarms, enables users to quickly and accurately identify the root causes of alarms, reduces operation and maintenance costs, and improves operation and maintenance capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988081A_ABST
    Figure CN119988081A_ABST
Patent Text Reader

Abstract

The invention discloses an alarm record analysis method and device, equipment and a storage medium, and the method comprises the steps: obtaining an alarm record, and determining a target business system name corresponding to the alarm record, the alarm record comprising a target host identifier; obtaining an alarm root event corresponding to the alarm record based on the target service system name and the target host identifier; and determining a disposal state corresponding to the alarm record according to the alarm root event. According to the method, the host performance data and the context log data are retrieved, the alarm root event is determined and the states of the alarm event and the related alarm event are automatically modified by enriching assets of the alarm records, tracing, merging and processing the alarm event and automatically associating, so that the fault positioning efficiency is improved, the problem of repeated troubleshooting of a plurality of alarm records is solved, and the fault positioning efficiency is improved. Through an automatic processing function, the processing states of the alarm root event and the related event are automatically modified, reasonable arrangement of operation and maintenance resources is facilitated, the operation and maintenance capability is improved, and stable operation of a service system is better ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of alarm tracing analysis, and in particular to an alarm record analysis method, device, equipment and storage medium. Background Art

[0002] With the rapid development of informatization, modern enterprises have an increasing demand for monitoring equipment, operating systems, and business systems. However, the resulting surge in alarm data often makes it difficult for users to locate and query alarm notifications after receiving them. Specifically, when handling alarms, it is difficult for users to quickly lock in the business system, source data, root cause logs, and host performance to which the alarm belongs, which prolongs the troubleshooting time. When multiple alarms are triggered at the same time, there is a lack of effective merge query capabilities, which further increases the operation and maintenance costs.

[0003] In the prior art, although some monitoring platforms provide the function of viewing the source data of alarms, the source tracing analysis of most alarms still needs to rely on manual elimination one by one. This method is not only cumbersome and time-consuming, but also leads to low troubleshooting efficiency and cannot meet the needs of rapid response. Summary of the invention

[0004] The present invention provides an alarm record analysis method, device, equipment and storage medium to solve the problem of long manual analysis and troubleshooting time, which can improve the alarm correlation analysis capability, enable users to quickly and accurately identify the root cause of the alarm, thereby significantly reducing fault handling time.

[0005] According to one aspect of the present invention, there is provided an alarm record analysis method, the method comprising:

[0006] Obtain an alarm record and determine the name of the target business system corresponding to the alarm record, wherein the alarm record includes a target host identifier;

[0007] Acquire the alarm root event corresponding to the alarm record based on the target business system name and the target host identifier;

[0008] Determine the handling status corresponding to the alarm record based on the alarm root event.

[0009] Optionally, obtaining an alarm record includes: obtaining an alarm monitoring configuration item, wherein the alarm monitoring configuration item includes a target host identifier and a monitoring item; determining a target host corresponding to the alarm monitoring configuration item based on the target host identifier; and monitoring the target host based on the monitoring item to obtain the alarm record.

[0010] Optionally, determining the target business system name corresponding to the alarm record includes: obtaining an asset database, wherein the asset database includes the business system name corresponding to each host identifier; matching the target host identifier through the asset database to determine the target business system name corresponding to the alarm record.

[0011] Optionally, an alarm root event corresponding to the alarm record is obtained based on the target business system name and the target host identifier, including: determining the alarm time of the alarm record, and determining the retrieval time range according to the alarm time; obtaining an alarm database, and obtaining relevant alarm data with the same name as the target business system and the target host identifier in the alarm database based on the retrieval time range; judging whether there is an uninterrupted alarm before the alarm time in the relevant alarm data, and if so, determining that the alarm record is a relevant alarm, and determining the alarm root event based on the preset uninterrupted time; otherwise, determining that the alarm record is an alarm root event.

[0012] Optionally, determining the disposal status corresponding to the alarm record based on the alarm root event includes: obtaining a log database, locating the source data and host performance data corresponding to the alarm root event in the log database; forming a root cause based on the source data or the host performance data; determining whether to enable automatic disposal, and if so, determining the disposal status of the alarm record and related alarm data as a root cause completion status based on the root cause; otherwise, determining the disposal status of the alarm record as a root cause completion status based on the root cause.

[0013] Optionally, forming a root cause based on source data or host performance data includes: locating context log data in a log database based on the source data; when the context log data triggers a preset abnormal keyword or the host performance data triggers a preset threshold, obtaining corresponding abnormal data and using the abnormal data as the root cause.

[0014] Optionally, the method further includes: when the context log data does not trigger a preset abnormal keyword or the host performance data does not trigger a preset threshold, determining that the handling status of the alarm record requires manual analysis.

[0015] According to another aspect of the present invention, there is provided an alarm record analysis device, the device comprising:

[0016] An alarm record acquisition module is used to acquire the alarm record and determine the name of the target business system corresponding to the alarm record, wherein the alarm record includes a target host identifier;

[0017] An alarm root event determination module is used to obtain an alarm root event corresponding to the alarm record based on the target business system name and the target host identifier;

[0018] The alarm record analysis module is used to determine the disposal status corresponding to the alarm record according to the alarm root event.

[0019] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0020] at least one processor;

[0021] and a memory communicatively coupled to the at least one processor;

[0022] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute an alarm record analysis method described in any embodiment of the present invention.

[0023] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement an alarm record analysis method described in any embodiment of the present invention when executed.

[0024] The technical solution of the embodiment of the present invention, through asset enrichment of alarm records, alarm event tracing and merging processing, and automatic association, retrieves host performance data and context log data, determines the alarm root event, and then automatically modifies the status of the alarm event and related alarm events, thereby improving the efficiency of fault location, solving the problem of repeated troubleshooting of multiple alarm records, and automatically modifies the disposal status of the alarm root event and related events through the automated processing function, which helps to reasonably arrange operation and maintenance resources, improve operation and maintenance capabilities, and better ensure the stable operation of the business system.

[0025] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0027] Figure 1 is a flow chart of an alarm record analysis method provided according to Embodiment 1 of the present invention;

[0028] Figure 2 is a flowchart of another alarm record analysis method provided according to Embodiment 2 of the present invention;

[0029] Figure 3is a structural diagram of an alarm record analysis device provided according to Embodiment 3 of the present invention;

[0030] Figure 4 It is a structural schematic diagram of an electronic device for implementing an alarm record analysis method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0031] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0033] Embodiment 1

[0034] Figure 1 A flowchart of an alarm record analysis method is provided for the first embodiment of the present invention. This embodiment is applicable to the case of alarm source tracing analysis. The method can be executed by an alarm record analysis device. The alarm record analysis device can be implemented in the form of hardware and / or software. The alarm record analysis device can be configured in a computer controller. Figure 1 As shown, the method includes:

[0035] S110: Acquire an alarm record, and determine the name of a target business system corresponding to the alarm record, wherein the alarm record includes a target host identifier.

[0036] Among them, the alarm record refers to the information record generated by the system when an abnormal situation is detected, including the target host ID, alarm time, alarm content, etc., which is used to remind operation and maintenance personnel to pay attention to potential problems. The business system is a collection of software systems built by the enterprise to achieve specific business functions. The name of the target business system is the name of the specific business system involved in the alarm, such as the financial system, e-commerce transaction system, etc., which is used to locate the business field to which the alarm belongs. The host is a physical or virtual computer device that runs the business system. The target host ID refers to the identifier that determines the alarm device, which can be an IP address, host name, etc., which can accurately locate the device that generates the alarm.

[0037] Specifically, in a modern complex information technology environment, a large number of alarm records are generated when various business systems are running. Alarm records can come from the server's operating system log, the monitoring module inside the application, the status feedback of network equipment, etc. For example, the server operating system will generate corresponding alarm information when the hardware fails, and the application will also record an alarm when it encounters a key business logic error. Obtaining alarm records can rely on special log collection tools or monitoring systems, which will collect alarm information from various data sources in a scheduled or real-time manner according to preset rules, and summarize it to a centralized storage location for subsequent processing. Through the pre-established mapping relationship, that is, which hosts belong to which business system, the name of the target business system corresponding to the alarm record can be determined. For example, in an enterprise's IT architecture, the IP address segment 192.168.1.0 / 24 is assigned to the server of the e-commerce business system. When an alarm record is obtained, its target host identifier is 192.168.1.10, it can be determined that the target business system corresponding to the alarm record is the e-commerce business system.

[0038] Optionally, obtaining an alarm record includes: obtaining an alarm monitoring configuration item, wherein the alarm monitoring configuration item includes a target host identifier and a monitoring item; determining a target host corresponding to the alarm monitoring configuration item based on the target host identifier; and monitoring the target host based on the monitoring item to obtain the alarm record.

[0039] Among them, the target host identifier is used to accurately locate the specific host that needs to be monitored. Through the target host identifier, the controller can filter out specific hosts from a large number of devices for monitoring. The monitoring items specify the specific indicators or events that need to be paid attention to for the target host. For example, for a Web server, monitoring items may include CPU usage, memory usage, network traffic, the operating status of specific services, etc. Alarm monitoring configuration items are stored in a dedicated configuration management system or database. Operation and maintenance personnel can pre-set the monitoring configuration of each host in the system according to business needs and host characteristics. When it is necessary to obtain an alarm record, the controller will read the alarm monitoring configuration item information from the centralized storage location.

[0040] Specifically, for different monitoring items, the controller can use corresponding technical means to perform data collection and status monitoring. For example, for performance indicator monitoring items such as CPU usage and memory usage, the system call interface provided by the operating system or a special monitoring agent can be used to collect data regularly. For specific service operation status monitoring items, the controller can determine whether the service is running normally by interacting with the service process. When the data collected or the monitored status of the monitoring item exceeds the normal range, the controller will automatically generate an alarm record. The alarm record can include the time when the alarm occurred, the target host identifier, the name of the monitoring item that triggered the alarm, the alarm description, etc.

[0041] Optionally, determining the target business system name corresponding to the alarm record includes: obtaining an asset database, wherein the asset database includes the business system name corresponding to each host identifier; matching the target host identifier through the asset database to determine the target business system name corresponding to the alarm record.

[0042] Specifically, when it is necessary to determine the target business system name corresponding to the alarm record, the controller will access the asset database, and then use the target host identifier as the query condition to search in the asset database. For each record in the asset database, the controller will compare its host identifier field with the target host identifier. If a completely matching record is found, the business system name in the record will be extracted and returned to the alarm analysis program. If no matching record is found, it may mean that the asset database information is incomplete or the target host identifier is incorrect. At this time, the controller may generate a corresponding prompt message to remind the operation and maintenance personnel to check and correct it. For example, if the target host identifier is the IP address "192.168.1.10", the system will search for a record that is exactly the same in the host identifier field of the asset database. Once a matching record is found, the value of the corresponding business system name field in the record is the target business system name corresponding to the alarm record.

[0043] S120: Acquire an alarm root event corresponding to the alarm record based on the target business system name and the target host identifier.

[0044] The root event of an alarm refers to the most fundamental cause of an alarm. Identifying the root event of an alarm is crucial to solving the problem, because only by processing it can related alarms be eliminated. For example, a network device failure triggers multiple host network connection alarms, and the network device failure is the root event of the alarm.

[0045] Specifically, in the actual operating environment, a fault may trigger multiple related alarms, but only one of these alarms is the root cause. For example, when a network switch fails, multiple servers connected to the switch may simultaneously report alarms of abnormal network connections, but the switch failure is the root event of the alarm.

[0046] Optionally, an alarm root event corresponding to the alarm record is obtained based on the target business system name and the target host identifier, including: determining the alarm time of the alarm record, and determining the retrieval time range according to the alarm time; obtaining an alarm database, and obtaining relevant alarm data with the same name as the target business system and the target host identifier in the alarm database based on the retrieval time range; judging whether there is an uninterrupted alarm before the alarm time in the relevant alarm data, and if so, determining that the alarm record is a relevant alarm, and determining the alarm root event based on the preset uninterrupted time; otherwise, determining that the alarm record is an alarm root event.

[0047] The alarm time in the alarm record accurately marks the moment when the alarm occurred. For example, an alarm record shows the alarm time as "2025-02-0810:30:00", indicating that at this time point, the controller detected an abnormal situation and generated an alarm. In order to fully obtain other alarm data that may be related to the alarm, it is necessary to determine a search time range based on the alarm time. The search time range will be based on the alarm time and extend forward and backward for a certain period of time. For example, if it extends forward by 1 hour and backward by 30 minutes, the search time range is from "2025-02-0809:30:00" to "2025-02-0811:00:00". The alarm database is where all alarm information is stored. It contains a large number of alarm records from various business systems and hosts of the enterprise. After obtaining the alarm database, the controller can perform accurate queries in the database based on the previously determined search time range, as well as the target business system name and target host identifier in the alarm record. For example, a database query statement may set the condition as: all alarm records with a target business system name of "financial business system" and a target host identifier of "192.168.1.10" within the time range of "2025-02-0809:30:00" to "2025-02-0811:00:00", in order to filter out other alarm data that is closely related to the current alarm record.

[0048] It should be noted that an uninterrupted alarm refers to an alarm that exists continuously and without interruption for a period of time. For example, if the network connection is interrupted during a certain period of time, the alarm will continue to be generated, forming an uninterrupted alarm. After obtaining the relevant alarm data, the controller will sort the data according to the alarm time, and then check from the alarm time of the current alarm record to determine whether there is such an uninterrupted alarm. Specifically, the system will check the alarm status field in the alarm data to see if there is an alarm that has been in a triggered state before the alarm time, and there is no record of the alarm being lifted. For example, if an alarm about the server temperature being too high is found, and looking back from the current alarm time, the alarm has been in an activated state without interruption in the past 10 minutes, then it can be determined that this is an uninterrupted alarm.

[0049] Specifically, if an uninterrupted alarm before the alarm time is found in the relevant alarm data, then the current alarm record is determined to be a relevant alarm, that is, it is caused by the persistent problem reflected by this uninterrupted alarm. Then, the controller can determine the alarm root event based on the preset uninterrupted time. The preset uninterrupted time is a time threshold preset based on business experience and system characteristics, and the preset uninterrupted time can be 1 minute. The controller can continue to find the earliest alarm record based on the preset uninterrupted time as the alarm root event. If the uninterrupted alarm before the alarm time is not found in the relevant alarm data, then the current alarm record is directly determined to be the alarm root event.

[0050] S130: Determine a handling status corresponding to the alarm record according to the alarm root event.

[0051] Among them, the disposal status is used to reflect the status information of the progress of handling alarms and related issues. It can be "unhandled", "being processed", "handled", etc. Determining the disposal status helps operation and maintenance personnel to clearly understand the handling status of each alarm and reasonably arrange work priorities.

[0052] The technical solution of the embodiment of the present invention, through asset enrichment of alarm records, alarm event tracing and merging processing, and automatic association, retrieves host performance data and context log data, determines the alarm root event, and then automatically modifies the status of the alarm event and related alarm events, thereby improving the efficiency of fault location, solving the problem of repeated troubleshooting of multiple alarm records, and automatically modifies the disposal status of the alarm root event and related events through the automated processing function, which helps to reasonably arrange operation and maintenance resources, improve operation and maintenance capabilities, and better ensure the stable operation of the business system.

[0053] Embodiment 2

[0054] Figure 2This is a flowchart of an alarm record analysis method provided in the second embodiment of the present invention. This embodiment adds a specific process of determining the corresponding disposal status of the alarm record according to the alarm root event on the basis of the above-mentioned first embodiment. Among them, the specific contents of steps S210-S220 are roughly the same as steps S110-S120 in the first embodiment, so they will not be repeated in this embodiment. Figure 2 As shown, the method includes:

[0055] S210: Acquire an alarm record, and determine the name of a target business system corresponding to the alarm record, wherein the alarm record includes a target host identifier.

[0056] Optionally, obtaining an alarm record includes: obtaining an alarm monitoring configuration item, wherein the alarm monitoring configuration item includes a target host identifier and a monitoring item; determining a target host corresponding to the alarm monitoring configuration item based on the target host identifier; and monitoring the target host based on the monitoring item to obtain the alarm record.

[0057] Optionally, determining the target business system name corresponding to the alarm record includes: obtaining an asset database, wherein the asset database includes the business system name corresponding to each host identifier; matching the target host identifier through the asset database to determine the target business system name corresponding to the alarm record.

[0058] S220: Acquire an alarm root event corresponding to the alarm record based on the target business system name and the target host identifier.

[0059] Optionally, an alarm root event corresponding to the alarm record is obtained based on the target business system name and the target host identifier, including: determining the alarm time of the alarm record, and determining the retrieval time range according to the alarm time; obtaining an alarm database, and obtaining relevant alarm data with the same name as the target business system and the target host identifier in the alarm database based on the retrieval time range; judging whether there is an uninterrupted alarm before the alarm time in the relevant alarm data, and if so, determining that the alarm record is a relevant alarm, and determining the alarm root event based on the preset uninterrupted time; otherwise, determining that the alarm record is an alarm root event.

[0060] S230: Acquire a log database, and locate source data and host performance data corresponding to the alarm root event in the log database.

[0061] Among them, the log database stores logs generated by various systems, applications, and devices. After determining the alarm root event, the controller can locate the source data in the log database according to the alarm time, host ID, alarm triggering rules, etc. of the alarm root event. Source data refers to event records directly related to the alarm root event, which can be exception logs thrown by the application, system configuration change records, etc. At the same time, the controller can associate the host performance data before and after the alarm time according to the alarm time and host ID of the alarm root event. Host performance data refers to the performance of the host before and after the alarm root event occurs, including CPU usage, memory usage, disk I / O rate, and network.

[0062] S240: Generate a root cause based on source data or host performance data.

[0063] Optionally, forming a root cause based on source data or host performance data includes: locating context log data in a log database based on the source data; when the context log data triggers a preset abnormal keyword or the host performance data triggers a preset threshold, obtaining corresponding abnormal data and using the abnormal data as the root cause.

[0064] Specifically, the controller can locate context log data in the log database based on the metadata log. The context log data can provide a detailed sequence of events before and after the alarm root event occurs, helping to analyze the background and process of the problem.

[0065] Among them, the preset abnormal keywords are specific words or phrases that are pre-set based on common failure modes and error types. When the preset abnormal keywords appear in the context log data, it indicates that the system may have corresponding problems. Host performance data such as CPU usage, memory occupancy, disk I / O rate, etc. all have their normal operating ranges. The preset threshold is a reasonable limit set for the host performance data. For example, the preset threshold of CPU usage is set to 80%. When the host performance data shows that the CPU usage exceeds 80%, it indicates that the host may have a performance bottleneck, which may be related to the alarm root event. By setting preset thresholds, the system can quantify the host performance status and detect performance anomalies in a timely manner.

[0066] Specifically, when the context log data triggers a preset abnormal keyword or the host performance data triggers a preset threshold, the controller can obtain the corresponding abnormal data. If the context log data triggers the abnormal keyword, the abnormal data may be the specific log line containing the keyword and several log lines before and after it. If the host performance data triggers the preset threshold, the abnormal data is the specific value of the performance indicator, the trigger time, and the relevant host identification information.

[0067] Optionally, the method further includes: when the context log data does not trigger a preset abnormal keyword or the host performance data does not trigger a preset threshold, determining that the handling status of the alarm record requires manual analysis.

[0068] Specifically, when neither the context log nor the host performance data shows obvious abnormalities, there may be some complex problems that have not been covered by the system's preset rules. At this time, the controller will set the handling status of the alarm record to require manual analysis and generate a prompt to notify the operation and maintenance personnel.

[0069] S250, determine whether automatic processing is enabled, if so, execute S260, otherwise, execute S270.

[0070] Specifically, operation and maintenance personnel can decide whether to enable the automatic handling function based on business needs and risk assessment. After enabling automatic handling, the controller can automatically perform corresponding processing operations for specific types of alarm root events based on preset rules and algorithms. For example, for alarms with a small impact range and relatively fixed processing methods, such as insufficient server disk space, the controller can automatically clean up temporary files to free up space. The controller can determine whether to enable automatic handling by reading a specific configuration file or querying the system settings table. For example, in the configuration file of the operation and maintenance management system, there is a switch parameter "auto-handling-enabled". If its value is "true", it means that automatic handling is enabled; if it is "false", it means that it is not enabled.

[0071] S260: Determine, based on the root cause, that the disposal status of the alarm record and related alarm data is a root cause completion status.

[0072] Specifically, when it is determined that automatic handling is turned on, the controller can update the handling status of the alarm record and related alarm data based on the root cause formed and in accordance with the preset automatic handling rules. For example, if the root cause is "insufficient user permissions leading to database connection failure", and the system presets the automatic handling rule for such problems as "automatically adjust user permissions", then after the system automatically executes the permission adjustment operation, the handling status of the alarm record and related alarm data is set to the root cause completion status, such as "root cause completed, need to be handled", indicating that the root cause problem has been handled and the problem corresponding to the relevant alarm has also been solved.

[0073] S270: Determine, based on the root cause, that the handling status of the alarm record is a root cause completion status.

[0074] Specifically, when it is determined that automatic handling is not enabled, the system will also set the handling status of the alarm record to the root cause completion status based on the root cause, such as "root cause completed, need to be handled". However, at this time, it only means that the system has analyzed the cause of the alarm root event, and the actual handling operation requires manual intervention.

[0075] The technical solution of the embodiment of the present invention, through asset enrichment of alarm records, alarm event tracing and merging processing, and automatic association, retrieves host performance data and context log data, determines the alarm root event, and then automatically modifies the status of the alarm event and related alarm events, thereby improving the efficiency of fault location, solving the problem of repeated troubleshooting of multiple alarm records, and automatically modifies the disposal status of the alarm root event and related events through the automated processing function, which helps to reasonably arrange operation and maintenance resources, improve operation and maintenance capabilities, and better ensure the stable operation of the business system.

[0076] Embodiment 3

[0077] Figure 3 This is a schematic diagram of the structure of an alarm record analysis device provided by Embodiment 3 of the present invention. Figure 3 As shown, the device includes: an alarm record acquisition module 310, which is used to acquire an alarm record and determine the name of the target business system corresponding to the alarm record, wherein the alarm record includes a target host identifier;

[0078] An alarm root event determination module 320 is used to obtain an alarm root event corresponding to the alarm record based on the target business system name and the target host identifier;

[0079] The alarm record analysis module 330 is used to determine the handling status corresponding to the alarm record according to the alarm root event.

[0080] Optionally, the alarm record acquisition module 310 specifically includes: an alarm record acquisition unit, used to: obtain alarm monitoring configuration items, wherein the alarm monitoring configuration items include target host identifiers and monitoring items; determine the target host corresponding to the alarm monitoring configuration items based on the target host identifier; monitor the target host based on the monitoring items to obtain alarm records.

[0081] Optionally, the alarm record acquisition module 310 specifically includes: a target business system name determination unit, used to: obtain an asset database, wherein the asset database includes the business system name corresponding to each host identifier; match the target host identifier through the asset database to determine the target business system name corresponding to the alarm record.

[0082] Optionally, the alarm root event determination module 320 is specifically used to: determine the alarm time of the alarm record, and determine the retrieval time range based on the alarm time; obtain the alarm database, and obtain relevant alarm data with the same name as the target business system and the target host identifier in the alarm database based on the retrieval time range; determine whether there is an uninterrupted alarm before the alarm time in the relevant alarm data, and if so, determine that the alarm record is a related alarm, and determine the alarm root event based on the preset uninterrupted time; otherwise, determine that the alarm record is an alarm root event.

[0083] Optionally, the alarm record analysis module 330 is specifically used to: obtain a log database, locate the source data and host performance data corresponding to the alarm root event in the log database; form a root cause based on the source data or the host performance data; determine whether to enable automatic disposal, and if so, determine the disposal status of the alarm record and related alarm data as the root cause completion status based on the root cause; otherwise, determine the disposal status of the alarm record as the root cause completion status based on the root cause.

[0084] Optionally, the alarm record analysis module 330 specifically includes: a root cause generation unit, used to: locate context log data in the log database based on source data; when the context log data triggers a preset abnormal keyword or the host performance data triggers a preset threshold, obtain the corresponding abnormal data and use the abnormal data as the root cause.

[0085] Optionally, the alarm record analysis module 330 further includes: a manual analysis unit, which is used to: when the context log data does not trigger a preset abnormal keyword or the host performance data does not trigger a preset threshold, determine that the handling status of the alarm record requires manual analysis.

[0086] The technical solution of the embodiment of the present invention, through asset enrichment of alarm records, alarm event tracing and merging processing, and automatic association, retrieves host performance data and context log data, determines the alarm root event, and then automatically modifies the status of the alarm event and related alarm events, thereby improving the efficiency of fault location, solving the problem of repeated troubleshooting of multiple alarm records, and automatically modifies the disposal status of the alarm root event and related events through the automated processing function, which helps to reasonably arrange operation and maintenance resources, improve operation and maintenance capabilities, and better ensure the stable operation of the business system.

[0087] An alarm record analysis device provided in an embodiment of the present invention can execute an alarm record analysis method provided in any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution method.

[0088] Embodiment 4

[0089] Figure 4A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0090] like Figure 4 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0091] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0092] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as an alarm record analysis method.

[0093] In some embodiments, an alarm record analysis method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the alarm record analysis method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform an alarm record analysis method in any other appropriate manner (e.g., by means of firmware).

[0094] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0095] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0096] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, device, or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0097] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).

[0098] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0099] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.

[0100] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0101] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A method for analyzing alarm records, characterized in that: include: Obtain an alarm record, and determine the name of the target business system corresponding to the alarm record, wherein the alarm record includes a target host identifier; Acquire an alarm root event corresponding to the alarm record based on the target service system name and the target host identifier; A handling status corresponding to the alarm record is determined according to the alarm root event.

2. The method according to claim 1, characterized in that The obtaining of alarm records includes: Obtain an alarm monitoring configuration item, wherein the alarm monitoring configuration item includes a target host identifier and a monitoring item; Determine a target host corresponding to the alarm monitoring configuration item based on the target host identifier; The target host is monitored based on the monitoring items to obtain alarm records.

3. The method according to claim 1, characterized in that: The determining the name of the target business system corresponding to the alarm record includes: Acquire an asset database, wherein the asset database includes a business system name corresponding to each host identifier; The target host identifier is matched through the asset database to determine the target business system name corresponding to the alarm record.

4. The method according to claim 1, characterized in that The acquiring the alarm root event corresponding to the alarm record based on the target service system name and the target host identifier includes: Determine the alarm time of the alarm record, and determine the retrieval time range according to the alarm time; Acquire an alarm database, and acquire relevant alarm data identical to the target business system name and the target host identifier in the alarm database based on the search time range; Determine whether there is an uninterrupted alarm before the alarm time in the relevant alarm data, and if so, determine that the alarm record is a relevant alarm, and determine the alarm root event based on the preset uninterrupted time; Otherwise, determine that the alarm record is an alarm root event.

5. The method according to claim 4, characterized in that The determining, according to the alarm root event, a handling status corresponding to the alarm record includes: Acquire a log database, and locate source data and host performance data corresponding to the alarm root event in the log database; forming a root cause according to the source data or the host performance data; Determine whether to enable automatic handling, and if so, determine the handling status of the alarm record and related alarm data as the root cause completion status based on the root cause; Otherwise, the handling status of the alarm record is determined to be a root cause completion status based on the root cause.

6. The method according to claim 5, characterized in that The forming of the root cause according to the source data or the host performance data includes: locating context log data in the log database according to the source data; When the context log data triggers a preset abnormal keyword or the host performance data triggers a preset threshold, corresponding abnormal data is obtained and the abnormal data is used as a root cause.

7. The method according to claim 6, characterized in that The method further comprises: When the context log data does not trigger a preset abnormal keyword or the host performance data does not trigger a preset threshold, it is determined that the handling status of the alarm record requires manual analysis.

8. An alarm record analysis device, characterized in that: include: An alarm record acquisition module is used to acquire an alarm record and determine the name of a target business system corresponding to the alarm record, wherein the alarm record includes a target host identifier; An alarm root event determination module, used to obtain an alarm root event corresponding to the alarm record based on the target business system name and the target host identifier; The alarm record analysis module is used to determine the handling status corresponding to the alarm record according to the alarm root event.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 7.

10. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method according to any one of claims 1 to 7 when executed.

Citation Information

Cited By

  • Automatic adaptation method, system and device for multi-source heterogeneous alarm, medium and program product

    CN121037202A