End-to-end monitoring system alarm compression processing method and device combined with large model, and medium
By combining a large model with an alarm compression processing method, the problems of redundant and invalid alarms in the end-to-end monitoring system are solved, alarm management and rapid fault location at the business level are achieved, and operation and maintenance costs are reduced.
Patent Information
- Application Number
- CN202510995409.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-09-09
AI Technical Summary
Existing end-to-end monitoring systems have a large number of redundant and invalid alarms when processing alarms, resulting in high operation and maintenance costs and difficulty in quickly locating the root cause of the problem. In addition, existing solutions lack business-level management of alarms.
By combining a large model with an alarm compression processing method, alarm events in the business dimension are generated through alarm level generation, homologous alarm merging, aggregation rule upgrades, and a fault knowledge base, and the large model is used to provide fault handling suggestions.
It realizes business-level alarm processing, reduces redundant alarms, lowers the operation and maintenance threshold, quickly locates the root cause of the problem, and improves fault handling efficiency.
Smart Images

Figure CN120614282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network monitoring technology, and in particular to an alarm compression processing method, device and medium for an end-to-end monitoring system combined with a large model. Background Art
[0002] An end-to-end monitoring platform is a monitoring system that integrates multiple monitoring components and tools. It can comprehensively monitor and manage the entire application system, from client to server, covering the entire application chain. It can help operations and maintenance personnel quickly locate and resolve system anomalies or failures. For an end-to-end monitoring system, accurately locating the source of the problem and obtaining accurate and effective alarm information is crucial. However, in reality, monitoring resources often generate batch alarms, invalid alarms, and duplicate alarms. Maintaining a large amount of alarm information significantly increases operations and maintenance costs and hinders locating the root cause of the problem. Furthermore, handling alarms requires sufficient skills from operations and maintenance personnel, which is a high barrier to entry.
[0003] Large models are machine learning models with large parameters and complex computational structures. These models are typically built using deep neural networks and are capable of handling complex tasks and data. Large models are widely used in various fields, including natural language processing, computer vision, speech recognition, and recommendation systems.
[0004] Existing alarm compression solutions for end-to-end monitoring systems have certain limitations. The compression and merging of alarms only stays at the transaction level, and there is a lack of management of alarms at the business level. With the rise of big model technology, how to lower the threshold for alarm processing and improve fault handling efficiency has also become a new direction for end-to-end monitoring. Summary of the Invention
[0005] In response to the deficiencies in the prior art, the present invention provides an end-to-end monitoring system alarm compression processing method, device and medium combined with a large model. By compressing and upgrading the alarms, business events of concern to operation and maintenance personnel are obtained, thereby helping operation and maintenance personnel to quickly locate the root cause of the problem. In addition, by building a fault knowledge base and combining it with a large model tool, general processing steps for corresponding fault types are given, thereby improving the efficiency of fault processing.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] A method for compressing and processing alarms in an end-to-end monitoring system combined with a large model comprises the following steps:
[0008] Periodically scan resource status. If the resource status meets the alarm condition, an alarm original event is generated according to the alarm level and marked as a triggered state. Otherwise, the alarm original event is marked as a recovered state.
[0009] When generating an alarm original event, determine whether there is an alarm event of the same source based on the alarm identifier. If there is an alarm event of the same source that has not been restored, then the current alarm original event will be included in the existing alarm event of the same source, and the alarm original event list will be updated, and the alarm status will be replaced. If there is no alarm event of the same source or the alarm event of the same source has been restored and can no longer accommodate a new alarm original event, then a new alarm will be generated based on this alarm original event.
[0010] When an alarm is triggered or updated, all currently enabled aggregation rules are traversed to find the aggregation rule associated with the alarm; the aggregation rule is used to aggregate the alarm event into a business-level event, and the business events under the aggregation rule are queried in turn to escalate the alarm and generate an alarm event at the business level;
[0011] Use the big model to provide troubleshooting suggestions for alarm events in the business dimension.
[0012] To optimize the above technical solutions, specific measures taken also include:
[0013] Furthermore, if the resource status satisfies the alarm condition, an alarm original event is generated according to the alarm level, and the alarm original event is marked as a trigger state. Specifically:
[0014] If the resource status meets the alarm trigger threshold or floating percentage judgment and is within the time window of the last trigger, the trigger count is increased by one. If the last trigger time has exceeded the time window range, the current trigger is updated to the first trigger and the trigger time is recorded; if the number of triggers has reached the configured threshold trigger count, an alarm original event is generated according to the alarm level, and the alarm organization, alarm indicator, alarm description and collection source information configured in the alarm rule are written as tags into the alarm original event, and a hash code is generated based on the tag information as the alarm identifier, marking this alarm original event as a triggered state.
[0015] Furthermore, the identifier of the alarm is specifically a hash code generated according to the tag information; the tag information includes the organization to which the alarm belongs, the alarm indicator, the alarm description and the collection source information configured by the alarm rule.
[0016] Furthermore, the method for determining the aggregation rule associated with the alarm is:
[0017] Determine whether the alarm tag matches the tag configured in the aggregation rule. If so, the alarm is associated with the aggregation rule.
[0018] Furthermore, the business events under the aggregation rules are sequentially queried and the alarm upgrade is specifically as follows:
[0019] Query the business events under the aggregation rules in sequence. If there is no business event under this aggregation rule, or there is an alarm event but the difference between the current time and the latest alarm update time in the business event has exceeded the time window range, then generate a new business event and include the alarm in this business event; on the contrary, if there is a business event and the current time is still in the time window of this business event, then update this alarm and associate it with the current business event; business event is the highest level of alarm aggregation, which is used to associate and aggregate alarms belonging to different resources of the same business, and generate business-dimensional event alarms for dispatch processing; when all associated alarms of a business event are in recovery state, the business event will be marked as recovery state, and the rest are in trigger state. The level of the business event is the highest alarm level among the associated alarms.
[0020] Furthermore, the use of the big model to provide fault handling suggestions for alarm events in the business dimension is specifically as follows:
[0021] When a work order triggered by an alarm event in the business dimension is closed and the work order allows knowledge generation, a knowledge title is generated based on the performance indicator that triggered the alarm and the trigger scenario configured in the alarm rule. The work order processing information, including the fault handling plan, fault contact person, fault duration, and fault impact range, is read as the knowledge content to obtain the knowledge data generated based on the closed work order.
[0022] The knowledge base is searched for existing knowledge data based on the knowledge title. If no such knowledge exists, it is directly stored in the knowledge base and marked as untrained. If duplicate titles are retrieved, the generated knowledge is stored in the pending knowledge base. Such knowledge is confirmed by the operation and maintenance personnel to determine whether to overwrite, ignore, or directly store it in the knowledge base. The knowledge in the knowledge base is used to train the big model. Specifically, the knowledge that needs to be synchronized to the big model is regularly read. Based on the title and content of the knowledge, data documents are generated in the form of question-answer pairs. The generated data documents are then imported into the big model for learning.
[0023] When troubleshooting suggestions are needed, the alarm event of the business dimension associated with the work order is read to generate a query title, and the big model interface query is called. A websocket connection is established using the query title as the question statement to obtain the troubleshooting suggestions given by the big model.
[0024] The present invention also proposes an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements the end-to-end monitoring system alarm compression processing method combined with a large model as described above.
[0025] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the above-mentioned end-to-end monitoring system alarm compression processing method combined with a large model.
[0026] The beneficial effects of the present invention are: 1. It can provide an alarm processing flow at the business level. 2. By setting the time window and the number of triggering times of the alarm triggering rules, a large number of redundant alarms are avoided. By setting the aggregation rules, alarms can be aggregated from the business level, which is convenient for operation and maintenance personnel to handle. 3. Combining the design of the big model and the knowledge base, it can provide assistance for repairing faults. Compared with the existing technology, the main creativity lies in: the design scheme of the present application is not limited to the alarm processing of a single resource, but provides an alarm processing scheme at the business level, which can effectively avoid redundant alarms and invalid alarms, and can also help operation and maintenance personnel quickly locate the cause of the problem and understand the service status. The validity of the fault record is guaranteed by the multi-level compression aggregation of the original event. While ensuring the quality of the alarm, combined with the big model, it can provide iterative assistance for repairing faults, further reducing the pressure of operation and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 This is an overall flow chart of the alarm compression processing method for an end-to-end monitoring system combined with a large model proposed by the present invention.
[0028] Figure 2 This is the flow chart for compressing and processing the original alarm event.
[0029] Figure 3 This is the alarm upgrade process flow chart.
[0030] Figure 4 Suggested flow chart for alerting. DETAILED DESCRIPTION
[0031] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0032] Example 1
[0033] The present invention proposes an end-to-end monitoring system alarm compression processing method combined with a large model. The overall process of the method is as follows: Figure 1 As shown in the figure, the system mainly includes four steps: alarm scanning, alarm compression, alarm escalation, and using the big model to provide troubleshooting suggestions for alarm events in the business dimension. The scanning process is used to trigger the original alarm event. The compression process generates or updates existing alarms based on the original alarm event. The escalation process aggregates related alarms into business events.
[0034] Before the method of the present invention is implemented, alarm triggering rules and aggregation rules need to be formulated.
[0035] Alarm rule settings: Configure corresponding alarm rules based on the metrics of managed resources. These include generating alarms for real-time data or interval trends, thresholds or floating percentages for alarm triggering and clearing, alarm levels corresponding to different thresholds or percentages, system properties of the rule itself, the rule's scanning interval, the alarm calculation time window, the number of times within the window that an alarm is triggered before it is generated, and supplementary descriptions after the alarm is triggered. Generate an alarm rule table.
[0036] Setting Alarm Aggregation Rules: Alarm aggregation rules aggregate and compress alarms when they are generated. Setting aggregation rules requires specifying the organizational scope, time window, alarm filtering, alarm rule scope, and whether to automatically generate knowledge. An aggregation rule table is generated. Users can associate different resources within the same business through aggregation rules. If problems with these resources cause business anomalies, events are generated for subsequent processing.
[0037] (1) Alarm scanning: Read the resource status monitored by the alarm rule scanning system according to the scanning interval period configured by the alarm rule. If the resource status meets the alarm trigger threshold or the floating percentage judgment and is within the time window range of the last trigger, the trigger count is increased by one. If the last trigger time has exceeded the time window range, the current trigger is updated to the first trigger and the trigger time is recorded. If the number of triggers has reached the configured threshold trigger count, an alarm original event is generated according to the alarm level, and the alarm organization, alarm indicator, alarm description, collection source and other information configured by the rule are written as tags into the original event. A hash code is generated based on the tag information as the alarm identifier, marking this alarm original event as a triggered state; on the contrary, if the threshold or interval floating percentage meets the recovery alarm threshold, the alarm original state is marked as a recovered state.
[0038] (2) Alarm Compression: Refer to the original alarm event compression processing flow chart Figure 2 When generating an original alarm event, the alarm identifier (i.e., the hash code generated based on the configured tag information) is used to determine whether there is an alarm data from the same source. If there is an unrecovered alarm from the same source, the current original alarm event is merged into the existing alarm event, and the original event list of the alarm is updated, replacing the alarm status (triggered, recovered, or alarm level updated). If there is no alarm from the same source or the alarm from the same source has been recovered and can no longer accommodate a new original event, a new alarm is generated based on this original event.
[0039] (3) Alarm upgrade: refer to the alarm upgrade process Figure 3. The role of the aggregation rule is to aggregate the alarm data in (2) into business-level events. The aggregation rule needs to configure the rule's attribution information, the time window for aggregation calculation, the alarm tag field for aggregation filtering, the scanning interval of the aggregation rule, and whether it is enabled. When an alarm is triggered or updated, all currently active aggregation rules are traversed to find the aggregation rule associated with this alarm. The method for determining whether the aggregation rule is associated with this alarm is: determine whether the alarm tag matches the tag configured in the aggregation rule. If no matching aggregation rule is found, ignore this aggregation rule and do not process it, and this alarm will not be upgraded to a business event. After finding all the associated aggregation rules, query the business events under this rule separately to determine whether there is a business event under this aggregation rule that can absorb the alarm. If there is no business event data under this aggregation rule, or there is an alarm event but the difference between the current time and the latest alarm update time in the business event has exceeded the time window range, a new business event is generated and the alarm is included in this business event; conversely, if there is a business event and the current time is still in the time window of this event, the alarm is updated and associated with the current business event.
[0040] Events are the highest level of alert aggregation. They aggregate alerts from different resources belonging to the same business, generating business-specific event alerts for dispatch and processing. A business event is marked as recovered only when all associated alerts are in the recovered state. All other alerts are in the triggered state. The severity of a business event is the highest severity among the associated alerts.
[0041] Users can query the event list for detailed information about business events, including a list of associated alarms, the source of the associated alarms, the trigger value, the trigger time, and the status of each alarm. They can also trigger dispatch actions and manual recovery operations on the page for subsequent processing.
[0042] (4) Use the big model to provide troubleshooting suggestions for alarm events in the business dimension: When the operation and maintenance personnel handle the work order, they can click on the repair suggestion on the details page to query the general repair plan for this type of fault. During the repair process, fill in the fault repair process, fault impact range, and other processing details according to the actual situation. After the fault repair is completed, the alarm scan will automatically restore the associated alarms. Before closing the work order, the operation and maintenance personnel choose whether to generate knowledge based on whether the suggestion is valid or needs to be updated (the configuration of the aggregation rule is checked by default). If the provided handling suggestions are insufficient, the knowledge generated after closing the work order will provide more appropriate suggestions the next time such a fault occurs through a new round of training process.
[0043] See also Figure 4 , the fault handling suggestions are divided into three steps: 1. Work order storage; 2. Semantic data generation; 3. Alarm suggestion generation. The specific process is as follows:
[0044] 1. When the work order triggered by the business-level alarm is closed and the work order allows knowledge generation (the work order inherits the setting of whether to generate knowledge configured by the aggregation rule by default, which can be modified), the background performs the knowledge base storage operation: the title of the knowledge base is generated according to the performance indicators that triggered the alarm and the triggering scenarios configured by the alarm rules (resource type, abnormal indicators, customer, alarm level, etc.). Read the processing information of the alarm work order, including fault handling plan, fault contact person, fault duration, fault impact range and other information as the content of the knowledge base, and obtain the knowledge data generated based on the closed work order. According to the newly generated knowledge title, the existing data in the database is retrieved. If there is no such knowledge, it is directly stored in the knowledge base and marked as untrained; if a duplicate title is retrieved, the generated knowledge is stored in the knowledge base to be confirmed. The operation and maintenance personnel confirm whether such knowledge items are overwritten (update the existing knowledge content and mark it as untrained), ignored or directly stored in the knowledge base (marked as untrained).
[0045] 2. The system's knowledge base query interface provides an interface for filtering untrained knowledge, allowing operations personnel to determine whether it needs to be synchronized to the big model for processing. Knowledge items that need to be synchronized to the big model are regularly read, and data documents are generated in the form of question-and-answer pairs based on the knowledge base title and content. These generated data documents are then imported into the big model for processing (building an internal database to prevent sensitive data leakage).
[0046] 3. When operations personnel obtain handling suggestions for an alarm ticket, they read the alarm information associated with the ticket to generate a query title, call the large model interface for a query, establish a websocket connection using this query title as the question statement, and obtain the troubleshooting suggestions provided by the large model. Before closing the ticket, operations personnel can choose whether to modify the generated knowledge configuration to maintain the timeliness and accuracy of the knowledge base and achieve closed-loop feedback on troubleshooting suggestions.
[0047] Example 2
[0048] The present invention proposes an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the end-to-end monitoring system alarm compression processing method combined with a large model as described in Example 1 is implemented.
[0049] Example 3
[0050] The present invention provides a computer-readable storage medium storing a computer program, wherein the computer program enables a computer to execute the alarm compression processing method of an end-to-end monitoring system combined with a large model as described in the first embodiment.
[0051] In the embodiments disclosed herein, computer storage media can be tangible media that can contain or store programs for use by or in conjunction with an instruction execution system, device, or apparatus. Computer storage media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. More specific examples of computer storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0052] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0053] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that for those skilled in the art, various improvements and modifications that do not depart from the principles of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A method for compressing alarms in an end-to-end monitoring system using a large model, characterized in that: The following steps are involved: Periodically scan resource status. If the resource status meets the alarm condition, an alarm original event is generated according to the alarm level and marked as a triggered state. Otherwise, the alarm original event is marked as a recovered state. When generating an alarm original event, determine whether there is an alarm event of the same source based on the alarm identifier. If there is an alarm event of the same source that has not been restored, then the current alarm original event will be included in the existing alarm event of the same source, and the alarm original event list will be updated, and the alarm status will be replaced. If there is no alarm event of the same source or the alarm event of the same source has been restored and can no longer accommodate a new alarm original event, then a new alarm will be generated based on this alarm original event. When an alarm is triggered or updated, all currently enabled aggregation rules are traversed to find the aggregation rule associated with the alarm; the aggregation rule is used to aggregate the alarm event into a business-level event, and the business events under the aggregation rule are queried in turn to escalate the alarm and generate an alarm event at the business level; Use the big model to provide troubleshooting suggestions for alarm events in the business dimension.
2. The alarm compression processing method for an end-to-end monitoring system combined with a large model according to claim 1, characterized in that: If the resource status meets the alarm condition, an alarm original event is generated according to the alarm level, and the alarm original event is marked as a trigger state. Specifically: If the resource status meets the alarm trigger threshold or floating percentage, and is within the time window of the last trigger, the trigger count is increased by one. If the last trigger time exceeds the time window, the current trigger is updated to the first trigger and the trigger time is recorded. If the number of triggers reaches the configured threshold, an alarm original event is generated according to the alarm level, and the alarm organization, alarm indicator, alarm description, and collection source information configured in the alarm rule are written as tags to the alarm original event. A hash code is generated based on the tag information as the alarm identifier, marking this alarm original event as a triggered state.
3. The alarm compression processing method for an end-to-end monitoring system combined with a large model according to claim 1, characterized in that: The identifier of the alarm is specifically a hash code generated according to the tag information; the tag information includes the organization to which the alarm belongs, the alarm indicator, the alarm description and the collection source information configured by the alarm rule.
4. The alarm compression processing method for an end-to-end monitoring system combined with a large model according to claim 1, characterized in that: The method for determining the aggregation rule associated with the alarm is: Determine whether the alarm tag matches the tag configured in the aggregation rule. If so, the alarm is associated with the aggregation rule.
5. The alarm compression processing method for an end-to-end monitoring system combined with a large model according to claim 1, characterized in that: The business events under the aggregation rules are sequentially queried and the alarm upgrade is specifically as follows: Query the business events under the aggregation rules in sequence. If there is no business event under this aggregation rule, or if there is an alarm event but the difference between the current time and the latest alarm update time in the business event is outside the time window range, generate a new business event and include the alarm in this business event. Conversely, if there is a business event and the current time is still within the time window of this business event, update the alarm and associate it with the current business event. Business events are the highest level of alarm aggregation, used to associate and aggregate alarms belonging to different resources of the same business, generating business-dimensional event alarms for dispatch processing. When all associated alarms of a business event are in the recovery state, the business event will be marked as the recovery state. The rest are in the trigger state. The level of the business event is the highest alarm level among the associated alarms.
6. The alarm compression processing method for an end-to-end monitoring system combined with a large model according to claim 1, characterized in that: The specific troubleshooting suggestions provided by the big model for business-dimensional alarm events are as follows: When a work order triggered by an alarm event in the business dimension is closed and the work order allows knowledge generation, a knowledge title is generated based on the performance indicator that triggered the alarm and the trigger scenario configured in the alarm rule. The work order processing information, including the fault handling plan, fault contact person, fault duration, and fault impact range, is read as the knowledge content to obtain the knowledge data generated based on the closed work order. Retrieve existing knowledge data from the knowledge base based on the knowledge title. If there is no such knowledge, store it directly in the knowledge base and mark it as untrained. If duplicate titles are retrieved, the generated knowledge is stored in the pending knowledge base. Operations and maintenance personnel will confirm whether such knowledge should be ignored or directly stored in the knowledge base. The knowledge in the knowledge base is used to train the big model. Specifically, the knowledge that needs to be synchronized to the big model is regularly read. Based on the title and content of the knowledge, data documents are generated in the form of question-answer pairs. The generated data documents are then imported into the big model for learning. When troubleshooting suggestions are needed, the alarm event of the business dimension associated with the work order is read to generate a query title, and the big model interface query is called. A websocket connection is established using the query title as the question statement to obtain the troubleshooting suggestions given by the big model.
7. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for compressing and processing alarms of an end-to-end monitoring system combined with a large model as described in any one of claims 1 to 6 is implemented.
8. A computer-readable storage medium storing a computer program, characterized in that: The computer program enables the computer to execute the end-to-end monitoring system alarm compression processing method combined with a large model as described in any one of claims 1-6.