Abnormal event processing method, electronic device, and storage medium
By acquiring and caching abnormal events within a preset time period, and utilizing the time bucket mechanism and convergence conditions to achieve bidirectional aggregation, the problem of difficulty in determining the aggregation range in the time dimension in existing technologies is solved, and the level of fault operation and maintenance is improved.
Patent Information
- Application Number
- CN202210678899.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-16
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2042-06-16
AI Technical Summary
In the existing technology, it is difficult to determine the time range in the time dimension of aggregation analysis, resulting in the inability to effectively identify the root cause of the fault caused by the previous operations before the alarm, low aggregation capability, and low fault operation and maintenance level.
By acquiring multiple abnormal events at the target location within a preset time period, the aggregation point is determined, and bidirectional aggregation is performed in the time dimension, including alarms, key performance indicator anomalies, and operation logs. The abnormal events are cached using a time bucket mechanism, and convergence conditions are set to control the caching, thereby achieving the aggregation of data before and after.
The data source aggregation capability has been improved, enabling forward and backward aggregation in the time dimension, clarifying the root cause of the fault and improving the level of fault operation and maintenance.
Smart Images

Figure CN117290133B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of, but is not limited to, communications technology, and in particular to an abnormal event processing method, electronic equipment, and storage medium. Background Art
[0002] With the development of mobile communications technology, network complexity, application diversity, and data explosion, the demand for intelligent operations and maintenance (O&M) is increasing. Aggregating relevant streaming data and then analyzing it is the primary method for identifying the root cause of a fault. However, related technologies often only aggregate data sources after an alarm occurs, and then aggregate them backwards. If an alarm is triggered by an operation before it occurs, such aggregation cannot determine the root cause of the fault. Consequently, aggregation capabilities are limited, leading to poor O&M. Summary of the Invention
[0003] The embodiments of the present invention provide an abnormal event processing method, an electronic device, and a storage medium, which realize bidirectional aggregation, can improve the aggregation capability of data sources, and improve the level of fault operation and maintenance.
[0004] In a first aspect, an embodiment of the present invention provides a method for handling abnormal events, the method comprising: obtaining multiple abnormal events at a target location within a preset time period, the abnormal events comprising at least one of an alarm, a key performance indicator abnormality, and an operation log; determining an aggregation point in the abnormal events; and performing aggregation based on the aggregation point and the abnormal events to obtain an aggregation result.
[0005] In a second aspect, an embodiment of the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the abnormal event handling method as described in any one of the embodiments of the first aspect of the present invention.
[0006] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the storage medium stores a program, and the program is executed by a processor to implement the abnormal event handling method as described in any one of the embodiments of the first aspect of the present invention.
[0007] The embodiments of the present invention include at least the following beneficial effects: the abnormal event processing method, electronic device and storage medium in the embodiments of the present invention, by executing the abnormal event processing method, can continuously obtain multiple abnormal events at the target location within a preset time period, the target location is a link, a network element or a computer room in space, the abnormal event includes at least one of an alarm, a key performance indicator abnormality and an operation log, thereby realizing the acquisition of multiple data sources, and then determining an aggregation point in the abnormal event, the aggregation point can be any calibrated abnormal event therein, when aggregating, the embodiments of the present invention can aggregate according to the aggregation point, and obtain an aggregation result according to the aggregation point and the abnormal event, so as to perform root cause analysis, because multiple abnormal events are obtained within a period of time, when aggregating, the required abnormal events can be aggregated before the time node of the aggregation point according to the time node and position of the aggregation point, so that the embodiments of the present invention can not only aggregate backward, but also aggregate forward, aggregate other events that may be the root cause of the fault, realize bidirectional aggregation, improve the aggregation capability of the data source, and improve the level of fault operation and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 This is a flow chart of a method for handling abnormal events provided by one embodiment of the present invention;
[0009] Figure 2 is a flowchart of a method for handling abnormal events provided by another embodiment of the present invention;
[0010] Figure 3 is a flowchart of a method for handling abnormal events provided by another embodiment of the present invention;
[0011] Figure 4 is a flowchart of a method for handling abnormal events provided by another embodiment of the present invention;
[0012] Figure 5 is a schematic diagram of a target cache area provided by one embodiment of the present invention;
[0013] Figure 6 is a flowchart of a method for handling abnormal events provided by another embodiment of the present invention;
[0014] Figure 7 is a flowchart of a method for handling abnormal events provided by another embodiment of the present invention;
[0015] Figure 8 is a flowchart of a method for handling abnormal events provided by another embodiment of the present invention;
[0016] Figure 9 This is a schematic diagram of forward and backward bidirectional aggregation using communication anomalies as aggregation points, provided by an embodiment of the present invention;
[0017] Figure 10 This is a schematic diagram of forward and backward bidirectional aggregation using a network disconnection alarm as an aggregation point, provided by an embodiment of the present invention;
[0018] Figure 11 is a flowchart of a method for handling abnormal events provided by another embodiment of the present invention;
[0019] Figure 12 is a flowchart of a method for handling abnormal events provided by another embodiment of the present invention;
[0020] Figure 13 is a flowchart of a method for handling abnormal events provided by another embodiment of the present invention;
[0021] Figure 14 is a flowchart of a method for handling abnormal events provided by another embodiment of the present invention;
[0022] Figure 15 is a flowchart of a method for handling abnormal events provided by another embodiment of the present invention;
[0023] Figure 16 FIG. 1 is a schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0025] In the description of the present invention, it should be understood that descriptions involving orientations, such as up, down, front, back, left, right, etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, they cannot be understood as limitations on the embodiments of the present invention.
[0026] It should be understood that in the description of the embodiments of the present invention, "several" means more than one, "multiple" means more than two, "greater than," "less than," and "exceed" are understood to exclude the number itself, while "above," "below," and "within" are understood to include the number itself. The use of terms such as "first" and "second" is solely for the purpose of distinguishing technical features and is not to be construed as indicating or implying relative importance, or implicitly indicating the number of the indicated technical features, or implicitly indicating the order of the indicated technical features.
[0027] In the description of the embodiments of the present invention, unless otherwise clearly defined, terms such as setting, installing, and connecting should be understood in a broad sense, and technicians in the relevant technical field can reasonably determine the specific meanings of the above terms in the embodiments of the present invention based on the specific content of the technical solution.
[0028] With the continuous advancement and development of 5G infrastructure, increasing network complexity, diverse applications, and data explosion, operators and equipment manufacturers are increasingly demanding automation and intelligence in the "planning, construction, maintenance, optimization, and operation" of autonomous networks. "Maintenance" refers to operations and maintenance, focusing on fault handling. Fault location has evolved from aggregated analysis of a single alarm data source to aggregated analysis of multiple data sources, such as logs, key performance indicators (KPIs), and alarms.
[0029] Aggregation analysis involves aggregating relevant streaming data and then analyzing it to identify the root cause of a fault. Aggregation occurs in two dimensions: time and space, also known as spatiotemporal aggregation. Spatial aggregation leverages topological resource dependencies, such as data from the same network element, link, or computer room. Temporal aggregation aggregates relevant data within a specific timeframe. Unlike spatial aggregation, temporal aggregation is more challenging because the timeframe is often difficult to determine.
[0030] For example, within a certain time window in a spatial dimension, alarm A causes alarm B. Operators and equipment vendors have developed some rules based on experience gained from historical data statistics. The following three rules are used as examples. The basic format of the rules can be:
[0031] 1) The first rule states that if a high link bit error rate alarm of a remote radio unit (RRU) and an abnormal optical module receive optical power occur simultaneously within a 5-minute window on the same network element, these two alarms are considered aggregated.
[0032] 1) The second rule states that if the optical module receive optical power anomaly and the RRU link down alarm occur simultaneously within a 10-minute window on the same network element, the two alarms are considered aggregated.
[0033] 2) The third rule is that for the same network element, within a 15-minute window, an RRU link disconnection alarm and a distributed unit (DU) cell out-of-service alarm are considered to be aggregated.
[0034] Based on the above rules, it is necessary to aggregate the RRU link high bit error rate alarm, optical module receive optical power abnormality alarm, RRU link disconnection alarm, and DU cell out-of-service alarm in the temporal and spatial dimensions. Finally, the root cause of the DU cell out-of-service caused by the high RRU link bit error rate can be found.
[0035] The applicant found that in the relevant technology, the time in the time dimension is difficult to determine, and a design concept of time step is required. It cannot be determined by the maximum time of several rules (15 minutes), nor can it be determined by the sum of the time of all relevant rules. Not only that, there is also a key dependence, that is, the time dimension is completely dependent on backward judgment, that is, the occurrence of alarm A causes the occurrence of alarm B, then the time of occurrence of alarm A will be before the occurrence of alarm B, so the time of occurrence of alarm B is predicted after the occurrence of alarm A.
[0036] The applicant found that there are many problems with aggregation in this case. In this case, the alarm occurrence time may be different, and it is even possible that alarm B is before alarm A. In addition, it is impossible to aggregate multiple data sources. Because if only alarms and key performance indicators are abnormal, it is relatively easy to identify the time when the abnormality occurs, that is, there is a clear abnormal data source. The alarm causes the key performance indicator to deteriorate or be abnormal, then the alarm occurs first, and the alarm is aggregated later. However, if a certain operation causes the relevant alarm, for example, the log of a certain operation is before the alarm, it is not convenient to aggregate backwards. Since the log is not convenient for identifying the abnormality, it cannot be perceived immediately. It is often the case that after the alarm or key performance indicator is abnormal, we look back for related logs. For example, for faults such as memory leaks, we have already discovered the memory leak or the leakage trend, and then we look back for related logs. This is post-aggregation.
[0037] Therefore, the solutions in the related technologies have technical defects. If the alarm is triggered by a certain operation before the alarm, such aggregation cannot clearly identify the root cause of the fault, so the aggregation capability is low, resulting in a low level of fault operation and maintenance.
[0038] Based on this, an embodiment of the present invention provides an abnormal event processing method, an electronic device, and a storage medium, which can achieve two-way aggregation, improve the aggregation capability of data sources, and improve the level of fault operation and maintenance.
[0039] The detailed description is given below.
[0040] The embodiment of the present invention provides a method for handling abnormal events, referring to Figure 1 As shown, the abnormal event handling method in the embodiment of the present invention includes but is not limited to steps S101 to S103.
[0041] Step S101: Acquire multiple abnormal events at a target location within a preset time period, where the abnormal events include at least one of an alarm, a key performance indicator abnormality, and an operation log.
[0042] Step S102, determining the aggregation point in the abnormal event.
[0043] Step S103: Aggregate based on the aggregation points and abnormal events to obtain an aggregation result.
[0044] In one embodiment, the abnormal event handling method in the embodiment of the present invention can be applied in a communication device. By executing the abnormal event handling method, bidirectional aggregation can be achieved, the aggregation capability of the data source can be improved, and the level of fault operation and maintenance can be improved. Specifically, in the embodiment of the present invention, multiple abnormal events at the target location can be obtained within a preset time period, and an aggregation point can be determined in the abnormal events obtained. The abnormal event includes at least one of an alarm, a key performance indicator (KPI) abnormality, and an operation log. In one embodiment, the abnormal event includes one of an alarm or a key performance indicator abnormality, and also includes an operation log. Alternatively, in another embodiment, the abnormal event includes an alarm, a key performance indicator abnormality, and an operation log. In the embodiment of the present invention, the above three are used as an example for explanation. The aggregation point is a point set according to the aggregation needs, and the specific type of the aggregation point can be specified by configuration.
[0045] The embodiment of the present invention performs aggregation based on aggregation points and abnormal events to obtain aggregation results. It can be understood that the aggregation point is one or more of the many abnormal events. Since the abnormal events are continuously obtained within a preset time period, the time at the aggregation point is in the middle of the preset time period. It can be understood that the aggregation based on the aggregation point can include multiple abnormal events before and after the acquisition time of the aggregation point. These abnormal events before and after the time can be at least one of alarms, key performance indicator abnormalities and operation logs. Therefore, in the embodiment of the present invention, when an alarm or key performance indicator abnormality is triggered by a certain operation before the aggregation point, the data before and after the aggregation point can be aggregated. The obtained aggregation result can be used to clarify the root cause of the fault and realize two-way aggregation, which can improve the aggregation capability of the data source and improve the fault operation and maintenance level.
[0046] It should be noted that the preset time period in the embodiment of the present invention can be set according to actual operation and maintenance needs. For example, the preset time period can be 20 minutes, 1 hour, 4 hours or longer. Starting from the start time of the preset time period, the embodiment of the present invention can start to obtain abnormal events as data sources, including obtaining at least one of alarms, key performance indicator abnormalities and operation logs, to achieve the acquisition of multiple data sources. In the embodiment of the present invention, the data source is obtained within the preset time period. By setting the length of the preset time period, the aggregation time can be clearly defined in the time dimension.
[0047] It should be noted that the target location in the embodiment of the present invention is a location in the spatial dimension. For example, the target location can be a network element, a computer room or a link. Through the final aggregation result, the root cause of the failure of the network element, computer room or link can be obtained through aggregation analysis.
[0048] Reference Figure 2 As shown, in one embodiment, the above step S101 may also include but not limited to step S201 and step S202.
[0049] Step S201 : establishing multiple time buckets in the target buffer area according to the total duration of the preset time period, wherein the time buckets are composed of timestamp intervals, the duration of each time bucket is the same, and the time of two adjacent time buckets is continuous.
[0050] Step S202 : continuously obtain multiple abnormal events at the target location, and cache them in a time bucket of a corresponding time according to the acquisition time of each abnormal event.
[0051] In one embodiment, the embodiment of the present invention implements a caching method for abnormal events by setting time buckets. Specifically, the embodiment of the present invention establishes multiple time buckets in the target cache area according to the total duration of the preset time period. The time buckets are composed of timestamp intervals. The duration of each time bucket is the same and the time of two adjacent time buckets is continuous. When multiple abnormal events of the target location are continuously obtained, they are cached in the time bucket of the corresponding time according to the acquisition time of each abnormal event to implement data caching. The target cache area is a cache area corresponding to the target location cache. A target location can correspond to multiple cache areas, or to a one-to-one corresponding cache area. The embodiment of the present invention uses a bidirectional time dimension for aggregation. Each cache area caches abnormal events according to the timestamp and a certain time interval as a time bucket. Therefore, after caching the abnormal event, it is not necessary to aggregate immediately. It is necessary to wait for a certain time until the time bucket is cached before preparing for aggregation.
[0052] It should be noted that in the target cache area, the duration of each time bucket is the same and the time of two adjacent time buckets is continuous. For example, when a target cache area obtains an abnormal event within a preset time period of 20 minutes, the duration of each time bucket can be set to 5 minutes, so 4 consecutive time buckets can be obtained, among which the time of the first time bucket is cached from the 0th minute to the 5th minute, the second time bucket is cached from the 5th minute to the 10th minute, the third time bucket is cached from the 10th minute to the 15th minute, and the fourth time bucket is cached from the 15th minute to the 20th minute. Longer preset time periods can be deduced in this way. The duration of each time bucket can be set according to actual operation and maintenance needs, and no specific restrictions are made here.
[0053] It is understandable that, in the embodiment of the present invention, the stopping of caching abnormal events can be controlled by setting a condition for no longer caching the time bucket in the target cache area.
[0054] Reference Figure 3 As shown, in one embodiment, the above step S202 may also include but not limited to steps S301 to S303.
[0055] Step S301: Obtain a convergence condition for stopping caching abnormal events within a preset time period.
[0056] Step S302 : continuously acquiring multiple abnormal events at the target location starting from the start time of the preset time period, and caching the abnormal events in a time bucket corresponding to the acquisition time of each abnormal event.
[0057] Step S303: When the cached abnormal events meet the convergence condition, stop caching the abnormal events.
[0058] In one embodiment, to address the problem that the time for aggregation in the time dimension is currently difficult to determine, an embodiment of the present invention caches abnormal events at a target location by setting a time bucket mode, and controls the time point when the time bucket stops caching by setting a convergence condition for the time bucket. Specifically, an embodiment of the present invention obtains a convergence condition for stopping caching abnormal events within a preset time period. When caching abnormal events, an embodiment of the present invention continuously obtains multiple abnormal events at the target location starting from the start time of the preset time period, and sequentially caches them in the time bucket of the corresponding time according to the acquisition time of each abnormal event. When the cached abnormal events meet the convergence condition, the abnormal events are stopped from being cached. When the convergence condition is met, it means that the target cache area has been cached, that is, the cache area is closed and no longer receives other abnormal events. The time bucket is encapsulated and prepared for aggregation. The target cache area can also be cleared after the convergence condition is met to wait for the subsequent abnormal event cache. In the embodiment of the present invention, by setting the convergence condition to control when to stop caching abnormal events, data collection time can be avoided, and the efficiency of aggregation in the time dimension can be improved, thereby improving the aggregation capability.
[0059] It should be noted that, from the perspective of the aggregation point, in the embodiment of the present invention, when the target cache area is closed to stop caching abnormal events based on meeting the convergence conditions, the currently cached data already includes abnormal events in both directions before and after the aggregation point time dimension, that is, the abnormal events before and after the aggregation point as the center have all entered the target cache area, thereby improving the data aggregation capability, thereby enabling bidirectional aggregation.
[0060] Reference Figure 4 As shown, in one embodiment, the convergence condition may include but is not limited to at least one of steps S401 to S403.
[0061] Step S401: The time of obtaining the abnormal event exceeds the end time of the preset time period.
[0062] Step S402: The decreasing rate of the number of cached abnormal events between a plurality of consecutive time buckets is less than a preset target decreasing rate.
[0063] Step S403: The number of abnormal events in the time bucket is less than a preset minimum threshold of the number of events in the bucket.
[0064] In one embodiment, there may be multiple convergence conditions in the embodiment of the present invention, and it can be determined in the time dimension when the data source collection is completed to improve the aggregation efficiency and aggregation capability in the time dimension. Specifically, the convergence conditions may include at least one of steps S401 to S403. It can be understood that when one of the convergence conditions in the above steps is met, it can be determined that the data source collection is completed, and thus the abnormal event is stopped from being cached.
[0065] It should be noted that determining whether the time for obtaining an abnormal event exceeds the end time of a preset time period is one of the convergence conditions. Specifically, the preset time period has a start time and an end time. When the time for obtaining an abnormal event exceeds the end time of the preset time period, it indicates that the cache time expires, that is, the time interval from the time of the last aggregation point in this cache to the end time. This time interval is the maximum value of the preset time period, which limits excessive waiting time. The convergence condition of an embodiment of the present invention is used to force the cache to end. In one embodiment, the maximum value of the aggregation time interval, that is, the preset time period, is set to 60 minutes. After this time, subsequent messages will no longer be waited for.
[0066] It should be noted that judging whether the decreasing rate of the number of cached abnormal events between multiple consecutive time buckets is less than the preset target decreasing rate is one of the convergence conditions. Specifically, when the number of abnormal events in three consecutive time buckets is set to decrease at a certain rate, it is judged that it is lower than the decreasing ratio of the number of events, that is, lower than the target decreasing rate, that is, the number of the latter bucket is lower than the number of the former bucket by a certain percentage. When it is lower than the target decreasing rate, it is judged that there is no need to cache abnormal events anymore. For example, when the target decreasing rate is 25%, the value of the target decreasing rate can be set according to actual needs. The convergence condition in the embodiment of the present invention ends the caching after the marginal effect decreases, such as Figure 5 As shown in the figure, each time bucket in the target cache area caches abnormal events such as user login log, configuration routing log, restart routing log, communication anomaly, network disconnection alarm, key performance indicator abnormality (KPI abnormality in the figure), business anomaly, business restart alarm, etc. Among them, communication anomaly, network disconnection alarm, key performance indicator abnormality and business restart alarm are aggregation points. Figure 5In the example, the number of abnormal events in time bucket 4 is one third of that in time bucket 3, which is higher than the set target deceleration rate (25%). Therefore, the cache cannot be ended at present and abnormal events continue to be received.
[0067] It should be noted that determining whether the number of abnormal events in a time bucket is less than a preset minimum threshold of the number of events in the bucket is one of the convergence conditions. Specifically, the number of abnormal events in a time bucket is less than the minimum threshold of the number of events in the bucket, that is, the minimum value of the number of events in the bucket. The minimum threshold of the number of events in the bucket can be set according to actual needs. When it is lower than the minimum threshold of the number of events in the bucket, it is determined that there is no need to cache abnormal events. The embodiment of the present invention realizes that the cache ends after the events converge within a certain period of time, and also refers to Figure 5 As shown, time bucket 4 receives only one event, which is less than the minimum threshold of the number of events in the bucket (assuming it is 2). That is, no more abnormal events need to be received, and the abnormal events are stopped from being cached. That is, time bucket 5 does not need to be received anymore, and the collection of data sources is finally completed.
[0068] In one embodiment, there are multiple target locations, refer to Figure 6 As shown, the above step S201 may also include but not limited to step S501 and step S302.
[0069] Step S501: Obtain the preset time periods corresponding to the respective target locations.
[0070] Step S502 : establishing target buffer areas corresponding to respective target locations, and establishing a plurality of time buckets in the corresponding target buffer areas according to the total duration of each preset time period.
[0071] In one embodiment, when there are multiple target locations, the embodiment of the present invention caches the data source according to different target locations respectively. Specifically, each different location can be set with a preset time period required for its own caching. The embodiment of the present invention obtains the preset time period corresponding to each target location respectively, and caches the data of each target location, and establishes target cache areas corresponding to each target location respectively. In one embodiment, the target cache areas correspond to the target locations one-to-one, and each target location has a corresponding target cache area. According to the total length of each preset time period, multiple time buckets are established in the corresponding target cache area, so that the abnormal time of each target location is cached in the corresponding time bucket.
[0072] The target location in the embodiment of the present invention is a location in the spatial dimension. For example, the target location can be a network element, a link or a computer room. Multiple target locations may include multiple network elements, links and computer rooms. By collecting abnormal events at different target locations, aggregation results of different target locations can be obtained so as to perform aggregation analysis on each target location. It can be understood that, in the embodiment of the present invention, aggregation results of each target location can be obtained, and the root cause of the fault of each target location itself can be analyzed through the aggregation results. The overall aggregation results of multiple target locations can also be obtained, and the root cause of the fault in multiple target locations can be analyzed through the aggregation results, thereby improving the aggregation capability and improving the level of fault operation and maintenance.
[0073] Reference Figure 7 As shown, in one embodiment, the above step S102 may also include but not be limited to step S601 and step S602.
[0074] Step S601, obtaining the filtering conditions of the aggregation points.
[0075] Step S601: Determine, among multiple abnormal events, abnormal events that meet the screening conditions as aggregation points.
[0076] In one embodiment, the embodiment of the present invention can obtain screening conditions of aggregation points, determine aggregation points from abnormal events, and determine which are major alarms or major key performance indicator anomalies from abnormal events through the screening conditions. A major alarm can be any one of the alarms in the abnormal event, and a major key performance indicator anomaly can be any one of the key performance indicator anomalies in the abnormal event. For example, aggregation points such as base station decommissioning and cell decommissioning are the real center of operation and maintenance. Aggregation centered on aggregation points of major alarms or major key performance indicator anomalies is suitable for actual operation and maintenance needs. Otherwise, the aggregation of a large number of ordinary alarms will waste a lot of time, resulting in a low level of operation and maintenance. The specific alarm type and key performance indicator anomaly type of the aggregation point can be specified through configuration.
[0077] It is understandable that in the embodiments of the present invention, the filtering conditions can be customized according to the actual operation and maintenance needs to determine the major alarms or major key performance indicator anomalies. The aggregation point is one or more of the many abnormal events. Since the abnormal events are continuously obtained within the preset time period, the time at the aggregation point is in the middle of the preset time period. It is understandable that the aggregation of the aggregation point can include multiple abnormal events before and after the acquisition time of the aggregation point. These abnormal events before and after the time can be at least one of alarms, key performance indicator anomalies and operation logs. Therefore, in the embodiments of the present invention, when an alarm or key performance indicator anomaly is triggered by a certain operation before the aggregation point, the data before and after the aggregation point can be aggregated. The obtained aggregation result can be used to clarify the root cause of the fault and realize two-way aggregation, which can improve the aggregation capability of the data source and improve the level of fault operation and maintenance.
[0078] It should be noted that in related technologies, backward aggregation is not performed using a certain operation log as the starting point. This is because there are too many logs and they are too frequent. In addition, most operation logs are only for recording and do not indicate abnormalities. Therefore, for operation logs, since it is not convenient to clearly identify abnormalities, they cannot be perceived immediately. Often, after an alarm or key performance indicator abnormality is detected, related operation logs are not searched for. For example, in the case of memory leaks, related operation logs are not searched for after a memory leak or a leakage trend has been discovered. This leads to poor aggregation capabilities.
[0079] Reference Figure 8 As shown, in one embodiment, the above step S103 may also include but is not limited to steps S701 to S703.
[0080] Step S701: determining a first target event and a second target event in an abnormal event, wherein the first target event is characterized as a noise event of an aggregation point, and the second target event is characterized as a correlation event of an aggregation point.
[0081] Step S702: clear the first target event and retain the second target event.
[0082] Step S703: Aggregate according to the aggregation point and the second target event to obtain an aggregation result.
[0083] In one embodiment, the embodiment of the present invention can denoise abnormal events, remove unnecessary events, and retain abnormal events related to the aggregation point, so as to improve the aggregation capability. Specifically, in the embodiment of the present invention, a first abnormal event and a second abnormal event can be determined in the abnormal event, the first target event is characterized as a noise event of the aggregation point, and the second target event is characterized as an associated event of the aggregation point. As a noise event, if it is aggregated with the aggregation point, the data volume of the final aggregation result will be too large, and there will be many abnormal events that are useless for the root cause analysis of the fault. Therefore, in the embodiment of the present invention, the noise event characterized as the aggregation point can be determined, that is, the first target event is determined, and the associated event characterized as the aggregation point, that is, the second target event characterization is determined, the first target event is cleared and the second target event is retained. Finally, aggregation can be performed according to the aggregation point and the second target event to obtain an aggregation result, which can improve the aggregation capability of the embodiment of the present invention and improve the fault operation and maintenance level.
[0084] It should be noted that, since the abnormal events in the embodiment of the present invention include operation logs, a large number of operation logs will exist during the actual operation and maintenance process. In the embodiment of the present invention, through two-way aggregation, the abnormal events before and after the aggregation point can be obtained to obtain the aggregation result, that is, the operation logs before and after the aggregation point can be obtained to obtain the aggregation result. Finally, the root cause of the fault can be found according to the aggregation result to find the operation log that causes the abnormality of the aggregation point. In order to solve the problem of too many operation logs and a large number of them are irrelevant to the aggregation point, the embodiment of the present invention clarifies the first target event and the second target event in the abnormal event, clears the first target event and retains the second target event, and finally ensures the aggregation ability and efficiency of the embodiment of the present invention.
[0085] by Figure 5 Taking the abnormal events collected in as an example, when the abnormal event of communication abnormality is used as the aggregation point, it can be based on Figure 9 As shown in the figure, communication anomalies are aggregated in both directions. Forward, abnormal events such as user login logs, routing configuration logs, and routing restart logs can be aggregated. Backward, abnormal events such as network failure alarms, key performance indicator abnormalities, and business abnormalities can be aggregated. When the abnormal event of network failure alarm is used as the aggregation point, Figure 10 As shown in the figure, communication anomalies are aggregated bidirectionally. Forward aggregation can aggregate abnormal events such as user login logs, routing configuration logs, routing restart logs, and communication anomalies. Backward aggregation can aggregate abnormal events such as key performance indicator anomalies, business anomalies, and business restart alarms.
[0086] Reference Figure 11 As shown, in one embodiment, the above step S103 may also include but not be limited to step S801 and step S802.
[0087] Step S801: Aggregate the aggregation points and abnormal events to obtain an aggregation package.
[0088] Step S802: identify the root cause of the aggregation packet and obtain the root cause identification result of the aggregation point by combining the abnormal events corresponding to each aggregation point.
[0089] In one embodiment, the present invention can perform root cause identification to obtain a root cause identification result. In the present invention, the aggregation point and the abnormal event are aggregated to obtain an aggregation package, the root cause identification is performed on the aggregation package, and the root cause identification result of the aggregation point is obtained in combination with the abnormal event corresponding to each aggregation point. In another embodiment, the present invention performs aggregation based on the aggregation point and the second target event to obtain an aggregation package. By aggregating the second target event, an aggregation package with higher aggregation efficiency is obtained. The root cause of these useful abnormal events is identified, and the second target event in the aggregation point and the knowledge base and other technologies can be used to analyze which abnormal event is the root cause event, thereby improving the level of fault operation and maintenance.
[0090] Reference Figure 12 As shown, in one embodiment, the above step S701 may also include but not limited to step S901 and step S902.
[0091] In step S901, the abnormal events are initialized to obtain initial data, and the initial data is input into a preset two-way aggregation model for probability calculation to obtain noise probability values of each abnormal event and the corresponding aggregation point.
[0092] Step S902: determining a first target event and a second target event in the abnormal event according to the noise probability value.
[0093] In one embodiment, the present invention determines the first target event and the second target event in the abnormal event by obtaining a preset bidirectional aggregation model. The bidirectional aggregation model is a data processing model obtained by training a neural network model. Specifically, the present invention obtains initial data by initializing the abnormal event, and inputs the initial data into the preset bidirectional aggregation model for probability calculation to obtain noise probability values of each abnormal event and the corresponding aggregation point. The input of the bidirectional aggregation model needs to match the corresponding initial data so that the bidirectional aggregation model can perform data processing. The noise probability value can represent the probability that the abnormal event is a noise event of the corresponding aggregation point. The probability represented by the noise probability value can determine whether the abnormal event is a noise event of the corresponding aggregation point, thereby determining the first target event and the second target event.
[0094] It can be understood that there can be multiple aggregation points in the embodiment of the present invention. When there are multiple aggregation points, each abnormal event can be probability calculated through the bidirectional aggregation model to obtain the noise probability value for each aggregation point. This is because some abnormal events have low probabilities for some aggregation points, but high probabilities for other aggregation points. Therefore, calculating the probability of each abnormal event with each aggregation point can avoid removing some high-probability abnormal events, which helps to aggregate all aggregation points.
[0095] Reference Figure 13 As shown, in one embodiment, the above step S902 may also include but not be limited to steps S1001 to S1003.
[0096] Step S1001, obtaining the first probability threshold and the second probability threshold of each aggregation point.
[0097] Step S1002 : determining abnormal events corresponding to noise probability values lower than all first probability thresholds as first target events.
[0098] Step S1003 : determining an abnormal event corresponding to a noise probability value higher than any second probability threshold as a second target event.
[0099] In one embodiment, the embodiment of the present invention screens abnormal events by setting a low probability threshold and a high probability threshold. The embodiment of the present invention can obtain the first probability threshold and the second probability threshold of each aggregation point. The first probability threshold is a low probability threshold, which is used to screen out the first target event among the abnormal events. Therefore, the abnormal event corresponding to the noise probability value lower than all the first probability thresholds is determined as the first target event, and the first target event is a low probability event. The second probability threshold is a high probability threshold, which is used to screen out the second target event. The abnormal event corresponding to the noise probability value higher than any second probability threshold is determined as the second target event, and the second target event is a high probability event.
[0100] It should be noted that, in the embodiment of the present invention, the mark of events below the low probability threshold as the first probability threshold of the low probability threshold can be set in the interface or configuration file and configured according to actual operation and maintenance needs. Abnormal events above the high probability threshold are placed in a high probability list. The second probability threshold as the high probability threshold can be set in the interface or configuration file and configured according to actual operation and maintenance needs. Its key value is the abnormal point and its value is a list. These abnormal events above the high probability threshold are saved in the list. Abnormal events below the low probability threshold are not temporarily excluded when analyzing each aggregation point, because some abnormal events have low probability for some aggregation points but high probability for other aggregation points. Therefore, in the embodiment of the present invention, when judging which abnormal events are the first target events, it is required that the noise probability value of the abnormal event is lower than the first probability threshold of all aggregation points to determine it as the first target event. When judging the second target event, the noise probability value of the abnormal event only needs to be higher than the second probability threshold of any aggregation point to be judged as the second target event.
[0101] In one embodiment, the first probability threshold is used when a certain aggregation point is used to predict the context-related events of the current aggregation point. If the probability of certain abnormal events is very low and the probability for all aggregation points is lower than the first probability threshold, such as 10%, then they can be treated as noise for denoising. The second probability threshold is used when a certain aggregation point is used to predict the context-related events of the current aggregation point. If the probability of certain related abnormal events is higher than the second probability threshold, such as 75%, then it can be considered that the correlation is very strong and can assist in subsequent root cause analysis.
[0102] Reference Figure 14 As shown, in one embodiment, the above step S901 may also include but is not limited to steps S1101 to S1103.
[0103] Step S1101 , performing one-hot encoding on the abnormal event to obtain initialized initial vector data.
[0104] Step S1102 , obtaining a preset bidirectional aggregation model, wherein the bidirectional aggregation model is obtained through unsupervised training based on sample abnormal events, sample target events characterized as noise events, and sample aggregation points in the acquired samples.
[0105] In step S1103 , the initial vector data is input into a preset bidirectional aggregation model to perform probability calculation, and noise probability values of each abnormal event and the corresponding aggregation point are obtained respectively.
[0106] In one embodiment, the abnormal events are initialized by vector conversion before being input into a preset bidirectional aggregation model for processing to obtain the required noise probability value. Specifically, the bidirectional aggregation model can be pre-established based on the data in the sample. In the embodiment of the present invention, since there are a large number of abnormal events in the aggregation, not all of them are useful for root cause analysis. Some events are noise events for the aggregation point, such as some daily operation logs and flash alarms that coincide with the abnormal points in a certain time window. Their existence interferes with the aggregation analysis. Therefore, through artificial intelligence (AI) training and filtering with a certain probability, the aggregation analysis can be made more accurate.
[0107] Just like the continuous skip-gram model (Skip-gram) of word vectorization (Word2vec) in natural language processing (NLP), which uses a central word to predict the probability of context words, the embodiments of the present invention use the same principle. After vectorizing abnormal events, the bidirectional aggregation model is used to obtain the probability of abnormal events in the time period before and after the aggregation point, and abnormal event denoising is performed using a probability threshold.
[0108] In an embodiment of the present invention, a trainer can be set to load historical data. The historical data may include sample abnormal events in the sample, sample target events characterized as noise events, and sample aggregation points. In the training stage, these data in the sample are one-hot encoded. Through unsupervised training, when the context probability is maximized and the loss function is minimized, the abnormal events can be vectorized, and a two-way aggregation model is obtained through training, which will be needed for subsequent applications.
[0109] In the embodiment of the present invention, when performing probability calculation, the trained bidirectional aggregation model is first loaded, the abnormal events are one-hot encoded, and the initialized initial vector data is obtained. The initial vector data is input into the preset bidirectional aggregation model for probability calculation, and the noise probability values of each abnormal event and the corresponding aggregation point are obtained respectively. The input has been expressed by initialization vector through one-hot encoding before the bidirectional aggregation model is input. Therefore, the bidirectional model obtained through training and release can be directly used to perform probability calculation on the event, and vector probability calculation can be performed on the abnormal events to obtain the noise probability values of each abnormal event and the corresponding aggregation point.
[0110] It can be understood that the implementation of the present invention can use but is not limited to the Skip-gram model of Word2vec for training, including the construction of a neural network to obtain the required two-way aggregation model. First, historical data in the sample is obtained, or it is used as a corpus for one-hot encoding according to a fault handling manual, etc., and then the Skip-gram model is used. Its loss function is to minimize all probabilities. At this time, the correlation between different abnormal events such as alarms and operation logs is trained to obtain an intermediate hidden layer. This is the final model required. The steps of training to obtain a two-way aggregation model are not specifically described in the embodiments of the present invention.
[0111] Reference Figure 15 As shown, in one embodiment, the above step S901 may also include but not be limited to step S1201 and step S1202.
[0112] Step S1201, sort multiple aggregation points by time and store them in an aggregation point list.
[0113] In step S1202, the initial data is input into a preset bidirectional aggregation model, and the probability of abnormal events is calculated according to each aggregation point in the aggregation point list to obtain the noise probability value of each abnormal event and the corresponding aggregation point.
[0114] In one embodiment, the present invention stores aggregation points by establishing an aggregation point list. After a target buffer stops caching abnormal events, the target buffer is closed and no longer receives further abnormal events. The time bucket is encapsulated into an initial packet in preparation for aggregation, and then the buffer is cleared to prepare for subsequent event caching. It should be emphasized that from the perspective of the aggregation point, the difference from the conventional one-way backward aggregation is that when the target buffer is closed, the present invention already includes both forward and backward events. That is, the forward and backward events centered around the aggregation point have entered the buffer. The bidirectional aggregation in the present invention is based on the aggregation point. Therefore, the aggregation points of the target buffer can be collected. If there are no aggregation points, the target buffer is directly recycled for the next caching. If a target buffer or even a bucket within the buffer may have multiple aggregation points, the aggregation points are first collected, sorted by time, and stored in the aggregation point list. The earliest and latest abnormal events in the buffer are then displayed. In addition, the location of the target buffer can also be displayed, such as the network element, computer room, or link.
[0115] It should be noted that, when performing probability calculation, the embodiment of the present invention first obtains a list of aggregation points, and then obtains the noise probability rate of other abnormal events in the data through a bidirectional model for the aggregation points in the list. Abnormal events below the low probability threshold are marked, and abnormal events above the high probability threshold are placed in a high probability list, whose key value is the abnormal point and the value is the list. These abnormal events above the high probability threshold are saved in the list. Finally, after denoising, the aggregation point can be attached to the high probability list to form an aggregation package and sent to the root cause identification.
[0116] In addition, the abnormal event processing method in the embodiment of the present invention can be applied to an abnormal event processing device, referred to as a processing device, which may include:
[0117] Buffer: Receives external abnormal events into the buffer area, assembles them into initial packets and sends them to the packager;
[0118] Packager: receives the initial packet, packages it into an encoded packet and sends it to the aggregator;
[0119] Aggregator: Receives encoded packets, performs context probability training and prediction with the aggregation point as the center, removes noise, obtains an aggregated packet, and sends it to the root cause analysis;
[0120] The trainer, the cache, the packager, the aggregator, and the trainer are in communication with each other, and when the abnormal event handling method in the above embodiment is executed by the processing device, the following four steps may be included:
[0121] Step 1: The trainer trains a bidirectional aggregation model to complete event vectorization.
[0122] Specifically, in aggregation, there may be a large number of abnormal events, but not all of them are useful for root cause analysis. Some events are noise events for the aggregation point, such as some daily operation logs and flash alarms that happen to coincide with abnormal points in a certain time window. Their existence interferes with the aggregation analysis. Therefore, through AI training and filtering with a certain probability, the aggregation analysis can be made more accurate.
[0123] Just like the Skip-gram model of Word2vec in NLP, which uses a central word to predict the probability of context words, the embodiments of the present invention use the same principle. After vectorizing abnormal events, a bidirectional aggregation model is used to obtain the probability of abnormal events in the time period before and after the aggregation point, and abnormal events are denoised using a probability threshold.
[0124] The trainer loads historical alarms, logs, key performance indicator anomalies, and troubleshooting manuals as a corpus and performs one-hot encoding. Then, through unsupervised training, when the context probability is maximized and the loss function is minimized, the abnormal events can be vectorized to obtain a bidirectional aggregation model, which is then released.
[0125] Step 2: The cache receives and caches streaming exception events.
[0126] Abnormal events are stream input, so abnormal events for a certain period of time need to be cached.
[0127] The cache sets different cache areas according to different spatial dimensions, that is, different locations. A cache area can only cache abnormal events of the same spatial dimension. Each cache area caches abnormal events according to the timestamp and a certain time interval as a time bucket. If the event is an aggregation point, it is marked.
[0128] After the aggregation point is found, aggregation does not need to be performed immediately. You need to wait for a certain period of time until the time bucket is fully cached before preparing for aggregation. A time bucket is a time interval, such as five minutes, in which abnormal events of the five minutes are cached. The next time bucket caches abnormal events of the next time interval, such as five minutes.
[0129] As reference Figure 5 ,After the streaming abnormal events enter, a cache area is placed at the same location. In the figure, a batch of abnormal events is cached in a time bucket every 5 minutes. Different time buckets may have different sizes.
[0130] A cache area consists of one or more time buckets. The key is when it ends, that is, when the cache is completed and can be aggregated. This embodiment of the present invention uses three dimensions as convergence conditions to complete the cache of the last time bucket:
[0131] The cache expiration time is the time interval from the last aggregation point of the cache to the expiration time. This time interval is the maximum time interval, which limits excessive waiting time. This approach is forced termination;
[0132] The number of abnormal events in three consecutive time buckets decreases at a certain rate, which is lower than the rate of decrease in the number of events. That is, the number of the latter bucket is lower than the number of the former bucket by a certain percentage, such as 25%. This number can be set. This approach ends after the marginal effect decreases. Figure 5 As shown in the figure, the number of events in time bucket 4 is one third of that in time bucket 3. Therefore, the process cannot be terminated yet and continues to receive abnormal events.
[0133] The number of abnormal events in the time bucket is less than the minimum threshold of the number of events in the bucket, that is, the minimum number of bucket events. This value can be set. This approach ends after the events converge within a certain period of time. Figure 5 As shown, time bucket 4 receives only one event, which is less than the minimum threshold of the number of events in the bucket (assuming it is 2), that is, no more events need to be received, that is, time bucket 5 does not need to receive any more abnormal data.
[0134] When any of the above three conditions is met, the cache area is completed, that is, the cache area is closed and no longer receives other abnormal events. The time bucket is encapsulated into an initial packet and sent to the packager for aggregation. Then the cache area is cleared to wait for the subsequent event cache.
[0135] It should be emphasized that, from the perspective of the aggregation point, the difference from the conventional one-way backward aggregation is that when the cache area is closed in the embodiment of the present invention, the forward and backward bidirectional events are already included, that is, the forward and backward events centered on the aggregation point have all entered the cache area.
[0136] Step 3: The packager packages the initial package.
[0137] After the caching is completed, the packager packages the initial packet into an aggregated packet. The bidirectional aggregation of the present invention is based on the aggregation point for bidirectional aggregation. Therefore, the packager first collects the aggregation points of the current cache area. If there is no aggregation point, the current cache area is directly recycled for the next caching. If there may be multiple aggregation points in a cache area or even a bucket in the cache area, the aggregation points are collected first and sorted by time. Then the earliest and latest abnormal events in the cache area are given, and the location of the cache area is given, such as the network element, computer room, or link. The abnormal events in the cache area are uniquely encoded. After completing the above operations, the packager completes the packaging and obtains the encoded package. The packager sends the encoded package to the aggregator for aggregation.
[0138] Step 4: Aggregation.
[0139] The aggregator loads the trained bidirectional aggregation model and performs forward and backward double denoising on the aggregation points in the aggregation package to complete the aggregation.
[0140] At the end of the third step, the aggregator receives the encoded package sent by the packager. Since it has been one-hot encoded, it can directly use the bidirectional model obtained through training and release to vectorize the events and calculate the vector probability of abnormal events in the package.
[0141] The aggregator first obtains a list of aggregation points, and then uses a bidirectional model to obtain the probability of other abnormal events in this package for the aggregation points in the list. It marks those below the low probability threshold (which can be set in the interface or configuration file), and puts abnormal events (which may also be aggregation points) above the high probability threshold in a high probability list, whose key value is the abnormal point and the value is a list. The list saves these abnormal events above the high probability threshold.
[0142] Note that abnormal events below the low probability threshold are not excluded temporarily when analyzing each aggregation point, because some abnormal events have low probability for some aggregation points but high probability for other aggregation points. After all aggregation points in the code packet are analyzed, all abnormal events marked with low probability are checked. If their probability for all aggregation points is lower than the minimum probability threshold, they are cleared.
[0143] After analyzing all aggregation points, the aggregator can also conduct a second check on the abnormal events marked below the low probability threshold to see whether they are below the low probability for each aggregation point. If not, they are retained; otherwise, denoising and clearing are performed. After denoising, the aggregator appends the high probability list to the aggregation points in the encoding package, combines them into an aggregation package, and sends it to the root cause identification. The aggregation is completed.
[0144] Figure 16 The electronic device 100 provided by an embodiment of the present invention is shown. The electronic device 100 includes: a processor 110, a memory 120, and a computer program stored in the memory 120 and executable on the processor 110. When the computer program is executed, it is used to execute the above-mentioned abnormal event processing method.
[0145] The processor 110 and the memory 120 may be connected via a bus or other means.
[0146] Memory 120, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer executable programs, such as the abnormal event handling method described in the embodiments of the present invention. Processor 110 implements the abnormal event handling method described above by executing the non-transitory software programs and instructions stored in memory 120.
[0147] The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store and execute the above-mentioned abnormal event handling method. In addition, the memory 120 may include a high-speed random access memory 120, and may also include a non-volatile memory 120, such as at least one storage device memory device, a flash memory device or other non-volatile solid-state memory device. In some embodiments, the memory 120 may optionally include a memory 120 remotely located relative to the processor 110, and these remote memories 120 may be connected to the electronic device 100 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0148] The non-transient software programs and instructions required to implement the above-mentioned abnormal event handling method are stored in the memory 120. When executed by one or more processors 110, the above-mentioned abnormal event handling method is executed, for example, Figure 1Steps S101 to S103 of the method, Figure 2 Steps S201 to S202 of the method, Figure 3 Steps S301 to S303 of the method, Figure 4 Steps S401 to S403 of the method, Figure 6 Steps S501 to S502 of the method, Figure 7 Steps S601 to S602 of the method, Figure 8 Steps S701 to S703 of the method, Figure 11 Steps S801 to S802 of the method, Figure 12 Steps S901 to S902 of the method, Figure 13 Steps S1001 to S1003 of the method, Figure 14 Steps S1101 to S1103 of the method, Figure 15 Method steps S1201 to S1202.
[0149] An embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the above-mentioned abnormal event processing method.
[0150] In one embodiment, the computer readable storage medium stores computer executable instructions, which are executed by one or more control processors, for example, Figure 1 Steps S101 to S103 of the method, Figure 2 Steps S201 to S202 of the method, Figure 3 Steps S301 to S303 of the method, Figure 4 Steps S401 to S403 of the method, Figure 6 Steps S501 to S502 of the method, Figure 7 Steps S601 to S602 of the method, Figure 8 Steps S701 to S703 of the method, Figure 11 Steps S801 to S802 of the method, Figure 12 Steps S901 to S902 of the method, Figure 13 Steps S1001 to S1003 of the method, Figure 14 Steps S1101 to S1103 of the method, Figure 15 Method steps S1201 to S1202.
[0151] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0152] Those skilled in the art will appreciate that all or some of the steps and systems in the method disclosed above can be implemented as software, firmware, hardware, and appropriate combinations thereof. Some physical components or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, and the computer-readable medium can include computer storage media (or non-transitory media) and communication media (or temporary media). As known to those skilled in the art, computer storage media include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technology, CD-ROM, digital versatile disks (DVD), or other optical disk storage, magnetic cassettes, magnetic tapes, storage device storage, or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, as is well known to those skilled in the art, communication media typically includes computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0153] It should also be understood that the various implementations provided in the embodiments of the present invention can be arbitrarily combined to achieve different technical effects.
[0154] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above implementation. Those skilled in the art can also make various equivalent modifications or substitutions under the shared conditions that do not violate the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present invention.
Claims
1. A method for handling abnormal events, characterized in that: The method comprises: Establishing multiple time buckets in the target buffer area according to the total duration of the preset time period, wherein the time buckets are composed of time stamp intervals, the duration of each time bucket is the same, and the time of two adjacent time buckets is continuous; Obtaining a convergence condition for stopping caching abnormal events within the preset time period, wherein the convergence condition includes at least one of the following: the time when the abnormal event is acquired exceeds the end time of the preset time period; the decreasing rate of the number of abnormal events cached between multiple consecutive time buckets is less than a preset target decreasing rate; the number of abnormal events in the time bucket is less than a preset minimum threshold of the number of events in the bucket; Continuously acquiring multiple abnormal events at the target location starting from the start time of the preset time period, and caching the abnormal events in the time bucket of the corresponding time according to the acquisition time of each abnormal event; When the cached abnormal event meets the convergence condition, stop caching the abnormal event, wherein the abnormal event includes at least one of an alarm, a key performance indicator abnormality, and an operation log; determining an aggregation point in said anomalous event; Aggregation is performed based on the aggregation point and the abnormal event to obtain an aggregation result.
2. The abnormal event handling method according to claim 1, characterized in that: There are multiple target locations, and establishing multiple time buckets in the target buffer area according to the total length of the preset time period includes: respectively obtaining a preset time period corresponding to each of the target locations; A target buffer area corresponding to each of the target locations is established respectively, and a plurality of time buckets are established in the corresponding target buffer area according to the total duration of each of the preset time periods.
3. The abnormal event handling method according to claim 1, characterized in that: Determining the aggregation point in the abnormal event includes: Get the filtering conditions of the aggregation points; The abnormal event that meets the screening condition is determined as the aggregation point among the multiple abnormal events.
4. The abnormal event handling method according to claim 1, characterized in that: The aggregating according to the aggregation point and the abnormal event to obtain an aggregation result includes: Determining a first target event and a second target event from the abnormal event, wherein the first target event is characterized as a noise event at the aggregation point, and the second target event is characterized as a correlation event at the aggregation point; clearing the first target event and retaining the second target event; Aggregation is performed based on the aggregation point and the second target event to obtain an aggregation result.
5. The abnormal event handling method according to claim 1 or 4, characterized in that: The aggregating according to the aggregation point and the abnormal event to obtain an aggregation result includes: Aggregating the aggregation point and the abnormal event to obtain an aggregation package; Root cause identification is performed on the aggregation packet, and the root cause identification result of the aggregation point is obtained by combining the abnormal events corresponding to each aggregation point.
6. The abnormal event handling method according to claim 4, characterized in that: Determining the first target event and the second target event in the abnormal event includes: Initializing the abnormal events to obtain initial data, and inputting the initial data into a preset bidirectional aggregation model to perform probability calculation, thereby obtaining noise probability values of each abnormal event and the corresponding aggregation point; A first target event and a second target event in the abnormal events are determined according to the noise probability value.
7. The abnormal event handling method according to claim 6, characterized in that: Determining the first target event and the second target event in the abnormal event according to the noise probability value includes: Obtain a first probability threshold and a second probability threshold for each of the aggregation points; Determine the abnormal event corresponding to the noise probability value lower than all the first probability thresholds as a first target event; The abnormal event corresponding to the noise probability value higher than any second probability threshold is determined as a second target event.
8. The abnormal event handling method according to claim 6, characterized in that: Initializing the abnormal events to obtain initial data, and inputting the initial data into a preset bidirectional aggregation model for probability calculation to obtain noise probability values of each abnormal event and the corresponding aggregation point, respectively, includes: Performing one-hot encoding on the abnormal event to obtain initialized initial vector data; Obtaining a preset bidirectional aggregation model, wherein the bidirectional aggregation model is obtained through unsupervised training based on sample abnormal events, sample target events characterized as noise events, and sample aggregation points in the acquired samples; The initial vector data is input into the preset bidirectional aggregation model to perform probability calculation, and the noise probability value of each abnormal event and the corresponding aggregation point is obtained respectively.
9. The abnormal event handling method according to claim 6, characterized in that: Inputting the initial data into a preset two-way aggregation model to perform probability calculation to obtain noise probability values of each abnormal event and the corresponding aggregation point, including: sorting the plurality of aggregation points according to time and storing the clustered points in an aggregation point list; The initial data is input into a preset bidirectional aggregation model, and the probability of the abnormal event is calculated according to each aggregation point in the aggregation point list to obtain the noise probability value of each abnormal event and the corresponding aggregation point.
10. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the abnormal event handling method according to any one of claims 1 to 9 is implemented.
11. A computer-readable storage medium, characterized in that The storage medium stores a program, and the program is executed by a processor to implement the abnormal event handling method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Log data processing method, device and equipment and computer readable storage medium
CN111274095A
Anomaly positioning method and device thereof
CN113986595A