Method and apparatus for anomaly detection

By grouping log data by context and using multiple machine learning models to refine anomaly flags, the method addresses the issue of high false positives in anomaly detection, achieving improved accuracy and efficiency in dynamic environments.

WO2025151055A1PCT designated stage expired Publication Date: 2025-07-17TELEFONAKTIEBOLAGET LM ERICSSON (PUBL)
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/SE2024/050016
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-09
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Existing anomaly detection methods in log data, particularly in dynamic environments like telecommunication networks, suffer from high false positives due to the lack of context in anomaly scores and the inability of static rules to scale across networks, leading to inefficiencies and inaccurate decision-making.

Method used

The method involves grouping log data by context identifiers and training distinct instances of machine learning models for each group, using decision boundaries from multiple models to refine anomaly flags, and initiating corrective actions based on refined anomaly scores.

Benefits of technology

This approach significantly reduces false positives and enhances the accuracy of anomaly detection by leveraging multiple models to validate decision boundaries, ensuring more precise classification of anomalies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SE2024050016_17072025_PF_FP_ABST
    Figure SE2024050016_17072025_PF_FP_ABST
Patent Text Reader

Abstract

A method (100) for anomaly detection in log data, where the log data comprises elements The method comprises grouping (101) the elements by a context identifier and training (102) one instance of a machine learning model for each context identifier, so that each group of elements is associated to a distinct instance of a machine learning model The method further comprises obtaining (103) an initial flagged anomaly and using (104) multiple instances to obtain a refined anomaly flag for a detected anomaly. The method further comprises, based on the refined anomaly flag, initiating (105) an action to correct the detected anomaly. Also disclosed are a related apparatus, computer program, and computer program product.figure for publication:
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD AND APPARATUS FOR ANOMALY DETECTION

[0002] TECHNICAL FIELD

[0003] The disclosure relates to a method for anomaly detection in log data. Further disclosed are a related apparatus, a computer program, and a computer program product.

[0004] BACKGROUND

[0005] Anomaly detection is widely applied in various industries and research areas where statistical deviations are expected to emerge. One way to implement anomaly detection is to use machine learning which has been proven to be effective especially in data intensive and dynamic environments.

[0006] In telecommunication networks, enormous data sources are available and various logs are commonly being used. These logs are often utilized to produce understanding of the status of the network and also various logics can be created to find and flag the logs of interest. However, the large volumes and dimensions, in dynamic network environments, introduce challenge where static rules may not apply from one network to another.

[0007] Static anomaly detection rules may not scale, i.e. static rules may not be transferrable from one network to another. Moreover, anomaly detection solutions tend to generate false positives. This is problematic especially in solutions where amount and dimension of data is large, e.g., telecommunication networks.

[0008] Additionally, anomaly detection solutions tend to utilize anomaly scores as a main factor in decision making. For example, in US 20210174253 A1 , systems and method for detecting anomalies in machine-generated logs are disclosed. The methods use anomaly scores of a single machine learning model to determine whether a log message is anomalous. However, anomaly scores lack in context, and individual machine learning models may flag relatively high numbers of anomalies to avoid being penalized for missing potential anomalies in the training phase. It is therefore of interest to develop a method for anomaly detection which avoid the high number of false positives generated by the state of the art. SUMMARY

[0009] It is an object of the present disclosure to present a more accurate anomaly detection method for log data.

[0010] According to a first aspect of the disclosure, there is a method for anomaly detection in log data. The log data comprises elements. The method comprises grouping the elements by a context identifier. The method comprises training one instance of a machine learning model for each context identifier, so that each group of elements is associated to a distinct instance of a machine learning model. The method comprises obtaining an initial, flagged anomaly. The method comprises using decision boundaries from multiple instances of a machine learning model to obtain a refined anomaly flag for a detected anomaly. The method comprises, based on the refined anomaly flag, initiating an action to correct the detected anomaly.

[0011] According to an embodiment of the first aspect, the log data is log data from a wireless communication network or vehicle communication network.

[0012] According to an embodiment of the first aspect, the method comprises applying a coder to the log data to obtain a numerical value.

[0013] According to an embodiment of the first aspect, the instances of machine learning models comprise instances of distinct classes of machine learning models.

[0014] According to an embodiment of the first aspect, obtaining an initial anomaly flag associated to a flagged element comprises obtaining information about which decision boundaries of the machine learning instance are breached by the flagged element.

[0015] According to an embodiment of the first aspect, using multiple instances of a machine learning model to improve the confidence of the initial anomaly flag comprises using each of the multiple instances to classify the flagged element as anomalous or not anomalous.

[0016] According to an embodiment of the first aspect, the method comprises obtaining, for each of the multiple instances, information regarding which decision boundaries of the decision boundaries of the multiple instances are violated by the flagged element. According to an embodiment of the first aspect, the method comprises statistically evaluating the number of times each decision boundary is violated by the different instances to obtain the refined anomaly flag.

[0017] According to an embodiment of the first aspect, initiating an action to correct the anomaly comprises one or more of: creating an internal alert event; creating an internal log event; and creating an external alert event.

[0018] According to a second aspect of the disclosure, there is an apparatus comprising processing circuitry and a memory. The apparatus is configured for anomaly detection in log data, the log data comprising elements. The apparatus is configured to group the elements by a context identifier. The apparatus is configured to train one instance of a machine learning model for each context identifier, so that each group of elements is associated to a distinct instance of a machine learning model. The apparatus is configured to obtain an initial flagged anomaly. The apparatus is configured to use decision boundaries from multiple instances of a machine learning model to obtain a refined anomaly flag. The apparatus is configured to, based on the refined anomaly flag, initiate an action to correct a detected anomaly.

[0019] According to embodiments of the second aspect, the apparatus is configured to perform a method according to any one embodiment of the first aspect.

[0020] According to an embodiment of the second aspect, the apparatus is a network node in a telecommunications network.

[0021] According to an embodiment of the second aspect the apparatus is a server host in an Internet of Things network.

[0022] According to a third aspect of the disclosure, there is a computer program comprising machine readable instructions which, when executed on a processor of an apparatus, cause an apparatus to perform a method according to any one embodiment of the first aspect.

[0023] According to a fourth aspect of the disclosure, there is a computer program product comprising a non-transient storage medium on which a computer program according to the third aspect is stored.

[0024] BRIEF DESCRIPTION OF THE DRAWINGS Fig. 1 depicts a flowchart of a method disclosed herein.

[0025] Fig. 2 depicts an example of log data.

[0026] Fig. 3 depicts a communication network in which a method disclosed herein can be performed.

[0027] Fig. 4 depicts a flowchart of a method disclosed herein.

[0028] Fig. 5 depicts a logical block diagram of a part of a method disclosed herein.

[0029] Fig. 6 depicts a logical block diagram of a part of a method disclosed herein.

[0030] Fig. 7 depicts a logical block diagram of a part of a method disclosed herein.

[0031] Fig. 8 depicts an apparatus according to the disclosure.

[0032] Fig. 9 depicts a network node according to the disclosure.

[0033] Fig. 10 depicts a server host according to the disclosure.

[0034] Fig. 11 depicts output data of a test of methods according to the disclosure.

[0035] DETAILED DESCRIPTION OF THE DRAWINGS

[0036] Fig. 1 depicts a method 100 according to the disclosure. The method is a method for anomaly detection in log data. The log data may, in a first embodiment, be log data from a telecommunications network. The first embodiment is an offline implementation of the method. An example of log data from a telecommunications network is shown in Fig. 2. Each row of the table of Fig. 2 corresponds to an element of log data. Each element of the log data of Fig. 2 comprises human-readable information about a kind of event, an application version associated to the element of log data, a level, a stage, a request uniform resource identifier, URI, a verb associated with the log data (e.g. “get”, “put”, “delete”, or “post”), a user associated to the log data, a source internet protocol, IP, address, a name of a user agent, and a response status. The method comprises grouping 101 the log data by a context identifier. In the first embodiment, the context identifier is a request application programming interface, API, associated to each element of the log data. Grouping the log data by the context identifier comprises creating a set of groups of log data, where the elements of a group are all associated with the same request API. The method further comprises training 102 an instance of a machine learning model for each context identifier, so that each group of log data is associated to a distinct instance of a machine learning model. In the first embodiment, training one instance of a machine learning model for each group comprises training one isolation forest model for each group of elements of log data. That is, each request API is associated to an instance of an isolation forest model trained on the corresponding group of log data elements.

[0037] The method further comprises obtaining 103 an initial anomaly score associated to a flagged anomaly. In the first embodiment, obtaining the initial anomaly score comprises analyzing a historical stream of log data by, for each element of the log data, retrieving the instance of the isolation forest associated to the same request API as the element of log data, and classifying the element of the log data using the retrieved instance of the isolation forest. Classifying the element of log data using the retrieved instance of the isolation forest comprises using the element of log data as input and letting the isolation forest classify the element as either an anomaly or not an anomaly. If the element is classified as no anomaly, the first embodiment comprises restarting with the next element of the stream of log data. If the retrieved machine learning model flags / classifies the element of log data as an anomaly, the first embodiment comprises obtaining an initial anomaly flag by retrieving, from the model, information regarding which decision boundaries of the model were breached by the element of log data.

[0038] The method further comprises using 104 decision boundaries from multiple machine learning models of the group of machine learning models to obtain a refined anomaly score. In the first embodiment, using decision boundaries from multiple machine learning models of the group of machine learning models to obtain a refined anomaly score comprises retrieving a plurality of instances of isolation forest models associated to other request APIs than the request API associated to the element of log data. In the first embodiment, the plurality of instances of isolation forest models is a fixed number. In the first embodiment, refining the anomaly score comprises classifying the log entry using each of the retrieved plurality of isolation forest models to classify the element of log data as an anomaly or not an anomaly. In the first embodiment, refining the anomaly score comprises determining whether the element of log data breaches a decision boundary for a majority of the isolation forest instances. If a decision boundary, such as the one associated to a source internet protocol, IP, address, is breached by the element of log data in a majority of the instances of the isolation forest, the refined anomaly score classes the element of log data as an anomaly with respect to that decision boundary. If a decision is not breached by the element of log data in a majority of the instances of the isolation forest, the refined anomaly score classes the element of log data as not an anomaly with respect to that decision boundary.

[0039] The method further comprises, based on the refined anomaly score, initiating (105) an action to correct a detected anomaly. In the first embodiment, where an element of log data has been found anomalous with respect to the decision boundary associated to the source IP address, an action to correct the anomaly comprises creating an event alert, wherein creating an event alert comprises one or more of providing an event alert to an API, creating a log event using an internal logging system, and / or creating an alert event to an integrated external computing system. If the element of log data was found to not be anomalous with respect to any decision boundary, the first embodiment comprises returning to the stream of log data until a new element of log data is flagged as an anomaly.

[0040] Methods according to the first embedment provide more accurate, but still computationally efficient classification of log data as anomalous or as not anomalous. Using the decision boundaries from multiple machine learning models allows for a more precise determination of whether a decision boundary is violated due to a statistical coincidence or due to an actual anomaly.

[0041] A second embodiment (not illustrated) of the method is an online embodiment of the method. The second embodiment comprises live data streams from a telecommunications network. According to the second embodiment, the method comprises the same type of log data as in Fig. 2 and the first embodiment, but streamed live from a network. Grouping 101 the data according to context identifier comprises routing the elements of log data into parallel streams by context identifier, which in this embodiment comprises request API.

[0042] The method further comprises training 102 one instance of a machine learning model for each context identifier. The second embodiment is live online and thus uses pre-trained instances. Training the instances comprises refining each instance of a pre-trained isolation forest using the live data stream corresponding to the instance. Refining the instances comprises using the live network data to periodically refine the isolation forest model.

[0043] The method further comprises obtaining an initial flagged anomaly. In the second embodiment, obtaining an initial flagged anomaly comprises feeding each parallel stream of elements of log data to the corresponding instance of an isolation forest. When an element of log data is classified by an instance of the isolation forest as an anomaly, it is flagged for further evaluation.

[0044] The method further comprises using 104 decision boundaries from multiple instances of the machine learning model to obtain a refined anomaly flag. In the second embodiment, using decision boundaries from multiple instances of the machine learning model to obtain a refined anomaly flag comprises obtaining, from the instance of the isolation forest which flagged the element of log data, the decision boundary or decision boundaries violated by the element of log data. In the second embodiment, the decision boundary which is violated may be a response time for an API call. The second embodiment further comprises adding the flagged element of log data to a plurality of other data streams, thereby obtaining any decision boundaries which the element of log data violates for the plurality of other instances of an isolations forest corresponding to the plurality of other data streams.

[0045] The plurality of other data streams are, in the second embodiment, selected randomly out of all data streams to obtain a specific number of other data streams.

[0046] The second embodiment further comprises refining the anomaly flag by checking, for each decision boundary, whether the decision boundary is violated for a majority of the instances. If at least one decision boundary is violated by the element of log data, the anomaly flag remains and an action to correct the detected anomaly is initiated. An action to correct the detected anomaly is, in the second embodiment, creating an event to an API as an event producer, thereby supporting an eventbased architecture and utilization of signals from an apparatus implementing the second embodiment by external systems. Alternatively, an action to correct the detected anomaly may comprise creating a log event using an internal event logging system, where the log event may comprise a warning level determined based on at least which decision boundaries were violated. Alternatively, an action to correct the anomaly may comprise creating an alert event to an integrated external system, using for example an HTTP protocol to post the anomaly information to the integrated external system. If no decision boundary is violated in a majority of the instances, the flag is cleared and the method returns to grouping 101 the log data to wait for another anomaly flag to be raised.

[0047] In embodiments of the method 100, the log data may originate from a telecommunications network such as the telecommunications network of Fig. 3.

[0048] In other embodiments of the method 100, the log data may originate from an Internet of Things, loT network. In some embodiments, an loT network may be incorporated in a telecommunications network 302 in the communication system 300 of Fig. 3. The communication system 300 comprises a host 301 in communication with the telecommunications network 302. The telecommunication network comprises a core network 306, and the core network comprising at least one core network node 308. The telecommunications network further comprises an access network 304, and the access network comprises at least one network node 310A, 310B, such as a radio network node. The network nodes provide access to the network to user equipments, UEs, 312A, 312B, or alternatively or in addition, to a hub 314, the hub providing access to the network to the UEs 312C, 312D.

[0049] In some embodiments of the method 100, the log data may originate from a network implementing a local area networking protocol / short-range radio communication protocol, such as Wi-Fi / IEEE 802.16.

[0050] In some embodiments of the method 100, the log data may originate from a hybrid network implementing some combination of telecommunications protocols, local area networking protocols, and short-range wireless protocols.

[0051] Fig. 4 depicts additional, optional, methods steps of the method 100. In some embodiments of the method 100, grouping 101 the log data by a context identifier comprises applying 106 a coder to each element of the log data to obtain a numerical value. Applying a coder to the log data comprises encoding the log data, which in embodiments is in human-readable form, into numerical form. An advantage of encoding the log data as a numerical value is very efficient evaluation of the decision boundaries of the instances of a machine learning model. In some embodiments of the method 100, the instances of machine learning modes may comprise instances of distinct classes of machine learning models. By way of example, elements associated to a first context identifier may have characteristics such that an unsupervised tree-type machine learning model performs very well for classifying the elements as anomalous or not anomalous. However, elements associated to a second context identifier may perform poorly with an unsupervised tree-type machine learning model, and may perform better with a supervised clustering-type machine learning algorithm. In such embodiments, the multiple instances used to obtain a refined flag may be selected exclusively from the instances using the same type of machine learning model as the machine learning model which raised the flag.

[0052] Obtaining 103 an initial anomaly flag associated to a flagged element may comprise obtaining 107 information about which decision boundaries of the machine learning instance are breached by the flagged element. When a machine learning algorithm is trained, the machine learning algorithm learns the boundaries of a multi-dimensional polyhedron, where each axis of the polyhedron corresponds to a dimension or an attribute of the training data. Information about which decision boundaries are breached by the flagged element therefore comprises identifying along which axes of the polyhedron the element is outside the polyhedron. The attributes associated to each violated decision boundary may be included in the initial anomaly flag. In embodiments where a coder is used to transform the training data into numerical form, determining which decision boundaries are breached is very computationally efficient.

[0053] Using 104 multiple instances of a machine learning model to obtain a refined anomaly flag for a detected anomaly may comprise using 108 each of the multiple instances to classify a flagged element as anomalous or not anomalous. The multiple instances may comprise a fixed number of instances selected from the trained instances. In embodiments, the fixed number of instances may be at least 100. In some embodiments, the multiple instances may be selected randomly. In some embodiments, a fixed proportion of the number of instances are selected. In some embodiments, instances which correspond to similar context identifiers may be selected. In some embodiments, a fixed number of instances may be selected, and the method may comprise checking whether there is convergence in the proportion of times the element breaches each decision boundary. If there is convergence, a decision on whether to keep the anomaly flag or clear the anomaly flag may be taken. If there is not convergence, additional instances may be selected until there is convergence.

[0054] Using 104 multiple instances of a machine learning model to obtain a refined anomaly flag for a detected anomaly may comprise obtaining 109, for each of the multiple instances, information regarding which decision boundaries of the decision boundaries of the instances are violated by the flagged element. In embodiments, a first group of the multiple instances may flag the element as anomalous and a second group of the multiple instances may not flag the element as anomalous. Obtaining information regarding which decision boundaries of the decision boundaries of the instances are violated by the flagged element may thus comprise obtaining, for each of the instances in the first group of instances, information regarding which decision boundaries are violated by the flagged element. The instances are trained on the same type of data, and therefore each comprise a decision polyhedron defined by the same axes, but different parameters, as the first instance which flagged the element as anomalous. Obtaining information about the specific decision boundaries allows for a more nuanced evaluation, since a non- anomalous element may be flagged along a single boundary of one instance by chance, but a non-anomalous element being flagged along the same boundary of multiple instances is less probable.

[0055] Using 104 multiple instances to obtain a refined anomaly flag may thus further comprise statistically evaluating 110 the number of times each decision boundary is violated by the element to obtain the refined anomaly flag. Statistically evaluating the number of times each decision boundary is violated by the element to obtain the refined anomaly flag may comprise determining whether any decision boundary is violated by a majority of the instances. Alternatively, statistically evaluating the number of times each decision boundary is violated may comprise determining whether any decision boundary is violated more than some threshold number of times. Alternatively, statistically evaluating the number of times each decision boundary is violated may comprise determining whether the total number of violations exceed some threshold number. Alternatively, statistically evaluating the number of times each decision boundary is violated may comprise assigning a relative importance to each decision boundary, and assigning different thresholds to each decision boundary depending on the relative importance of the decision boundary.

[0056] Obtaining the refined anomaly flag may thus comprise determining, using statistical test, whether any one decision boundary is violated sufficiently many times to retain the anomaly flag. If no decision boundary is violated sufficiently many times, the anomaly flag may be cleared and the next flagged element of log data evaluate. If at least one decision boundary is violated in sufficiently many instances, the initial anomaly flag remains and is now a refined anomaly flag.

[0057] The method further comprises initiating 105 an action to correct the anomaly. Initiating an action to correct the anomaly may comprise creating an alert event to an API, with the apparatus performing the method acting as an event producer / API provider supporting an event-based architecture. Initiating an action to correct the anomaly may comprise creating a log event using an internal log event logging system. The internal logging system may flag elements using an information level when an event is found to not be anomalous, and using a warning level when an event is found to be anomalous. The precise information level and the precise warning level may be determined using the exact number of violated boundaries.

[0058] Initiating an action to correct the anomaly may comprise creating an alert event to an integrated external system. Creating an alert event to an external system may comprise using, for example, a hypertext transfer protocol POST message to post the anomaly information to the external system.

[0059] Fig. 5 depicts a logical arrangement of computers / nodes performing an offline embodiment of a part of a method as disclosed herein. Fig. 5 depicts a repository 501 for historic log data, an apparatus 502 which may request or receive the historic log data from the repository 501 and use the historic data to train the multiple instances of the machine learning model, and a repository 503 for storing the trained machine learning models.

[0060] Fig. 6 depicts a logical arrangement of nodes performing an online embodiment of a part of a method as disclosed herein. The arrangement of Fig. 6 comprises an apparatus 601 , receiving a live log data stream 602. For each log data element, the apparatus may retrieve an instance of a machine learning model from a repository 604, the instance corresponding to the same context as the element. The apparatus may use the instance to classify the element as anomalous or not anomalous. If the retrieved instance classifies the element as anomalous, the apparatus may request from the repository additional instances of the machine learning model and use each of the additional instances to classify the element as anomalous or not anomalous. The apparatus may further extract information regarding which decision boundaries were violated for each instance, and apply statistical analysis methods to determine whether the element should remain flagged with a refined anomaly flag.

[0061] If the element is classed as not anomalous, the apparatus may alert an anomaly monitor 603 that the element is not anomalous. If the element is flagged as anomalous, the apparatus may inform the anomaly monitor that the element has a refined anomaly flag. The monitor may then take further actions associated with correcting anomalous elements.

[0062] Fig. 7 depicts a logical architecture associated with an online implementation of a method according to the disclosure. The logical architecture of Fig. 7 comprises a repository 701 of audit logs, originating from for example Kubernetes. The logical architecture further comprises a node 702 tasked with routing the elements of the log depending on an associated context. Elements associated to a first context may be routed to a first training data collection node 703, which further transmits the collected training data to a training node 704 associated to the first context, which uses the training data to train the instance associated to the first context, and passes the trained first instance to a repository 705 from which the trained first instance can be retrieved. Similarly, there is a second training data collection node 706, which collects and transmits data associated to a second context to a second training node 707, which uses the training data to train the instance associated to the second context, and passes the trained second instance to a repository 708 from which the trained second instance may be retrieved. Similarly, there is an nthtraining data collection node 709, which collects and transmits data associated to an nthcontext to an nthtraining node 710, which uses the training data to train an nthinstance associated to the nthcontext, and passes the trained nthinstance to a repository 711 from which the trained nthinstance may be retrieved. With reference to the logical architectures and the method as described above, a third embodiment of the method as disclosed herein will be described. The third embodiment is a combined offline / online embodiment where the log data is log data from a telecommunications network. As a first step of the third embodiment, a repository of historical log data is accessed. The historical log data is grouped by API endpoint and the historical log data is further grouped by a user identity associated to the log data. The second embodiment comprises training one instance of a local outlier factor model for each group of log data to obtain one trained instance of a local outlier factor model for each group of historical log data.

[0063] The second embodiment further comprises storing the trained instances in a repository of instances. The second embodiment further comprises a live portion, where a logical entity receives a live log data stream from a network. The logical entity is a node capable of routing the elements of the live log data stream according to an associated API endpoint and according to an associated user identity to a group of logical evaluation entities, each logical evaluation entity corresponding to one group of log data. The logical evaluation entity retrieves the instance of the machine learning model corresponding to the API endpoint and user identity and uses the instance to evaluate whether the element is an anomaly of not. If the instance classifies the element as not an anomaly, the element is assumed to be not an anomalous and the method returns to scan the next element of log data. The logical evaluation entities of the third embodiment process data in parallel. If the element is classed as not anomalous, the element receives an initial anomaly flag. The initial anomaly flag comprises information from the evaluation entity regarding which decision boundary or which decision boundaries of the local outlier factor model instance are violated by the element.

[0064] If the element receives an initial anomaly flag, the third method further comprises retrieving, from the repository, a plurality of additional instances. The plurality of instances comprises, in the third embodiment, 10% of the total number of instances selected at random. The third embodiment further comprises using each of the plurality of instances to evaluate the element and retrieving, from each instance which classifies the element as anomalous, information about which decision boundary or which decision boundaries of the decision boundaries are violated by the element. In the third embodiment, different attributes of the log data are associated to different sensitivities. For example, the attribute request URI is considered more sensitive and has an associated threshold of only 25% of the instances finding the associated boundary to be violated to cause a flag. The attribute API version, on the other hand, is considered less sensitive and 60% of the instances must find the associated boundary to be violated to cause a flag.

[0065] The third embodiment thus further comprises obtaining, from each of the instances classifying the element as anomalous, information regarding which decision boundary or which decision boundaries of the instance are violated by the element. The third embodiment comprises statistically evaluating the number of times each decision boundary was violated to determine whether the element should retain a refined anomaly flag, in relation to each of the specific thresholds for the attributes.

[0066] If the element is found to not violate any one boundary sufficiently many times, the anomaly flag is cleared and the next element in the stream of log elements can be processed.

[0067] If the element is found to violate one or more thresholds, the flag remains as a refined anomaly flag. The third embodiment then generates an alert using the hypertext transfer protocol to the system from which the log events originate.

[0068] Embodiments of the method as disclosed herein may be performed by an apparatus, as in Fig. 8. The apparatus 800 comprises processing circuitry 801 and a memory 802. The memory may further comprise a computer program 803 or a computer program product 804, the computer program product comprising the computer program 803. The apparatus may be contained in a network node in a communications system such as the radio access node of Fig. 9. In addition to the apparatus 800, the radio access node comprises a communication interface 906, the communication interface comprising an antenna 910, radio front-end circuitry 318 comprising a filter 920 and an amplifier 922, and a port / terminal 916. The radio access node comprises a power source 908. The radio access node comprises processing circuitry 905, the processing circuitry 905 comprising radio frequency transceiver circuitry 912 and baseband circuitry 914. The radio access node further comprises a memory 904. In some embodiments of the method described herein, the apparatus performing the method is a server host in an loT network, the loT network comprising at least one server host and a plurality of loT devices, wherein at least one server host may be configured to perform a method according to the disclosure. An loT device may be a device for use in one or more application domains, these domains comprising, but not limited to, home, city, wearable technology, extended reality, industrial application, and healthcare. Fig. 10 depicts an example of a server host which may receive log data from loT devices according to Fig. 10 and perform embodiments of the method as depicted herein. The server host 1100 comprises processing circuitry 1102, an input / output interface 1106, a network interface 1108, and a power source 1110. The server host further comprises a bus 1104. The server host further comprises a memory 1112, the memory comprises host application programs 1114 and data 1116.

[0069] Fig. 11 depicts results when applying methods according to the disclosure to log data from a real telecommunications system. The dataset held over 600 000 events, of which 30 were known to be anomalous. The first set of columns, denoted A, depicts the accuracy (A1 ), precision (A2), recall (A3), and F1 -score (A4) when applying a simple isolation forest model to find anomalous results. The second set of columns, denoted B, depicts the accuracy (B1 ), precision (B2), recall (B3), and F1- score (B4) when applying an embodiment of the method presented herein to identify anomalous results. As can be seen, the accuracy, recall, and F1 -score all improve with the method presented herein. Moreover, simply applying an isolation forest predicts 2 627 anomalies, while the method of the disclosure identifies 14, several orders of magnitude closer to the ground truth of the data set. Using a local outlier factor model provides even more impressive results. The third set of columns, denoted C, depicts the accuracy (C1 ), precision (C2), recall (C3), and F1 score (C4) when naively applying a local factor model to the data. All values are higher for the naive local factor model than for the naive isolation forest. In the fourth set of columns in depicted using a local outlier factor model with the method according to the disclosure. As can be seen, the local factor model gives a near-perfect accuracy (D1 ), precision (D2), recall (D3), and F1 -score (D4) on the test data. Moreover, the total number of flagged anomalies with the naive local outlier factor model was 796, while the total number of flagged anomalies using the method according to the disclosure was 31 , including all of the 30 known anomalies.

Claims

CLAIMS1 . A method for anomaly detection in log data, the log data comprising elements, the method comprising: grouping (101) the elements by a context identifier; training (102) one instance of a machine learning model for each context identifier, so that each group of elements is associated to a distinct instance of a machine learning model; obtaining (103) an initial, flagged anomaly; using (104) decision boundaries from multiple instances of a machine learning model to obtain a refined anomaly flag for a detected anomaly; based on the refined anomaly flag, initiating (105) an action to correct the detected anomaly.

2. The method (100) according to claim 1 , wherein the log data is log data from a wireless communication network or vehicle communication network.

3. The method (100) according to claim 1 or 2, comprising applying (106) a coder to the log data to obtain a numerical value.

4. The method (100) according to any one of claims 1-3, wherein the instances of machine learning models comprises instances of distinct classes of machine learning models.

5. The method (100) according to any one of claims 1-4, wherein obtaining (103) an initial anomaly flag associated to a flagged element comprises obtaining (107) information about which decision boundaries of the machine learning instance are breached by the flagged element.

6. The method (100) according to any one of claims 1-5, wherein using (104) multiple instances of a machine learning model to improve the confidence of the initial anomaly flag comprises using (108) each of the multiple instances to classify the flagged element as anomalous or not anomalous.

7. The method (100) according to claim 6, comprising obtaining (109), for each of the multiple instances, information regarding which decision boundaries ofthe decision boundaries of the multiple instances are violated by the flagged element.

8. The method (100) according to claim 7, comprising statistically evaluating (110) the number of times each decision boundary is violated by the different instances to obtain the refined anomaly flag.

9. The method (100) according to any one of claims 1-8, wherein initiating (105) an action to correct the anomaly comprises one or more of:- creating an internal alert event;- creating an internal log event; and- creating an external alert event.

10. An apparatus (800) comprising processing circuitry (801 ) and a memory (802), the apparatus configured for anomaly detection in log data, the log data comprising elements, the apparatus configured to: group (101 ) the elements by a context identifier; train (102) one instance of a machine learning model for each context identifier, so that each group of elements is associated to a distinct instance of a machine learning model; obtain (103) an initial flagged anomaly; use (104) decision boundaries from multiple instances of a machine learning model to obtain a refined anomaly flag; based on the refined anomaly flag, initiate (105) an action to correct a detected anomaly.11 . The apparatus (800) according to claim 10, configured to perform a method according to any one of claims 2-9.

12. The apparatus (800) according to claims 10 or 11 , wherein the apparatus is a network node in a telecommunications network (302).

13. The apparatus (800) according to claims 10 or 11 , wherein the apparatus is a server host (1100) in an Internet of Things network.

14. A computer program (803) comprising machine readable instructions which, when executed on a processor of an apparatus, cause the apparatus to perform a method according to any one of claims 1-9.

15. A computer program product (804) comprising a non-transient storage medium on which a computer program (803) according to claim 14 is stored.

Citation Information

Patent Citations

  • Analysis of system log data using machine learning

    US20210174253A1

  • Hybrid Machine Learning to Detect Anomalies

    US20210281592A1

  • Unsupervised anomaly detection with self-trained classification

    WO2022251462A1