Enhanced anomaly detection in computing environment

By identifying the dominant mode of non-exception blocks in the computing environment and generating exception vectors, the problem of high consumption of log data processing in the prior art is solved, and efficient exception detection and prediction are achieved.

CN120429155APending Publication Date: 2025-08-05ORACLE INT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510638022.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-10-03
Filing Date
2020-10-02
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The prior art requires a large amount of log data processing when detecting and predicting abnormal behavior in a computing environment, resulting in high consumption of bandwidth and computing resources, limiting the effectiveness and deployment of abnormal analysis.

Method used

By receiving log data at edge nodes, identifying dominant patterns of non-exception blocks, generating exception vectors, and distributing them to nodes to detect exceptions, reducing the amount of data to be transmitted and computational complexity.

Benefits of technology

It realizes more efficient anomaly detection and prediction, reduces bandwidth requirements and computing overhead, and can detect abnormal events in a timely manner and take mitigation measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429155A_ABST
    Figure CN120429155A_ABST
Patent Text Reader

Abstract

The invention relates to enhanced anomaly detection in a computing environment. An exception service receives log data from a node in a computing environment, which includes a sequence of information indicative of log messages generated by the node. The anomalous service identifies a dominant pattern in the sequence of information representing non-anomalous blocks of the log message. Upon identifying the dominant mode, the service can extract non-anomalous blocks from the log data to reveal anomalous blocks that do not conform to the dominant mode. The service may then generate an anomaly vector based on the anomaly blocks, which can be distributed to the nodes to detect anomalies.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the patent application for invention titled "Enhanced Anomaly Detection in a Computing Environment" with the application number 202080069864.2 and the filing date of October 2, 2020. Technical Field

[0002] The solutions disclosed herein relate to block-based detection and prediction of low-probability behaviors (e.g., anomalous behaviors) in a computing environment. Background Art

[0003] Modern information services process large amounts of data and generate large amounts of log data in that processing. Log data is generated by code at runtime to provide a record of the state of one or more components of the service. Log data can be useful in troubleshooting or otherwise maintaining the service. Examples of log data include log statements and operational metrics, both of which can be analyzed to predict anomalous conditions or events such as processing failures, server failures, and computer hardware failures.

[0004] Anomalies can be detected and predicted by analyzing large amounts of log data for patterns related to past anomalies in the statements or metrics. Then a pattern can be deployed against the data such that the presence of a given pattern in the data triggers a mitigation action or an alert. Unfortunately, this pattern extraction requires large amounts of log data, which is very expensive both to transport and to process.

[0005] In fact, the amount of log data that needs to be sent from the edge to the central server to successfully extract useful patterns can easily approach the amount of normal operation data being sent in the same direction - assuming there is bandwidth to do so. Additionally, the amount of computation required to find patterns within a reasonable time frame can exceed the amount of computation initially allocated for normal operation. To date, these limitations have hindered the development and deployment of effective anomaly analysis. Summary of the Invention

[0006] Techniques for improving anomaly condition detection and prediction in a computing environment are disclosed herein. In various embodiments, an anomaly service receives log data from edge nodes in a computing environment, which includes information sequences indicative of log messages generated by the nodes. The anomaly service identifies a dominant pattern in the information sequences representing non-anomalous blocks of the log messages. After identifying the dominant pattern, the service is able to extract non-anomalous blocks from the log data to reveal anomalous blocks that do not conform to the dominant pattern. Then, the service can generate anomaly vectors based on the anomalous blocks, and the anomalous blocks can be distributed to the edge nodes to detect anomalies.

[0007] In the same or other embodiments, one or more nodes in a computing environment receive an exception vector. When an event occurs in the computing environment, a log message is generated, and in response, the node generates a corresponding sequence of hash values. A sequence vector is generated based on the sequence of hash values, which can be compared with or otherwise evaluated against the exception vector to determine whether the log message indicates the occurrence of one or more exception events.

[0008] This summary is provided to introduce some concepts in a simplified form that will be further described in the technical disclosure below. It is understood that this summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Many aspects of the present disclosure can be better understood with reference to the following drawings. Although several embodiments are described in conjunction with these drawings, the present disclosure is not limited to the embodiments disclosed herein. On the contrary, it is intended to cover all alternatives, modifications, and equivalents.

[0010] Figure 1 Illustrates the operating system architecture in an implementation.

[0011] Figure 2 Illustrates the log generation process in an implementation.

[0012] Figure 3 Illustrates the exception block extraction process in an implementation.

[0013] Figure 4 Illustrates the exception detection process in an implementation.

[0014] Figure 5 Illustrates the operating scenario in an implementation.

[0015] Figure 6 Illustrates the log generation scenario in an implementation.

[0016] Figure 7 Illustrates the exception block extraction scenario in an implementation.

[0017] Figure 8 Further illustrates the exception block extraction scenario.

[0018] Figures 9A - 9C Illustrates the exception detection scenario in an implementation.

[0019] Figure 10 Illustrates a computing system suitable for implementing the enhanced exception detection techniques disclosed herein, including any architecture, process, operating scenario, and operation sequence shown in the drawings and discussed in the detailed description below. DETAILED DESCRIPTION

[0020] The solutions disclosed herein relate to block - based detection and prediction of low - probability behaviors (e.g., anomalous behaviors) in a computing environment. An anomaly service receives log data from edge nodes in one or more computing environments, where the edge nodes can be remote from the service and / or co - located relative to the service. The log data includes a sequence of information indicative of log messages generated by the nodes at runtime.

[0021] The anomaly service identifies a dominant pattern in the sequence of information of non - anomalous blocks representative of log messages. The service extracts or otherwise ignores non - anomalous blocks from the log data to reveal anomalous (or low - probability) blocks that do not conform to the dominant pattern. Then, the service generates anomaly vectors based on the anomalous blocks and distributes them to the nodes.

[0022] One or more nodes in the (one or more) computing environments receive the anomaly vectors and use them to detect anomalous behaviors in the flow of log messages. A given node generates a sequence of hash values based on a sequence of log messages associated with events in the computing environment. The node also generates a sequence vector based on the hash values. The sequence vector includes numbers corresponding to a set of possible hash values, for example, by position.

[0023] Then, the node performs a comparison of the sequence vector with the set of anomaly vectors to determine whether at least a portion of the hash values indicates the occurrence of one or more anomalous events. If so, then one or more actions can be taken to mitigate the risk of the problematic or anomalous events predicted by the comparison.

[0024] Returning to the anomaly service, the service can identify the dominant pattern by finding potential patterns within the information sequence and scoring them. Then the dominant pattern can be selected based on the scoring. In some scenarios, the scoring function can be used to elevate a subset of potential patterns that occur frequently within the information sequence relative to different subsets of potential patterns that occur less frequently within the information sequence. For each potential pattern among the potential patterns, the scoring function determines the relative advantage of the potential pattern based on the description length of the information sequence when encoded using a compressed representation of the potential pattern.

[0025] In some embodiments, the information sequence includes a sequence of log codes, and each log code can be a symbolic representation of a possible hash value. In this case, the log data will also include a codebook or dictionary that maps each unique log code to a different hash value in the set of possible hash values.

[0026] In various embodiments, each anomaly vector includes a string of numbers that correspond by position to a set of possible log codes. The anomaly service sets the value of each number in the string to indicate the presence or absence of the corresponding log code in the possible log codes in the anomalous block represented by the anomaly vector.

[0027] Referring again to the client node, a given node generates a sequence vector by computing a weighted value for each digit of a numeric string. The node computes the weighted value by multiplying a weighting factor by a binary value to obtain the weighted value. The binary value indicates whether the sequence of hash values includes the hash value corresponding to the position of a given digit in the string, and the weighting factor indicates the relative importance of the log message compared to other log messages.

[0028] In some embodiments, the node performs a comparison of the sequence vector with a set of anomaly vectors by applying a cosine similarity function to the inner dot product of the sequence vector with each anomaly vector in the set of anomaly vectors. The results can be compared to determine whether the sequence vector sufficiently matches one or more of the anomaly vectors to indicate an anomalous state of the computing environment.

[0029] It can be appreciated that the anomaly detection techniques disclosed herein allow for more efficient detection of events that deviate from standards, normal, or expected, with less operational overhead. For example, the log codes sent by the client node to the anomaly service require less bandwidth than the complete log statements themselves, and distributing the anomaly vectors to the edge allows anomalies to be detected at their source.

[0030] It can further be appreciated that low-probability events - not just anomalous events - can also be predicted and / or detected. In such embodiments, the service receives log data from nodes in the computing environment, where the log data includes an information sequence indicating log messages generated by the nodes. The service identifies a dominant pattern in the information sequence that represents high-probability blocks of the log messages, and extracts the high-probability blocks from the log data based on the dominant pattern, thereby revealing low-probability blocks that do not conform to the dominant pattern.

[0031] Then, the service generates a low-probability vector based on at least the low-probability blocks and distributes the data to the nodes to detect low-probability events. Then, the nodes can deploy the low-probability vectors (including in the data) for new log message streams and other such data.

[0032] Turning to the drawings, Figure 1 An operational environment 100 in an implementation is illustrated. The operational architecture 100 includes an anomaly service 110 that communicates with a computing environment 101 and a computing environment 111. The computing environment 101 includes one or more edge computing nodes represented by computing nodes 103 and 105. The computing environment 111 also includes one or more computing nodes represented by computing nodes 113 and 115. The anomaly service 110 receives log data from the computing environments 101 and 111 and sends anomaly vectors to the computing environments 101 and 111. The anomaly service 110, as well as the computing environments 101 and 111, can be implemented with Figure 10 one or more server computers represented by the computing system 1001 in.

[0033] In operation, computing environments 101 and 111 employ a log generation process 200 to generate log data for an exception service 110. The exception service 110 employs an extraction process 300 relative to the log data to extract exception blocks and generate an exception vector that is later supplied to the computing environments 101 and 111. Then, the computing environments 101 and 111 employ a detection process 400 to detect, predict, or otherwise identify an exception condition in their respective environments.

[0034] The log generation process 200 can be implemented in program instructions in the context of any software application, module, component, or other such program element of the computing environments 101 and 111. For example, the log generation process 200 can be implemented in the context of computing nodes 103 and 105 in the computing environment 101 and in the context of computing nodes 113 and 115 in the computing environment 111. In addition to or instead of their respective computing nodes, the log generation process 200 can also be implemented on any other computing element of the computing environments 101 and 111. The program instructions direct the underlying physical or virtual computing system(s) to operate as follows, ostensibly referencing the Figure 2 steps shown in

[0035] First, the log generation process 200 directs the computing system(s) to identify a stream of log messages or statements related to the operating state, application, service, process, or any other aspect of a given computing environment (step 201). Examples include log statements that describe the start of a process, the end of a process, a process failure, the state of a database, the state of an application, or any other event that can be logged. The log messages can be identified by the same or different components that generate the messages.

[0036] Next, the log generation process 200 directs the computing system(s) to generate an encoded version of the message, hereinafter referred to as a log code (step 203). For example, the log code can be a hash value produced by a hash function that takes the log statement as input and produces a hash value as output. The computing system(s) can modify the log message prior to hashing the log message, e.g., by replacing unique identifying portions of the message with a common placeholder, such as a server name or a process identifier. This technique will allow the hash function to produce the same hash value for log messages that would be identical if they did not have their unique identifiers.

[0037] In some scenarios, the log code can be further encoded and / or compressed by replacing each of the various hash values produced by the hash function with a different symbol. For example, the hash function can output a string of numbers corresponding to a given log statement. The string of numbers can be replaced with a shorter alphanumeric symbol, thereby reducing the size of the log code.

[0038] Finally, the log generation process 200 instructs one or more computing systems to send log data to an exception service (e.g., exception service 110) that includes log codes (step 205). The dictionary can include log data that maps various log codes in their common form to their corresponding log statements. In scenarios where the log code is a compressed representation of a hash value, the dictionary can map the compressed representation to log statements and / or their corresponding hash values.

[0039] Figure 3 The extraction process 300 shown in can be implemented in program instructions within the context of any software application, module, component, or other such program element that includes the exception service 110. The program instructions instruct one or more underlying physical or virtual computing systems of the exception service 110 to operate as follows, ostensibly referring to the steps shown in Figure 3 the steps shown in.

[0040] The extraction process 300 begins by instructing one or more computing systems to receive log data from a computing node (step 301). The log data includes a sequence of information, such as log codes generated and timestamped by the computing node.

[0041] Next, the extraction process 300 instructs one or more computing systems to identify a dominant pattern in the sequence of information (step 303). Examples of dominant patterns include patterns of log codes that are consecutive with each other, which occur repeatedly within the sequence of information or have a certain frequency that is greater than the frequency of other less dominant patterns. For example, a snapshot of log data can include one thousand log codes, where a block of three consecutive log codes repeats hundreds of times. Such a pattern can be considered a dominant pattern and can thus be considered non - anomalous. That is, a dominant pattern may represent a pattern of log messages indicating a normal operating state rather than a state of exception that warrants an alert or other mitigation.

[0042] Figure 3 Steps 303A - 303D in represent a sub - process for identifying a dominant pattern in the sequence of information, but other techniques are possible and within the scope of the present disclosure. At step 303A, the extraction process 300 instructs one or more computing systems to identify a root block within the sequence of information. The root block can be, for example, the log code that appears most (or most frequently) within the sequence of information relative to any other log code.

[0043] Next, extraction process 300 instructs the (one or more) computing systems to grow the root block by at least one code or symbol at each position in the sequence where it occurs (step 303B) and compute a metadata length (MDL) score for the new root block (step 303C). The MDL score is computed by replacing the new root block at each position in the sequence with a symbol and then counting how much metadata is needed to describe the modified information sequence. In cases where the information sequence already includes symbols rather than the hash values themselves, the symbol replacing the new root block will replace the block of consecutive symbols that make up the new root block.

[0044] Then, extraction process 300 continues to compare the new value of the MDL score with the previous value of the MDL score before expanding the root block (step 303D). If the MDL score is less than the previous MDL score, then the sub - process advances to step 303B to grow the root block at each possible position in the sequence. If the root block cannot be expanded any further, then the process returns to step 303A to identify the next root block.

[0045] If the MDL score is not less than the previous MDL score, then the process also returns to step 303A to identify the next root block. If there are no more root blocks to identify and grow, then the sub - process returns the resulting non - anomalous blocks to the main process. Alternatively, the sub - process can return to the main process after finding a threshold number of non - anomalous blocks, even if there may be one or more root blocks remaining.

[0046] After identifying the non - anomalous blocks, extraction process 300 instructs the (one or more) computing systems to extract the non - anomalous blocks from the information sequence, thereby revealing the remaining blocks that can be considered anomalous (step 305). Then, extraction process 300 instructs the (one or more) computing systems to generate an anomaly vector based on the anomalous blocks (step 307) and distribute the vector to the computing environment to facilitate their detection of anomalous conditions.

[0047] In some embodiments, the revealed blocks can be submitted to an annotation and review process to confirm that a given block represents an anomaly. The annotation and review process can be performed manually by a human operator, autonomously, or through a combination of manual and autonomous reviews. In other embodiments, no review is required. In cases where a review process is employed, extraction process 300 serves to greatly reduce the amount of log statements to be manually reviewed, autonomously reviewed, or reviewed through a combination of manual and autonomous means. Optionally, the results of the review process can be provided as feedback to extraction process 300 to further enhance its ability to identify anomalous blocks. Blocks that are revealed as not being non - anomalous but are rejected as non - anomalous during the review phase can be re - classified as non - anomalous and identified to extraction process 300 for future use.

[0048] An abnormal condition can be identified by adopting a detection process 400. The detection process 400 can be implemented in program instructions in the context of any software application, module, component, or other such program elements of the computing environments 101 and 111. For example, the detection process 400 can be implemented in the context of the computing nodes 103 and 105 in the computing environment 101 and in the context of the computing nodes 113 and 115 in the computing environment 111. In addition to or instead of their respective computing nodes, the detection process 400 can also be implemented on any other computing element of the computing environments 101 and 111. The program instructions instruct the underlying physical or virtual computing system(s) to operate as follows, ostensibly referring to Figure 4 the steps shown in

[0049] The detection process 400 first instructs the computing system(s) to generate a hash value based on the log messages generated at runtime (step 401). Then, the process instructs the computing system(s) to generate a sequence vector based on the hash value (step 403). The sequence vector is a vectorized version of the sequence of hash values generated from the log messages. In other words, the hash values are put in a vectorized form. In some scenarios, the sequence vector can be generated directly from the log messages rather than the hash values or from a compressed representation of the hash values.

[0050] After generating the sequence vector, the detection process 400 instructs the computing system(s) to compare the sequence vector with each of the abnormal vectors in the set of abnormal vectors provided by the abnormal service (step 405). At step 407, the detection process 400 instructs the computing system(s) to determine whether the sequence vector matches one or more of the abnormal vectors. If so, an abnormality is detected (step 409). If not, the process continues with the next information sequence.

[0051] An exact match between the sequence vector and the abnormal vectors is not required for the abnormality to be identified. More precisely, an approximate match may be sufficient to be considered an abnormality. A similarity function can be employed in some scenarios to determine which of the abnormal vectors sufficiently matches the sequence vector. For example, a cosine similarity function can be applied to determine whether an abnormal vector sufficiently matches the sequence vector. In addition, it can be recognized that the detection process 400 proceeds continuously when generating new log messages.

[0052] Figure 5 A brief operation scenario 500 is included to illustrate the sequence of operations in the example. In operation, the computing environment 101 generates log codes and sends the log codes in the log data to the abnormal service 110. The abnormal service 110 also receives log data from the computing environment 111. The abnormal service 110 processes the log codes in the log data to identify abnormal blocks and generates abnormal vectors using the abnormal blocks.

[0053] An exception vector is sent by an exception service 110 to a computing environment 101 and a computing environment 111. The corresponding computing environment can use the exception vector for a stream of log messages to detect anomalies in the stream. In this way, various exception events such as server failures, processing failures, etc. can be detected and / or predicted.

[0054] Refer back to Figure 1 , operation scenarios 120, 130, and 140 illustrate example embodiments of a log generation process 200, an extraction process 300, and a detection process 400. It can be appreciated that while operation scenario 120 is illustrated with respect to computing environment 101, it can also occur in the context of computing environment 111. Similarly, operation scenario 140 is depicted with respect to computing environment 111, but operation scenario 140 can also occur in the context of computing environment 101.

[0055] In operation scenario 120, one or more computing nodes in computing environment 101 generate a stream 121 of log messages. The logs in the stream are identified by their line numbers such as line x - 1, line x, and line x + 1. A log generation process running in computing environment 101 performs an encoding process 123 on the log statements to generate log codes. For example, the process can anonymize each log statement by replacing server - or process - specific information with placeholder information. Then the anonymized log statements can be hashed to produce hash values, and the hash values are further encoded into symbols, words, or other such compressed representations of the hash values.

[0056] As an example, once anonymized, a total of k different log statements are possible. Thus, a deterministic encoding algorithm such as hashing will produce the same output code for each identical input statement. Thus, statement_1 will be encoded as code_1, statement_2 will be encoded as code_2, and so on until statement_k. The codes are then further encoded into symbols or words. For example, code_1 maps to alpha, code_2 maps to bravo, and so on until code_k, but any symbol is possible.

[0057] Then, the encoded stream 125 is sent from computing environment 101 to the exception service 110. The encoded stream 125 in this example includes at least one continuous block of symbols - echo, alpha, and bravo - which respectively trace back to code_4, code_2, and code_1, which themselves respectively trace back to anonymized statement_4, statement_1, and statement_2. The encoded stream can be sent along with other log data such as a dictionary, where the dictionary provides a mapping from symbols to codes and possibly to anonymized statements.

[0058] In operation scenario 130, the anomaly service 110 receives log data generated by the computing environment 101. The anomaly service 110 may also receive log data from the computing environment 111. The log data 131 represents data that can be received from the computing environment 101, the computing environment 111, or both.

[0059] The log data includes a log code flow mined by the anomaly service 110 for the anomaly block 133. For example, the anomaly service 110 identifies a non - anomaly block that includes the log codes alpha, bravo, delta. For illustrative purposes, assume that the block occurs with sufficient frequency (e.g., three times) such that it can be considered a dominant block.

[0060] Removing the (one or more) dominant blocks results in anomaly blocks 134 (echo) and anomaly blocks 135 (lima, zulu). The revealed blocks can optionally be submitted to an annotation and review process to confirm that a given block represents an anomaly. The annotation and review process can be performed manually by a human operator, autonomously, or through a combination of manual and autonomous reviews.

[0061] The anomaly service 110 continues to generate anomaly vectors based on the revealed anomaly blocks. Anomaly block 134 results in anomaly vector 138, while anomaly block 133 results in anomaly vector 139. Then, the anomaly service 110 can send the anomaly vectors to one or both of the computing environments 101 and 111 for deployment in the advancement of anomaly detection / prediction.

[0062] In operation scenario 130, the computing environment 111 has received the anomaly vectors generated by the anomaly service 110. Then, when a new log statement is generated, the computing environment 111 is able to vectorize the log stream and compare it with the anomaly vectors to determine if an anomalous condition exists.

[0063] The log stream 141 represents a log stream generated by one or more elements in the computing environment 111. A stream vector 143 is generated from the log stream 141 so that it can be compared with the anomaly vector 145. The system state 147 can be determined based on the result of the comparison. If there is a match between the stream vector 143 and one of the anomaly vectors 145, then an anomalous state or condition is considered to exist. If not, then the system state is considered normal or at least non - anomalous.

[0064] Figure 6Illustrates an operational scenario 600 in a more detailed example of log generation processing. In operation, a computing system generates a log message stream 601 related to the state of an information service. Each line of the log stream describes an event that has occurred regarding one or more elements of the service. For example, line_1 indicates that a new process has started. This process is identified by a process ID, and the element on which the process has started is identified by a host ID and a client ID. Line_2 describes the start of a different process on the same computing element as line_1.

[0065] Then, at line_3, a different event is described. The process described in line_2 is restarted by the combination of elements described in line_2. Line_4 describes the same process that has failed due to a network disconnection. Finally, line_5 describes the failure of the restart of the same process described in line_4.

[0066] It can be recognized that all five exemplary log messages belong to the same host / client combination but refer to two different processes. The host identifier uniquely identifies a machine from other machines in a computing environment, and the client identifier uniquely identifies a client from other clients on a machine. The process identifier also uniquely identifies a process related to other processes on a machine. However, the identifiers make it difficult to detect and / or predict abnormal conditions from log statements because they make each statement completely different from every other statement. (As mentioned herein, the term "machine" can represent a physical computer, a virtual machine, a container, or any other type of computing resource, variant, or combination thereof.)

[0067] To mitigate these effects, the log generation processing (e.g., 200) removes the process, machine, and client identifiers from a copy of the log message and replaces them with placeholder information. This step can be performed on all log messages generated by all other machines in a given computing environment, such that the aggregated stream of log messages is cleared of such differentiating identifiers. The original version of the log message with the identifiers intact can be retained for other purposes.

[0068] The modified log stream 603 provides an example where the process identifiers in lines 1 - 5 have been replaced by the acronym "PH", but any word or symbol is possible. Similarly, the combination of the host and client identifiers has been replaced by the same acronym. Overall, anonymizing or flattening log statements across all machines can increase their usefulness for anomaly detection and prediction.

[0069] The modified log stream 603 is then fed into a hash function 605. The hash function, which can be implemented in program instructions on a computing system, takes each line as input to generate a corresponding hash value. Since the hash function is deterministic, the same input always produces the same output. In combination with the modified log messages, this feature can produce the same hash value for multiple log statements that were initially different but are the same after modification. For example, line_1 and line_2 of the original log stream were different because of their processing IDs, but they are the same after replacing their processing IDs (and machine / client IDs) with placeholders. Thus, the hash values generated by the hash function 605 are the same (e.g., hash_1). The same result is produced for the hash value of line_5, but the hash values of line_3 and line_4 are different (hash_2, hash_3) because their corresponding log statements are different from the other statements after placeholder replacement.

[0070] Then, the encoded log stream 607 generated by the hash function 605 can be sent to an exception service (e.g., 110). The encoded log stream 607 includes hash values in an encoded / compressed form (e.g., hash_1, hash_2) relative to their actual values. A dictionary for decoding / decompressing the symbols can also be provided to the exception service. Optionally, the actual hash values themselves can be sent to the exception service.

[0071] Figure 7 An operational scenario 700 in a more detailed example of an exception block extraction process (e.g., 300) in an implementation is illustrated. In operation, the exception service receives one or more encoded log streams represented by an encoded log stream 701 from one or more computing environments. For illustrative purposes, the encoded log stream 701 includes five lines, each line having a log code including a symbolic / encoded representation of a hash value as discussed Figure 6 above.

[0072] The exception service feeds the log codes into a block extraction function 703, which can be implemented in program instructions on the (one or more) computing systems of the service. The block extraction function 703 examines the sequence of log codes to identify contiguous blocks of codes that occur more frequently than other codes. For example, for illustrative purposes, assume that the function finds a candidate block 705 that occurs 200 times in the examined log code sequence. For comparison purposes, a candidate block 707 is also found, but it only occurs once in the log code sequence.

[0073] Then the candidate block and its corresponding count are fed into an annotation function 709, which can also be implemented in program instructions on the service's (one or more) computing systems. The annotation function 709 determines whether to classify a given candidate block as a non-anomalous block or an anomalous block, at least in part based on its count. The candidate block 705 is classified as a non-anomalous block 711 and is thus removed from the sequence. The candidate block 707 is classified as an anomalous block 712 and is fed as an input into a vectorization function 713, which vectorizes the (one or more) log codes included in the anomalous block.

[0074] To vectorize the anomalous block 712, the vectorization function 713 identifies which log codes are included in the block. The vectorization function 713 can also be implemented in program instructions on the service's (one or more) computing systems. In this example, only one log code (hash_3) is included in the block. Then, the vectorization function 713 supplies a value to the position in the anomalous vector 715 corresponding to the log code represented in the block. For example, the third position from the right to the left in the anomalous vector 715 is supplied with the value "1" to indicate that the anomalous block 712 includes the log code "hash_3". For illustrative purposes, assume that the anomalous block also includes the log code "hash_5", then the fifth position in the anomalous vector will be supplied with the value "1". The remaining positions in the vector are supplied with the value "0" to indicate that no corresponding log code exists in the block.

[0075] Figure 8 An example scenario 800 of the block extraction process in an implementation is illustrated. In the example scenario 800, the processing table 801 steps through the order of operations in an example of identifying anomalous blocks in a log sequence, from top to bottom.

[0076] The first column of the processing table 801 describes the shorthand replacement of letter pairs for hash codes. The replacement does not need to be actually performed but is shown merely for explanatory purposes. The log code hash_1 is represented by A, hash_2 by B, and so on. For illustrative purposes, assume that there are twenty-six unique log codes in the encoded log stream, but any number of log codes is possible.

[0077] Next, the processing table 801 indicates the current root block being evaluated, as well as the characters and / or symbols in the log sequence. At the beginning, the sequence is analyzed to find the most frequent log code to set as the initial value of the root block (r-block). Using A as the root block, the metadata length (MDL) score of the root block is initially twenty-three.

[0078] Therefore, the evaluation continues to grow the root block A by one symbol at all its positions in the sequence. Thus, the first potential root block (p-block) is found to be AB. The consecutive symbols AB are replaced by the symbol α at all positions in the sequence, and the MDL score is calculated as nineteen. Since the MDL score of the potential block is less than the MDL score of the root block, the evaluation continues to replace the root block with the potential block.

[0079] Therefore, the root block becomes α, and the evaluation attempts to grow the root block at all positions where the root block is in the sequence. In cases where multiple potential blocks can be grown from the root block, the evaluation selects the most frequent potential block. For example, α can be grown by adding C at four positions and D at three positions. Thus, the evaluation proceeds to αC as the next potential block and restores α at its other positions to its original character set AB. Then, the potential block αC is replaced by the symbol λ at all positions where it is in the sequence, and an MDL score of seventeen is calculated. Since the MDL score of the potential block λ is less than the MDL score of the root block α, the evaluation continues to replace the root block with the potential block.

[0080] Then, the evaluation attempts to grow the root block at all positions of the root block. However, the root block λ cannot grow at any position in the sequence. Therefore, the root block λ is considered a non-anomalous block, and its characters or values can be removed from the sequence. Thus, the root block is restored to its previous value α and the evaluation starts again, but replaces ABC with λ at all positions.

[0081] In the case where the root block is set to α, the evaluation attempts to grow the root block at all its positions and identifies new potential blocks δ = αD or ABD. Then, the potential block ABD is replaced by the symbol δ at all positions where it is in the sequence, and an MDL score of fifteen is calculated. Since the MDL score of the potential block δ is less than the MDL score of the root block α, the evaluation continues to replace the root block with the potential block.

[0082] Then, the evaluation attempts to grow the root block δ at all its positions, and after determining that it cannot grow, continues to identify the next root block. The next root block is Y, whose MDL score is 14. Y can be grown at only one position by adding Z, and the MDL score of the potential block YZ is 16. Since 16 is greater than 14, YZ fails as a potential root block and is thus revealed as an anomalous block.

[0083] Then the anomalous block can be vectorized and distributed to the client for detecting and / or predicting an anomalous operation state. Figures 9A - 9C The figure illustrates one such scenario, ranging from the generation of an anomaly vector in an anomaly service to the deployment of the vector on the client. In Figure 9AIn the first part of 900A of the scenario shown, the exception service has identified an exception block 901 from the log code stream generated by one or more clients in one or more computing environments. The exception block 901 includes block 903, block 905, and block 907. Block 903 includes two log codes represented by hash_1 and hash_5. Block 905 includes only one log code represented by hash_3. Block 907 includes two log codes represented by hash_3 and hash_n.

[0084] The leftmost column of the vector table 909 lists all possible n hash codes from hash_1 to hash-n. The top row identifies the exception blocks from the first block b_1 to the kth block b_k. Each cell defined by a given hash code / block combination contains a value indicating whether the block includes the hash code. Zero indicates that the block does not include the hash code, while one indicates that the block does include the hash code. Regarding the exception block 901, block 603 is defined at the cells in the vector table 909 defined by hash_1 and b_1 and at the cells defined by hash_5 and b_1. The remaining cells in column b_1 have a zero value because block 903 does not contain other hash codes.

[0085] The vector values of the exception blocks b_2 and b_3 are filled in the same way. That is, block 905 is defined in the vector table 909 by a one at the cell defined by hash_3 and b_2 and zeros at all other positions. Block 907 is defined in the vector table 909 by a one at the cell defined by hash_3 and b_3 and at the cell defined by hash_n and b_3. The remaining cells in column b_3 have a zero value because block 907 does not contain other hash codes.

[0086] The resulting exception vector 911 is a string of binary values at positions corresponding to the values in the cells of each block in the vector table 909 from right to left. For example, the exception vector 913 corresponding to block 903 includes the string 0...0010001. The "1" in the rightmost position corresponds to the value in the cell defined by hash_1 and b_1, and the "1" in the fifth position from the right corresponds to the value in the cell defined by hash_5 and b_1. All other values in the exception vector 913 are "0".

[0087] The exception vector 915 corresponding to block 905 follows the same convention and includes the string 0...0000100. The "1" in the third position from the right corresponds to the value in the cell defined by hash_3 and b_2, and all other values are zero.

[0088] The exception vector 917 corresponding to block 907 includes the string 1...0000100. The "1" in the leftmost position corresponds to the value in the cell defined by hash_n and b_3, and the "1" in the third position from the right corresponds to the value in the cell defined by hash_3 and b_3. All other values in the exception vector 917 are "0".

[0089] The exception service distributes the vector to one or more computing environments to facilitate exception detection and prediction, as Figures 9A - 9C shown in the second part 900B of the described operating scenario.

[0090] Figure 9B including a log stream 921 representing a stream of log codes generated by the client computing system. The log codes in this example are hash values, which the client inputs into the vectorization function 923. The vectorization function vectorizes the log stream 921 to produce a vector stream 925. Similar to the vectorization process Figure 9A shown in, each log code can be vectorized by indicating a "1" at the position in the vector corresponding to the hash value represented by the log code.

[0091] As an example, the set of log codes in the middle of the log stream 921 is shown in bold, and their corresponding vectors are provided in the vector stream 925. According to this convention, the log code represented by hash_4 is transformed into a vector having a "1" at the fourth position from the right and zeros in all other positions. Thus, the log code represented by hash_3 is transformed into a vector having a "1" at the third position from the right and zeros in all other positions. The log code represented by hash_2 is transformed into a vector having a "1" at the second position from the right and zeros in all other positions. Finally, the log code represented by hash_n is transformed into a vector having a "1" at the nth position from the right and zeros in all other positions.

[0092] The client inputs the vector stream 925 into a weighting function 927. The weight function 927 applies the weighted vectors 928 to each individual vector in the vector stream 925. The weighted vectors 928 include weight values between 0 and 1 at each position in the vector, which correspond to the positions in the vectors of the vector stream 925. For example, a "1" in a position of the weighted vector 928 indicates that the corresponding position in the vector stream 925 should be given its full value, while a ".5" in a position of the weighted vector indicates that the corresponding position in the vector stream 925 is at half of its full value.

[0093] The weighting function 927 produces a weighted vector stream 929, which is the product of each position in the weighted vector 928 and its corresponding position in each vector of the vector stream 925. As an example, the middle vectors in the vector stream 925 and the weighted vector stream 929 are shown in bold to highlight the product of the weight value (.7) given to hash_3 multiplied by the value (1) at the hash_3 position in the middle vector. As another example, the second vector from the bottom in the weighted vector stream 929 from the vector stream 929 is shown in bold to highlight the zero weight given to hash_2 in the weighted vector 928.

[0094] Finally, Figure 9C illustrates Figures 9A - 9C the final part 900C of the operational scenario shown in. In Figure 9C the exception vectors 911 generated by the exception service have been distributed to the clients in their respective computing environments. Each client generates its own vector stream represented by the weighted vector stream 929.

[0095] The comparison function 931 receives the exception vector 911 and the weighted vector stream 929 as inputs and performs a comparison calculation of each vector in the weighted vector stream 929 for each vector in the exception vector 911. The comparison function can be, for example, a cosine similarity function, which evaluates the dot product of each exception vector multiplied by each vector in the vector stream. Then the comparison function analyzes the results of the calculation to determine whether one of the weighted exception vectors 911 includes a sufficient match with one or more of the exception vectors 911. If so, then an abnormal condition is considered to exist or be predicted. If not, then a normal state exists.

[0096] Figure 10 Illustrated is a computing system 1001, which represents any system or collection of systems in which the various processes, programs, services, and scenarios disclosed herein can be implemented. Examples of the computing system 1001 include, but are not limited to, server computers, routers, web servers, cloud computing platforms, and data center equipment, as well as any other type of physical or virtual server machine, physical or virtual router, container, and any variant or combination thereof.

[0097] The computing system 1001 can be implemented as a single device, system, or equipment, or can be implemented in a distributed manner as multiple devices, systems, or equipment. The computing system 1001 includes, but is not limited to, a processing system 1002, a storage system 1003, software 1005, a communication interface system 1007, and an optional user interface system 1009. The processing system 1002 is operatively coupled to the storage system 1003, the communication interface system 1007, and the user interface system 1009.

[0098] Processing system 1002 loads and executes software 1005 from storage system 1003. Software 1005 includes and implements exception processing 1006, which represents the processes discussed with respect to the preceding figures, including log generation process 200, extraction process 300, and detection process 400. When executed by processing system 1002 to provide block-based anomaly detection and prediction, software 1005 instructs processing system 1002 to operate as described herein for the various processes, operational scenarios, and sequences discussed in at least the preceding embodiments. Computing system 1001 may optionally include additional devices, features, or functionality not discussed for the sake of brevity.

[0099] Still refer to Figure 10 , processing system 1002 may include a microprocessor and other circuitry that retrieves and executes software 1005 from a storage system 1003. Processing system 1002 may be implemented within a single processing device, but may also be distributed across multiple processing devices or subsystems that cooperate to execute program instructions. Examples of processing system 1002 include general-purpose central processing units, graphics processing units, special-purpose processors and logic devices, and any other type of processing device, combination, or variation thereof.

[0100] The storage system 1003 may include any computer-readable storage medium that can be read by the processing system 1002 and that can store the software 1005. The storage system 1003 may include volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Examples of storage media include random access memory, read-only memory, magnetic disks, optical disks, optical media, flash memory, virtual and non-virtual memory, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other suitable storage medium. In any case, a computer-readable storage medium is not a propagated signal.

[0101] In addition to computer-readable storage media, in some embodiments, storage system 1003 may also include computer-readable communication media on which at least some of software 1005 may be transferred internally or externally. Storage system 1003 may be implemented as a single storage device, but may also be implemented across multiple storage devices or subsystems that are co-located or distributed relative to each other. Storage system 1003 may include additional elements, such as a controller, that can communicate with processing system 1002 or possibly other systems.

[0102] The software 1005 (including the exception handling 1006) can be implemented in program instructions and, among other functions, can particularly instruct the processing system 1002 to operate as described regarding the various operation scenarios, sequences, and processes herein when executed by the processing system 1002. For example, the software 1005 can include program instructions for implementing the log generation process, the exception block extraction process, and the exception detection process as described herein.

[0103] Specifically, the program instructions can include various components or modules that cooperate or otherwise interact to perform the various processes and operation scenarios described herein. The various components or modules can be implemented in compiled or interpreted instructions, or in some other variant or combination of instructions. The various components or modules can execute in a synchronous or asynchronous manner, serially or in parallel, in a single-threaded environment or in a multi-threaded environment, or according to any other suitable execution paradigm, variant, or combination thereof. The software 1005 can include additional processes, programs, or components, such as operating system software, virtualization software, or other application software. The software 1005 can also include firmware or some other form of machine-readable processing instructions executable by the processing system 1002.

[0104] Generally speaking, when loaded into the processing system 1002 and executed, the software 1005 can transform a suitable device, system, or apparatus (represented by the computing system 1001) from a general-purpose computing system as a whole into a dedicated computing system customized to provide block-based exception handling as described herein. In fact, the encoded software 1005 on the storage system 1003 can transform the physical structure of the storage system 1003. The specific transformation of the physical structure can depend on various factors in different embodiments of this description. Examples of these factors can include, but are not limited to, the technology used for implementing the storage medium of the storage system 1003 and whether the computer storage medium is characterized as a primary storage or a secondary storage device, and other factors.

[0105] For example, if the computer-readable storage medium is implemented as a semiconductor-based memory, the software 1005 can transform the physical state of the semiconductor memory when the program instructions are encoded therein, such as by transforming the states of transistors, capacitors, or other discrete circuit elements that make up the semiconductor memory. Similar transformations can occur for magnetic or optical media. Without departing from the scope of this specification, other transformations of the physical media are possible, and the foregoing examples are provided merely for the purpose of facilitating this discussion.

[0106] The communication interface system 1007 may include communication connections and devices that allow communication with other computing systems (not shown) via a communication network (not shown). Examples of the connections and devices that together allow inter-system communication may include network interface cards, antennas, power amplifiers, RF circuits, transceivers, and other communication circuitry. The connections and devices may communicate via a communication medium to exchange communications with other computing systems or a network of systems, such as metal, glass, air, or any other suitable communication medium. The media, connections, and devices mentioned above are well known and need not be discussed in detail here.

[0107] Communication between the computing system 1001 and other computing systems (not shown) may occur over one or more communication networks and in accordance with various communication protocols, combinations of protocols, or variants thereof. Examples include intranets, the Internet, the Internet, local area networks, wide area networks, wireless networks, wired networks, virtual networks, software-defined networks, data center buses and backplanes, or any other type of network, combination of networks, or variants thereof. The communication networks and protocols mentioned above are well known and need not be discussed in detail here.

[0108] As will be recognized by those skilled in the art, aspects of the present invention may be implemented as a system, method, or computer program product. Accordingly, aspects of the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment (including firmware, resident software, microcode, etc.), or an embodiment combining software and hardware aspects, which embodiments generally may all be referred to as “circuitry,” “module,” or “system.” Furthermore, aspects of the present invention may take the form of a computer program product implemented in one or more computer-readable media having computer-readable program code embodied thereon.

[0109] The included description and drawings depict specific embodiments that teach those skilled in the art how to make and use the best mode. Some conventional aspects have been simplified or omitted for the purpose of teaching the principles of the invention. Those skilled in the art will recognize variations of these embodiments that fall within the scope of the present disclosure. Those skilled in the art will also recognize that the above features may be combined in various ways to form multiple embodiments. Therefore, the present invention is not limited to the above specific embodiments, but is only limited by the claims and their equivalents.

Claims

1. A method for identifying anomalies in a computing environment, the method comprising: receiving log data from a plurality of nodes in a computing environment, wherein the log data includes a sequence of information indicative of log messages generated by the plurality of nodes; identifying, in the information sequence, a dominant pattern representing non-anomalous blocks of log messages; extracting non-anomalous blocks from the log data based on the dominant pattern to reveal abnormal blocks that do not conform to the dominant pattern; generating an exception vector based on at least the exception block; as well as Anomaly data is distributed to at least one node of the plurality of nodes to detect anomalies, wherein the anomaly data includes an anomaly vector.

2. The method according to claim 1, wherein Identifying dominant patterns within the information sequence includes: identifying underlying patterns within the information sequence; and A dominant pattern is selected from the potential patterns based at least on a scoring function applied to one or more of the potential patterns.

3. The method of claim 2, wherein the scoring function promotes a subset of latent patterns that occur frequently within the information sequence relative to a different subset of latent patterns that occur less frequently within the information sequence.

4. The method according to claim 3, wherein: For each of the latent patterns, the scoring function determines the relative dominance of the latent pattern when encoded using a compressed representation of the latent pattern based on a description length of the information sequence.

5. The method according to claim 4, wherein: The information sequence comprises a log code sequence, and wherein each log code in the log code sequence comprises a symbolic representation of a possible hash value, and wherein the log data comprises a codebook mapping each unique log code to a different hash value in a set of possible hash values.

6. The method according to claim 5, wherein: Each of the exception vectors includes a string of numbers corresponding by position to a set of possible log codes.

7. The method according to claim 6, wherein: Generating the exception vectors based on at least the exception block includes, for each exception vector, setting a value of each digit in the series of digits to indicate the presence or absence of a corresponding log code in a set of possible log codes in the exception block represented by the exception vector.

8. A computing device comprising: one or more computer-readable storage media; one or more processors operatively coupled to the one or more computer-readable storage media; as well as Program instructions stored on the one or more computer-readable storage media, which, when executed by the one or more processors, instruct the computing device to at least: receiving log data from a plurality of nodes in a computing environment, wherein the log data includes a sequence of information indicative of log messages generated by the plurality of nodes; identifying a dominant pattern in the information sequence that represents a high probability block of log messages; Extract high-probability blocks from log data based on dominant patterns to reveal low-probability blocks that do not conform to the dominant patterns; generating a low probability vector based on at least the low probability block; as well as Data is distributed to at least one node of the plurality of nodes to detect a low probability event, wherein the data includes a low probability vector.

9. The computing device of claim 8: in, The high probability blocks include non-anomalous blocks, the low probability blocks include abnormal blocks, the low probability vectors include abnormal vectors, the data include abnormal data, and the low probability events include abnormal events; and In which, in order to identify the dominant pattern within the information sequence, the program instructions instruct the computing device to identify potential patterns within the information sequence and select a dominant pattern from the potential patterns based at least on a scoring function applied to one or more of the potential patterns.

10. The computing device of claim 9, wherein the scoring function promotes a subset of latent patterns that occur frequently within the information sequence relative to a different subset of latent patterns that occur less frequently within the information sequence.

11. The computing device of claim 10, wherein: For each of the latent patterns, the scoring function determines the relative dominance of the latent pattern when encoded using a compressed representation of the latent pattern based on a description length of the information sequence.

12. The computing device of claim 11, wherein: The information sequence comprises a log code sequence, and wherein each log code in the log code sequence comprises a symbolic representation of a possible hash value, and wherein the log data comprises a codebook mapping each unique log code to a different hash value in a set of possible hash values.

13. The computing device of claim 12, wherein: Each of the exception vectors includes a string of numbers corresponding by position to a set of possible log codes.

14. The computing device of claim 13, wherein: Generating exception vectors based on at least the exception block includes, for each of the exception vectors, setting a value of each digit in the string to indicate the presence or absence of a corresponding log code in a set of possible log codes in the exception block represented by the exception vector.

15. A method for identifying anomalies, the method comprising: receiving log data from one or more nodes in a cloud computing environment, wherein the log data includes a log code sequence generated by the one or more nodes; identifying a code block within the log code sequence; Separate code blocks into at least normal and exception blocks; generating an exception vector based on at least the exception block; as well as Anomaly data is distributed to the one or more nodes to detect anomalies, wherein the anomaly data includes an anomaly vector.

16. The method of claim 15, wherein: Identifying code blocks within the logging code sequence includes: identifying potential code blocks within the logged code sequence; selecting a code block from the potential code blocks based at least on a scoring function applied to one or more of the potential code blocks; The scoring function improves the dominant subsequence of log codes within the log code sequence relative to the non-dominant subsequence of log codes.

17. The method of claim 16, wherein: Separating code blocks into at least normal and exception blocks involves: identifying a subset of code blocks that appear frequently within the log code sequence as normal blocks; and Different subsets of rarely occurring code blocks within the log code sequence are identified as exception blocks.

18. The method of claim 17, wherein: For each of the potential code blocks, the scoring function determines a relative dominance of the potential code block based on a description length of the log code sequence when encoded using the compressed representation of the potential code block.

19. The method of claim 18, wherein: Each log code in the sequence of log codes comprises a symbolic representation of a possible hash value, and wherein the log data comprises a codebook mapping each unique log code to a different hash value in a set of possible hash values.

20. The method of claim 19, wherein: Each of the exception vectors includes a string of numbers corresponding by position to a set of possible log codes; and Generating the exception vector based on at least the exception block includes setting a value of each digit in the series of digits to indicate the presence or absence of a corresponding log code in a set of possible log codes in the exception block represented by the exception vector.