Encoding log data using neural networks
By using a neural network trained with contrastive learning and similarity loss, log codes that can be shared across multiple tasks are generated, solving the problem of inconsistent log recording between different systems and improving log analysis efficiency and anomaly detection capabilities.
Patent Information
- Application Number
- CN202510549016.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-05-08
- Filing Date
- 2025-04-28
- Publication Date
- 2025-10-31
AI Technical Summary
Existing technologies lack a consistent logging standard across different systems, resulting in low efficiency in automatically parsing and analyzing log information. Furthermore, conventional training methods require a large amount of labeled data and a fixed vocabulary, making it impossible to share encodings across multiple tasks.
By employing contrastive learning and similarity loss training methods, a pre-trained neural network is used to minimize triplet loss and cosine similarity loss, generating log codes that can be shared across different tasks, and combining them with telemetry information for anomaly detection.
It enables code sharing across different tasks, reduces reliance on labeled data, improves the efficiency and accuracy of log analysis, and can effectively combine telemetry information for anomaly detection.
Smart Images

Figure CN120874752A_ABST
Abstract
Description
[0001] Priority requirements
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 640,061, filed April 29, 2024, entitled “Using Contrastive Learning to Train Neural Networks,” the entire contents of which are incorporated herein by reference.
[0003] For all purposes, the entire disclosure of the following co-pending applications is incorporated by reference: U.S. Patent Application No. 18 / 658,284 entitled “Using Contrastive Learning to Train Neural Networks”, U.S. Patent Application No. 18 / 658,324 entitled “Using Similarity Loss to Train Neural Networks”, and U.S. Patent Application No. 18 / 658,508 entitled “Using Neural Networks to Classify Logs”, all filed concurrently with this application. Technical Field
[0004] At least one embodiment relates to a neural network for encoding at least one log message. For example, at least one embodiment relates to encoding at least one log message at least in part by: encoding first type information in at least one log message to obtain a first encoding, encoding second type information in at least one log message to obtain a second encoding, and obtaining a resulting encoding at least in part by combining at least the first encoding and the second encoding. In at least one embodiment, a computing system (e.g., within a data center) implements various novel techniques described herein. Background Technology
[0005] System and / or service logs may include information related to these systems and / or services, such as descriptors of events that change over time and / or other useful information. The techniques used for logging are not consistently standardized across different systems (e.g., different domains with different terminology). This can make it challenging to automatically parse and / or analyze logs to extract and / or detect information contained within them that can be used for a variety of tasks. Automatic log parsing and / or analysis consumes significant amounts of memory, time, or computational resources. The amount of memory, time, sensor input, or computational resources used for automatic log parsing and / or analysis can be improved. Attached Figure Description
[0006] Figure 1 It is a block diagram illustrating a system for encoding and / or classifying log data according to at least one embodiment;
[0007] Figure 2 This is a block diagram illustrating a system for generating result encodings to encode at least one log message;
[0008] Figure 3 It is a block diagram illustrating a system for encoding at least one log message based at least in part on one or more types of information, according to at least one embodiment;
[0009] Figure 4 This is a flowchart illustrating the process of encoding the results of providing logs according to at least one embodiment;
[0010] Figure 5 This is a block diagram illustrating a system for training one or more neural networks to encode one or more logs, according to at least one embodiment.
[0011] Figure 6 This is a block diagram illustrating a system for training one or more converter encoders to encode one or more log sequences, according to at least one embodiment.
[0012] Figure 7 This is a block diagram illustrating a system for embedding vectors representing one or more logs according to at least one embodiment;
[0013] Figure 8 The diagram illustrates a system for training one or more neural networks based at least in part on triplet loss, according to at least one embodiment.
[0014] Figure 9 This is a flowchart illustrating the process of training a neural network to encode at least one vector associated with a log sequence, according to at least one embodiment.
[0015] Figure 10This is a block diagram illustrating a system for performing a neural network to classify one or more logs according to at least one embodiment;
[0016] Figure 11 An exemplary process for classifying at least one log by combining log information and telemetry information according to at least one embodiment is shown;
[0017] Figure 12 This is a flowchart illustrating the process of classifying one or more log messages according to at least one embodiment;
[0018] Figure 13 This is a flowchart illustrating a process for classifying one or more logs, at least in part, by an encoder using similarity loss to determine the classification, according to at least one embodiment.
[0019] Figure 14 This is a block diagram illustrating a system including an encoder for generating one or more logs based at least in part on a similarity loss, according to at least one embodiment.
[0020] Figure 15 The diagram illustrates a system for training one or more encoders based at least in part on cosine similarity loss, according to at least one embodiment.
[0021] Figure 16 This is a flowchart illustrating the process of determining and providing indicators for indicating similarity according to at least one embodiment;
[0022] Figure 17A An example of a system including a driver and / or runtime according to at least one embodiment is shown, wherein the driver and / or runtime includes one or more libraries for providing one or more application programming interfaces (APIs);
[0023] Figure 17B It is a block diagram illustrating examples of processors and modules according to at least one embodiment;
[0024] Figure 18A The logic according to at least one embodiment is shown;
[0025] Figure 18B The logic according to at least one embodiment is shown;
[0026] Figure 19 An example data center system according to at least one embodiment is shown;
[0027] Figure 20 This is a block diagram illustrating a computer system according to at least one embodiment;
[0028] Figure 21The training and deployment of a neural network according to at least one embodiment are illustrated;
[0029] Figure 22 Components of a system for accessing a large language model according to at least one embodiment are shown; and
[0030] Figure 23 This is a flowchart illustrating, according to at least one embodiment, the process of training a second neural network, at least in part, based on a first neural network, to encode a log sequence. Detailed Implementation
[0031] Various techniques have been described in the preceding and following sections. For illustrative purposes, specific configurations and details have been elaborated to provide a thorough understanding of the possible ways to implement these techniques. However, it will also be apparent that the techniques described below can be practiced in different configurations without specific details. Furthermore, well-known features may be omitted or simplified to avoid obscuring the described techniques.
[0032] Figure 1 This is a block diagram illustrating a system 100 for encoding, classifying, and / or otherwise processing log data according to at least one embodiment. System 100 may execute one or more neural networks (e.g., encoder 114, neural network NN1, neural network NN2, and / or classifier 122) for encoding and / or classifying log data. System 100 includes one or more processors 110 connected to memory 130 via one or more connections 134. In at least one embodiment, memory 130 (e.g., one or more non-transitory processor-readable media) stores machine-executable instructions 132 that, when executed by processor 110, implement topology function 106, telemetry function 108, preprocessing function 112, initial encoder function 113, encoder function 115, classification function 116, position encoder function 111, aggregation function 117, downstream function 118, and / or other functions. Processor 110 can receive or acquire input (e.g., one or more logs 104) and generate output 120 based at least in part on that input.
[0033] Logs of systems and / or services (e.g., within a data center) can include information related to these systems and / or services that changes over time, such as event descriptors. For example, if an event occurs, log messages or entries can be entered into one or more logs. Log sequences can include more than one log entry (e.g., concatenated together). One or more log entries can be stored as text including one or more letters, numbers, and / or symbols, combinations of which can indicate useful information (e.g., event description, timestamp, numeric counter value, identifier, etc.). Multiple combinations can be used to indicate different information in a single log. As an example, the information contained in logs can be used for various tasks such as anomaly detection, incident prediction, root cause analysis, and observation generation. However, the techniques for logging may not be uniformly standardized across different systems (e.g., different domain terminology), making it challenging to automatically parse and / or analyze logs to extract and / or detect the information contained within them.
[0034] As mentioned above, log entries can be stored as text and can include various types of information, such as one or more letters, one or more numbers, and / or one or more symbols. Furthermore, at least a portion of the log can be categorized into one or more distinct classes. If the system excludes one or more types of available data when encoding the log, one or more downstream processes (e.g., anomaly detection, event prediction, root cause analysis, and / or observation generation) may be negatively impacted because the encoding may ignore information useful to those downstream processes.
[0035] exist Figure 1In the example shown, log 104 includes text data 104A, numerical data 104B, and / or category data 104C. Preprocessing function 112 can remove information from log 104 that is not used for classifying and / or reformatting log 104 (e.g., changing letter case) (e.g., punctuation marks, spaces, etc.). Preprocessing function 112 can associate network devices or nodes (e.g., computing devices, routers, switches, etc.) with each log message included in log 104. For example, preprocessing function 112 can receive topology information from topology function 106, which includes node identifiers, and can associate the node identifiers with each log message. Preprocessing function 112 can divide each log line or entry into a separate data SD (e.g., stored in a separate data structure) for individual processing by initial encoder function 113. For example, preprocessing function 112 can create a data structure (e.g., a string) for each of text data 104A, numerical data 104B, and / or categorical data 104C, and provide one or more data structures to initial encoder function 113. Preprocessing function 112 can use one or more neural networks to partition each log line or entry into data SD.
[0036] Initial encoder function 113 encodes data SD to create initial encoding EL1 (e.g., one or more vectors), and the initial encoding EL1 is used as input to encoder function 115, which further encodes the initial encoding EL1 into encoding EL2 for use by classification function 116 (e.g., as input to one or more machine learning models, such as one or more neural networks NN2). Classification function 116 can use encoding EL2 to perform one or more tasks (e.g., anomaly detection, event prediction, root cause analysis, and / or observation generation). Initial encoder function 113 includes one or more encoders 114 (e.g., encoders 114A-114C) for encoding data SD to obtain the initial encoding EL1. Encoders 114 can be implemented using one or more neural networks. For example, one or more of encoders 114A-114C can be implemented using one or more neural networks.
[0037] The initial encoding EL1 of log 104 generated by initial encoder function 113 can encode the numerical data (e.g., numerical data 104B) and additional types of information (e.g., text data 104A, category data 104C, and / or other types of information) included in log 104. System 100 can encode log 104 using category data 104C (e.g., metadata).
[0038] In at least one embodiment, system 100 includes Figures 2 to 4One or more systems shown or otherwise are Figures 2 to 4 One or more systems are shown, for example, for executing process 400 (see...) Figure 4 In at least one embodiment, the initial encoder function 113 performs the process of encoding the text data, numerical data, and / or categorical data of each of one or more log entries using one or more encoders 114A-114C (e.g., Figure 4 The process 400 shown (e.g., in parallel, serial, or a combination of both) combines the outputs of these encoders 114A-114C to produce a uniform representation or encoding of log entries (e.g., initial encoding EL1). The text encoder 114A encodes any text data 104A information included in log 104, the numeric encoder 114B encodes information about any numeric data 104B in log 104, and the category encoder 114C encodes any category information (e.g., category data 104C), which may include metadata. Examples of metadata include event priority or message type.
[0039] An initial code EL1 generated by the initial encoder function 113 is provided to the encoder function 115. The position encoder function 111 may provide the position code POS of the log 104 to the encoder function 115. The encoder function 115 encodes the initial code EL1 and the position code POS to produce a code EL2 (e.g., one or more vectors). The encoder function 115 may use one or more neural networks NN1 (e.g., one or more transformer encoders) to produce the code EL2 at least in part based on the initial code EL1 and the position code POS.
[0040] Classification function 116 receives or obtains encoded EL2 and generates a classification or encoded EL3 (e.g., classifying encoded EL2 into one or more classes). Classification function 116 may use a neural network NN2 (e.g., one or more large language models (LLM)) to generate encoded EL3 at least in part based on encoded EL2.
[0041] Aggregation function 117 receives or obtains encoded EL3 and combines the information provided by topology function 106 and / or telemetry function 108 with the encoded EL3 (e.g., a classification indicating "IGNORE" or "ALERT") to create aggregated data AD. Aggregation function 117 may use one or more neural networks to generate aggregated data AD. Downstream function 118 receives or obtains aggregated data AD and generates output 120 based at least in part on the aggregated data AD. Downstream function 118 may use one or more neural networks to generate output 120.
[0042] While neural networks can be used to analyze logs, using supervised learning to analyze logs may require labeled training data. Creating such training data can be time-consuming and / or expensive, given the potentially vast amounts of information stored in logs. Self-supervised techniques can assume that logs do not contain anomalies, and if anomalies do appear in the training data, the model's performance may be negatively impacted. Self-supervised techniques may require a fixed vocabulary, and developers can add new messages to the logs that the model cannot encode. Furthermore, neither conventional supervised nor self-supervised training techniques may be able to encode entire sequences in a way that can be shared across multiple tasks, because such training techniques train neural networks to produce encodings specific to one or more specific tasks, and these encodings cannot generalize to other tasks. For example, encodings produced by a neural network trained to encode logs for anomaly detection may not be suitable for other tasks (e.g., event prediction) because different tasks require different labels for training. Therefore, techniques can be trained separately for each type of task using different training datasets, and / or using training datasets that include many labels for different tasks. However, the former technique can result in pre-trained neural networks being useless for other types of tasks (e.g., because it is trained on task-specific datasets), while the latter technique requires datasets with a large number of labels.
[0043] As described above, encoder function 115 can use a neural network NN1 (e.g., one or more transducer encoders) to generate encoded EL2 at least in part based on the initial encoded EL1 and the positional encoded POS. In at least one embodiment, system 100 includes Figures 5 to 9 One or more systems shown or otherwise are Figures 5 to 9 One or more systems are shown, for example, for executing process 900 (see Figure 9 In at least one embodiment, system 100 performs the process of pre-training neural network NN1 to encode an initial encoding EL1 (which encodes log 104) without a task-specific label, and using (e.g., minimizing) the triplet loss of the vector encoding produced by neural network NN1 (e.g., process 900, see process 900). Figure 9System 100 may use process 900 to pre-train neural network NN1 without using task-specific labels, such that after such pre-training, neural network NN1 can be easily trained to encode an initial encoding EL1 (which encodes logs) with minimal labeling for different types of tasks (e.g., as input to other neural networks (e.g., neural network NN2) and / or machine learning processes). System 100 performing process 900 may use a contrastive learning method that minimizes triplet loss to pre-train neural network NN1 to encode the initial encoding EL1 (which encodes log 104) at least in part based on a dataset with omitted or missing task-specific labels.
[0044] In at least one embodiment, system 100 includes Figures 14 to 16 One or more systems shown or otherwise are Figures 14 to 16 One or more systems are shown, for example, for executing process 1600 (see...). Figure 16 In at least one embodiment, system 100 performs a process of fine-tuning a neural network (e.g., neural network NN1 and / or neural network NN2) to encode encoded ELs (e.g., encoded EL1 and / or encoded EL2, which are the encodings of log 104) using similarity scores (e.g., process 1600, see process 1600). Figure 16The neural network (e.g., neural network NN1 and / or neural network NN2) can be trained and / or fine-tuned using (e.g., minimizing) the similarity loss of one or more vector codes generated by the neural network (e.g., code EL2 generated by neural network NN1 and / or code EL3 generated by neural network NN2). Since such a fine-tuned neural network (e.g., neural network NN1 and / or neural network NN2) can be trained using training data without task-specific labels, the neural network (e.g., neural network NN1 and / or neural network NN2) can then be trained to encode different types of tasks (e.g., as input to other neural networks and / or machine learning processes) using codes (code EL1 and / or code EL2, which encode log 104), with minimal labeling for small event sets. Process 1600 may include training on semantic similarity using one or more pairs of log entries and minimizing the cosine similarity loss. For example, neural network NN1 and / or neural network NN2 can be used to implement a log event classification model that can classify results as "ignored" or "alarmed". If previously unseen log entries are encoded, the log event classification model can classify the encoded unseen log entries into the same class as similar log entries that have been seen and encoded previously. Therefore, unlike self-supervised learning, a fixed vocabulary may not be required. Neural network NN1 can output encoded EL2 to neural network NN2, which can process encoded EL2 and output encoded EL3. The encoded EL3 output by neural network NN2 can be used as input to downstream processes (e.g., aggregation function 117, downstream function 118, one or more neural networks, etc.) that can classify the encoded log entries. For example, downstream function 118 can implement classifier 122, which can classify one or more log entries 104 as including evidence of anomalies.
[0045] While logs include information generated by software running on the computing system (e.g., error messages), telemetry includes information about the computing system itself, such as bit error rate (BER), CPU utilization, memory utilization, disk I / O, temperature, etc., as the computing system executes its software. Telemetry involves measuring data transmission from remote sources, such as physical or electrical data. Telemetry data can be collected using sensors or other devices, such as temperature sensors, counters (e.g., for counting anomalous events over time), etc. Both telemetry and logs include information that can be used to evaluate the computing system (data center), but current methods do not combine the use of logs and, for example, telemetry data to detect anomalies. Therefore, many current methods do not use at least some types of available data when performing anomaly detection, which negatively impacts the capabilities of downstream processes (e.g., event prediction, root cause analysis, and / or observation generation) because useful information may be lacking, making it impossible to detect anomalies.
[0046] In at least one embodiment, system 100 includes Figures 10 to 13 One or more systems shown or otherwise are Figures 10 to 13 One or more systems are shown, for example, for executing process 1100 (see...) Figure 11 ), 1200 (see Figure 12 ), 1300 (see Figure 13 (or a combination thereof). In at least one embodiment, system 100 performs at least one or more parts (e.g., processes 1100, 1200, and / or 1300; see below) of a process that combines log information (e.g., in the form of encoded EL3) and telemetry information (e.g., combining encoded vectors or aggregated data AD). Figures 11 to 13 For example, aggregation function 117 may use one or more neural networks (e.g., encoders) to aggregate information received from topology function 106, telemetry function 108, and / or encoding EL3, and then provide the aggregated data AD to downstream function 118, which may, for example, use the aggregated data AD as input to one or more neural networks (e.g., classifier 122) to detect one or more anomalies within a computing system (e.g., a data center). Topology function 106 may encode network topology information (e.g., devices, physical connections, or locations) in combination with or separately from telemetry information (e.g., provided by telemetry function 108). Topology function 106 may provide the encoded topology information to aggregation function 117 as part of performing anomaly detection. For example, when system 100 is used to perform anomaly detection, system 100 may be characterized as or implement an anomaly detection pipeline. Topology function 106 and / or telemetry function 108 may receive information from one or more external data sources and provide such information to aggregation function 117.
[0047] In at least one embodiment, system 100 includes a collection of one or more hardware and / or software computing resources having instructions that, when executed, perform one or more communication processes, such as those described herein. In at least one embodiment, system 100 is a software program, an application program, and / or variations thereof, executing on computer hardware. In at least one embodiment, one or more processes of system 100 are executed by any suitable processing system or unit (e.g., a graphics processing unit (GPU), a general-purpose GPU (GPGPU), a parallel processing unit (PPU), a central processing unit (CPU), a data processing unit (DPU) (described below)) in any suitable manner, including sequentially, in parallel, and / or variations thereof. In at least one embodiment, system 100 uses machine learning training frameworks (e.g., PYTORCH, TENSORFLOW, BOOST, CAFFE, MICROSOFT COGNITIVE TOOLKIT / CNTK, MXNET, CHAINER, KERAS, DEEPLEARNING4J) and / or other training frameworks to implement and perform the operations described herein to encode and / or classify log data and / or otherwise perform the operations described herein. In at least one embodiment, as an example, training a neural network model includes using a server (e.g., an NVIDIA DGX server) that further includes at least a GPU (e.g., AMD MI200, VEGAL10, VEGO20, and ARCTURUS), an optimizer (e.g., ADAMOPTIMIZER), or a discriminator architecture (e.g., a converter encoder architecture using sentence embeddings of a converter-based sentence bidirectional encoder representation (SBERT) trained at least partially with cosine similarity loss, or a discriminator architecture trained using one or more loss operations described herein).
[0048] In at least one embodiment, system 100 includes modules (e.g., modules 1724-1730, see below). Figure 17BThe system 100 executes a neural network to encode and / or classify log data. In at least one embodiment, the module includes any combination of any type of logic (e.g., software, hardware, firmware) and / or circuitry configured to perform the function. In at least one embodiment, the module includes one or more circuits that form part of a larger system (e.g., an integrated circuit (IC), a system-on-a-chip (SoC), a central processing unit (CPU), a graphics processing unit (GPU), a data processing unit (DPU), etc.). In at least one embodiment, the controller includes any combination of any type of logic (e.g., software, hardware, firmware) and / or circuitry configured to perform the function. In at least one embodiment, the software includes software packages, code, programming languages, drivers, instructions, instruction sets, or some combination thereof. In at least one embodiment, the hardware includes hardwired circuitry, programmable circuitry, state machine circuitry, fixed-function circuitry, execution unit circuitry, firmware storing instructions executed by the programmable circuitry, or some combination thereof.
[0049] In at least one embodiment, system 100 includes one or more logic units. In at least one embodiment, the logic unit includes firmware logic, hardware logic, or some combination thereof, configured to provide any of the functions further described herein. In at least one embodiment, the logic unit includes circuitry that forms part of a larger system (e.g., IC, SoC, CPU, GPU, DPU). In at least one embodiment, the logic unit includes logic circuitry for implementing firmware and / or hardware to execute a neural network to encode and / or classify log data.
[0050] In at least one embodiment, system 100 includes one or more engines. In at least one embodiment, an engine includes modules and / or logic units as further described herein. In at least one embodiment, a component includes modules and / or logic units as further described herein. In at least one embodiment, an engine includes software logic, firmware logic, hardware logic, or some combination thereof, configured to provide any functionality as further described herein. In at least one embodiment, a component includes software logic, firmware logic, hardware logic, or some combination thereof, configured to provide any functionality as further described herein. In at least one embodiment, operations performed by hardware and / or firmware may alternatively be implemented via software modules, which may be embodied as software packages, code, and / or instruction sets. In at least one embodiment, a logic unit may also utilize a portion of software to implement its functionality.
[0051] In at least one embodiment, system 100 includes one or more processors 110 for executing one or more neural networks, such as neural network NN1, neural network NN2, classifier 122, and / or others. Processor 110 may receive one or more inputs 102, such as one or more logs 104, topology information provided to topology function 106 by one or more topology data sources, and / or telemetry information provided to telemetry function 108 by one or more telemetry data sources. Inputs 102 may include one or more inputs 202, 302, 1402, and / or 1502 (see respective examples). Figure 2 , Figure 3 , Figure 14 and Figure 15 The input 102 of one or more logs 104 may include one or more logs 204, log lines 304, log sequence 508, tags 610 representing one or more log event codes (e.g., which form or define log sequence 612) and position codes (e.g., which form or define position sequence 614), logs 702 and / or embedding vectors 704, log line stream 1002A, input to the original log 1102, log line 1404, log line pair 1504A, or combinations thereof (see [link to documentation]). Figures 2 to 16 One or more logs 104 may include information such as text data 104A (e.g., text data 206A and / or 306A), numerical data 104B (e.g., numerical data 206B and / or 306B), and / or category data 104C (e.g., category data 206C and / or 306C). Input 102 for topology information may include topology data 1012 and / or topology and metadata information 1002C. Input 102 for telemetry information 108 may include topology data 1012.
[0052] Processor 110 may execute one or more neural networks, such as encoder 114, neural network NN1, neural network NN2, classifier 122, and / or one or more others. One or more encoders 114 may include text encoder 114A (e.g., encoder 208 and / or text encoder 308A), numerical encoder 114B (e.g., encoder 208 and / or numerical encoder 308B), class encoder 114C (e.g., encoder 208 and / or class encoder 308C), neural network NN1 trained using triplet loss with one or more vector encodings (e.g., model 510 and / or neural network 608), and / or neural network NN1 trained using similarity loss with one or more vector encodings (e.g., log event classification model 1006, encoder 1408, and / or encoder 1512).
[0053] Processor 110 may execute encoder 114, neural network NN1, neural network NN2, classifier 122, and / or one or more of the others to generate one or more log codes EL1, EL2, and / or EL3, which may include one or more vector codes, vectors of combined codes, result codes 216, result codes 314, embedding vectors 704, generated semantic codes 1412, vector codes 1516, or combinations thereof. In at least one embodiment, the vector code is a tensor representing information (e.g., information type) associated with one or more logs.
[0054] Processor 110 may execute one or more neural networks (e.g., neural network NN1 and / or neural network NN2), which may include one or more classifiers, such as log event classification model 1006, encoder 1408, encoder 1512, and / or LLM. The classifier (e.g., neural network NN2) may perform one or more tasks, such as anomaly detection (e.g., anomaly detection 1414A and / or model 1014), event prediction (e.g., event prediction 1414B), root cause analysis (e.g., root cause analysis 1016 and / or 1414C), observation generation (e.g., observation generation 1414D), and / or one or more other downstream tasks and / or applications 1414 described herein (see [link to documentation]). Figure 14 The processor 110 executing one or more neural networks can generate one or more outputs 120, such as classifications associated with one or more inputs 102 and / or one or more outputs as described herein.
[0055] In at least one embodiment, processor 110 includes one or more circuits that execute at least a portion of instructions 132 stored in memory 130 (e.g., implementing encoder 114, neural network NN1, neural network NN2, classifier 122, other machine learning processes, topology function 106, telemetry function 108, preprocessing function 112, initial encoder function 113, encoder function 115, classification function 116, position encoder function 111, aggregation function 117, downstream function 118, and / or other functions). In at least one embodiment, processor 110 includes one or more parallel processing units (“PPUs”), such as one or more graphics processing units (“GPUs”), one or more massively parallel GPUs, one or more accelerators, and / or others. In at least one embodiment, a massively parallel GPU refers to one or more GPUs or any suitable collection of processing units that can be used to execute various processes in parallel. In at least one embodiment, the processor 110 is implemented using, for example, a main central processing unit (“CPU”) complex, one or more microprocessors, one or more microcontrollers, a PPU (e.g., an accelerator, GPU, and / or others), one or more data processing units (“DPU”), one or more arithmetic logic units (“ALU”), and / or others. In at least one embodiment, one or more processors 110 are implemented using... Figures 17A to 22 The image shown and / or about Figures 17A to 22 This is implemented using one or more devices as described. In at least one embodiment, any circuitry used to implement one or more processors 110 is implemented using... Figures 17A to 22 The image shown and / or about Figures 17A to 22 It can be implemented using any circuit described.
[0056] In at least one embodiment, memory 130 (e.g., one or more non-transitory processor-readable media) is implemented using, for example, volatile memory (e.g., dynamic random access memory (“DRAM”)) and / or non-volatile memory (e.g., hard disk drive, solid-state drive (“SSD”) and / or others). In at least one embodiment, memory 130 (e.g., one or more non-transitory processor-readable media) is implemented using... Figures 17A to 22 Shown and / or about Figures 17A to 22 It is implemented using one or more memory devices as described.
[0057] In at least one embodiment, the memory 130 and the processor 110 communicate with each other via one or more connections 134 (e.g., buses, peripheral component interconnect rapid (“PCIe”) connections (or buses) and / or others). In at least one embodiment, the connection 134 is used in… Figures 17A to 22Shown and / or about Figures 17A to 22 It is implemented by describing one or more structures.
[0058] In at least one embodiment, system 100 includes one or more processors for executing one or more neural networks to encode one or more logs, classify one or more logs, and / or otherwise perform the operations described herein. In at least one embodiment, system 100 is included Figures 1 to 17B and / or Figure 23 The system shown and / or otherwise includes Figures 1 to 17B and / or Figure 23 The system shown is used to execute one or more neural networks to encode one or more logs, classify one or more logs, and / or otherwise perform the operations described herein. In at least one embodiment, system 100 performs... Figures 1 to 17B and / or Figure 23 The processes shown herein include, for example, one or more neural networks for encoding one or more logs, classifying one or more logs, and / or otherwise performing the operations described herein. In at least one embodiment, system 100 includes... Figures 17A to 22 The hardware shown may be, for example, used to execute one or more neural networks to encode one or more logs, classify one or more logs, and / or otherwise perform the operations described herein.
[0059] Figure 2 This is a block diagram illustrating a system 200 for generating result encodings to encode at least one log message. System 200 may be implemented at least in part by an initial encoder function 113. In at least one embodiment, system 200 includes one or more encoders 208 (e.g., Figure 1 The encoders 114A-114C shown can generate one or more codes 210 (e.g., codes 210A-210C). Encoder 208 may include one or more attention layers 212 and / or communicate with one or more attention layers 212 to combine the codes 210A-210C generated by encoder 208, for example, to generate a resulting code 216 (e.g., initial code EL1, which may be a vector of combined codes). One or more codes 210 may each correspond to a type of information 206 included in log 204. For example, a type of information 206 (e.g., text data 206A) may correspond to one or more encoders 208 that generated the first code 210A.
[0060] System 200 may use at least one neural network (e.g., one or more encoders 208) to encode text data 206A, numerical data 206B, and / or category data 206C of one or more log entries (e.g., obtained as data SD) to generate a first code 210A, a second code 210B, and an Nth code 210C, and combine the outputs of the neural network to produce a result code 216. For example, the result code 216 may be a uniform representation or encoding of the log entries (e.g., obtained as data SD). The first encoder in encoder 208 may be a text encoder (e.g., text encoder 308A and / or 114A) that can encode any text information; the second encoder may be a numerical encoder (e.g., numerical encoder 308B and / or 114B) that can encode information related to any numerical value in the log; and the third encoder may be a category encoder (e.g., category encoder 308C and / or 114C) that can encode any category information, which may include metadata. Examples of metadata are event priority or message type.
[0061] Before using the first, second, and third encoders, preprocessing operations (e.g., performed by preprocessing function 112) can partition log entries into separate data (e.g., data SDs) based on the type of information 206 (e.g., text data 206A and numeric data 206B). For example, the preprocessing operation can duplicate the log entry, remove numeric data from a first copy of the log entry to create text data, and remove text data from a second copy of the log entry to create numeric data. The preprocessing operation can also identify any categories and / or metadata associated with the log entry as category data. Category data can be stored in a data structure (e.g., strings, arrays, etc.). Data SDs can include separate text data, numeric data, and / or category data.
[0062] As an example, one or more encoders 208 may include a semantic encoder (e.g., a sentence converter pre-trained on text) that receives text data 106A (e.g., included in data SD) obtained from entries in log 204 through preprocessing operations and generates encoding 210 by encoding text segments within text data 106A into information related to the natural language in text data 106A. Log entries may include descriptors of events over a period of time. A non-limiting example of such a descriptor is “INFO dfs.DataBlock Scanner:Verification Succeeded for…”, which may be divided into text segments to be encoded, such as “dfs” and “DataBlockScanner”. The encoding output by the semantic encoder includes values representing each text segment in the text data, which are combined to define a vector representing text data 106A, such as one of encodings 210.
[0063] As an example, one or more encoders 208 may include a sine encoder that receives numerical data 106B (e.g., included in data SD) obtained from entries in log 204 through preprocessing operations and encodes the numerical values (e.g., timestamps, counters, object identifiers, etc.) within the numerical data 106B using a sine function (e.g., sine and / or cosine), such as one of encodings 210. As an example, the sine encoder may encode location information using one or more sine functions and / or one or more cosine functions. The sine encoder may represent timestamps, counters, or other time-series data as one of encodings 210. Time-series information may be encoded using scaling and / or quantization, for example by one or more time-series prediction models (e.g., Chronos) and / or by extracting one or more Fourier features and applying one or more neural network layers. The code 210 generated by the sine encoder (e.g., one or more encoders 208, 114B and / or 308B) includes a value representing each value in the numerical data, which is combined to define a vector representing the numerical data 206B.
[0064] As an example, one or more encoders 208 may include an embedding encoder that receives category data 206C obtained from entries in log 204 through a preprocessing operation and encodes the category data 206C into a vector, such as one of encodings 210. As an example, the category data 206C may be ordinal data, where an ordered relationship exists (e.g., "first", "second", and "third"). The category data 206C may include one or more labels. Labels can be encoded by mapping labels to integers (e.g., integer encoding), mapping labels to binary vectors (e.g., one-hot encoding), or learning embeddings (e.g., a distributed representation of the category). As an example, one or more encoders 208 generate vector embeddings for priority information (e.g., one of encodings 210) included in log entries. As an example, entries in log 204 may include text descriptors (e.g., INFO or WARN) associated with an event that can be categorized by an anomaly detector into a priority level (e.g., low priority or high priority) (e.g., "INFO" = low priority, "WARN" = high priority) (e.g., categorized into category data 206C). If a log entry includes the text "INFO", the preprocessing operation may include information in category data 206C indicating low priority, and an embedding (or one of encodings 210) generated by an embedding encoder (e.g., one or more encoders 208) indicating a low-priority classification. If log 204 includes the text "WARN", the preprocessing operation may include information in category data 206C indicating high priority, and an embedding (e.g., one of encodings 210) generated by an embedding encoder indicating a high-priority classification.
[0065] Once text data 206A, numerical data 206B, and category data 206C are encoded into one or more codes 120A-120C, the individual outputs of each of the three encoders can be fed to attention layer 212 (e.g., a single attention layer of a converter encoder) to combine (e.g., fuse) one or more outputs of encoder 208 (e.g., codes 210A-210C). Attention layer 212 can assign one or more weights to each feature embedding in one or more outputs of encoder 208 (e.g., codes 210A-210C) and use these weights to compute result code 216 (e.g., a weighted average of one or more outputs). As an example, the output 214 of attention layer 212 (which includes a vector) (e.g., initial code EL1) can be provided to and used by downstream processes. For example, result code 216 can be used to generate training data to train one or more neural networks, such as neural network NN1 and / or... Figures 5 to 9 The encoder shown.
[0066] In at least one embodiment, system 200 includes one or more processors for encoding at least one log message at least partially by: encoding information of a first type in the at least one log message to obtain a first encoding; encoding information of a second type in the at least one log message to obtain a second encoding; obtaining a result encoding at least partially by combining the first encoding and the second encoding; and / or otherwise performing the operations described herein. In at least one embodiment, system 200 is included Figures 1 to 17B and / or Figure 23 The system shown and / or otherwise includes Figures 1 to 17B and / or Figure 23 The system shown is used to encode at least one log message at least partially by: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code at least partially by combining at least the first code and the second code; and / or otherwise performing the operations described herein. In at least one embodiment, system 200 performs... Figures 1 to 17B and / or Figure 23 The processes shown herein include, for example, encoding at least one log message by at least partially: encoding information of a first type in at least one log message to obtain a first encoding; encoding information of a second type in at least one log message to obtain a second encoding; obtaining a result encoding by at least partially combining the first encoding and the second encoding; and / or otherwise performing the operations described herein. In at least one embodiment, system 200 includes... Figures 17A to 22 One or more pieces of hardware shown herein are used, for example, to encode at least one log message at least in part by: encoding information of a first type in at least one log message to obtain a first encoding; encoding information of a second type in at least one log message to obtain a second encoding; obtaining a result encoding at least in part by combining at least the first encoding and the second encoding; and / or otherwise performing the operations described herein.
[0067] Figure 3A block diagram of a system 300, according to at least one embodiment, is shown that encodes at least one log message based at least in part on one or more types of information. System 300 may be implemented at least in part by an initial encoder function 113. Logs can provide a rich source of information about the lifecycle of systems and services. The large scale of log generation and its inherent characteristics (e.g., lack of standardization and use of domain-specific terminology) can make manually extracting meaningful insights challenging. Encoding log lines in a way that captures semantic meaning and relationships can improve the performance of downstream log analysis tasks that operate on individual log lines (e.g., a cluster of one or more log messages) and / or log sequences (e.g., a combination of log messages). When encoding a single log line, one or more types of information 306 reported in the log line (e.g., numerical data 306B) may be ignored, or a separate model may be used to analyze each type of information 306, which may not take into account categorical data 306C, such as event priority or event type.
[0068] In at least one embodiment, system 300 executes a general feature tokenization model that operates on and integrates different types of log line information 306, which may include: "clean" text 306A (e.g., message / template), numerical data 306B (e.g., duration, telemetry reported in the log), and categorical data 306C (e.g., event priority, event type). System 300 may include an encoder 308 for encoding one or more types of information (e.g., feature types to be identified) using a dedicated encoding model, and then fusing them with an attention-based layer 310 (e.g., a single layer of a converter encoder). System 200 may be coupled with a log-based analytics model to provide a more complete encoding of log information.
[0069] System 300 may receive one or more inputs 302, such as log lines 304. In at least one embodiment, one or more inputs 302 of system 300 may include one or more characters from log lines 304, one or more log lines, one or more sequences of log lines, one or more encodings of one or more log lines (e.g., vectors representing said log lines), text, symbols, previous input, one or more scripts for training one or more neural networks, information represented as data, and / or other inputs described herein. In at least one embodiment, one or more inputs 302 are transmitted to one or more processors via signals. In at least one embodiment, one or more inputs 302 are information represented as one or more data packets. In at least one embodiment, one of the inputs 302 is received by a software process, for example, in conjunction with... Figures 1 to 16The software process described in any of the diagrams. In at least one embodiment, at least one of the inputs 302 is received by one or more pieces of hardware, such as in conjunction with... Figures 17A to 22 The hardware described by any of the diagrams in the diagram.
[0070] Log line 304 may include one or more information types 306, such as text 306A, numeric data 306B, and / or categorical data 306C (e.g., metadata 306D). As an example, one or more encoders may correspond to one of the information types 306 and may be used to encode log line 304. For example, Figure 3 One or more encoders 308 (e.g., text encoder 308A, numeric encoder 308B, and category encoder 308C) are shown for information type 306. Each encoder 308 corresponding to one of the information types 306 can generate an encoding corresponding to that information type in log line 304. As a non-limiting example, text encoder 308A generates an encoding corresponding to text 306A, numeric encoder 308B generates an encoding corresponding to numeric data 306B, category encoder 308C generates an encoding corresponding to category data 306C, metadata encoder generates an encoding corresponding to metadata, and / or one or more other encoders can each generate encodings for other types of information. Each encoding generated by encoder 308 is combined (e.g., fused) by attention layer 310 into a resulting encoding 314 (e.g., a vector representing the combined encodings, such as a mean or weighted mean). In at least one embodiment, encoder 308 (e.g., text encoder 308A, numeric encoder 308B, and / or category encoder 308C) includes one or more encoders 114 and / or 208. One or more encodings generated by encoder 308 may include one or more encodings 210. Attention-based layer 310 may be implemented by one or more attention layers 212, which may generate outputs 312 and / or 214.
[0071] System 300 may generate and provide one or more outputs 312, such as result encoding 314 (e.g., initial encoding EL1). In at least one embodiment, one or more outputs 312 of system 300 may include one or more embeddings (e.g., encodings) representing one or more information types 306 included in log lines 304, one or more tensors (e.g., vectors), one or more log lines 304 (e.g., log messages), one or more sequences of log lines 304, one or more encodings of one or more log lines 304 (e.g., vectors representing said log lines), text, symbols, previous inputs, one or more weights, one or more representations of the logs, one or more classifications of the logs, information represented as data, and / or other outputs described herein. In at least one embodiment, one or more outputs 312 are transmitted to one or more processors via signals. In at least one embodiment, one or more outputs 312 are information represented as one or more data packets. In at least one embodiment, at least one output 312 is generated by a software process, for example, in conjunction with... Figures 1 to 16 The software process described in any of the diagrams. In at least one embodiment, at least one output 312 is generated or received by one or more hardware components, such as in conjunction with... Figures 17A to 22 The hardware described by any of the diagrams in the diagram.
[0072] In at least one embodiment, system 300 includes one or more processors configured to encode at least one log message at least partially by: encoding information of a first type in the at least one log message to obtain a first encoding; encoding information of a second type in the log message to obtain a second encoding; obtaining a result encoding at least partially by combining at least the first encoding and the second encoding; and / or otherwise performing the operations described herein. In at least one embodiment, system 300 is included Figures 1 to 17B and / or Figure 23 The system shown and / or otherwise includes Figures 1 to 17B and / or Figure 23 The system shown is used to encode at least one log message at least partially by: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code at least partially by combining at least the first code and the second code; and / or otherwise performing the operations described herein. In at least one embodiment, system 300 performs... Figures 1 to 17B and / or Figure 23The processes shown herein, for example, are used to encode at least one log message at least partially by: encoding information of a first type in at least one log message to obtain a first encoding; encoding information of a second type in at least one log message to obtain a second encoding; obtaining a result encoding at least partially by combining at least the first encoding and the second encoding; and / or otherwise performing the operations described herein. In at least one embodiment, system 300 includes Figures 17A to 22 One or more pieces of hardware shown herein are used, for example, to encode at least one log message at least in part by: encoding information of a first type in at least one log message to obtain a first encoding; encoding information of a second type in at least one log message to obtain a second encoding; obtaining a result encoding at least in part by combining at least the first encoding and the second encoding; and / or otherwise performing the operations described herein.
[0073] Figure 4 This is a flowchart illustrating a process 400 for providing the result encoding of logs according to at least one embodiment. Process 400 may be executed at least in part by an initial encoder function 113 (e.g., when executed by processor 110). Process 400 may begin when otherwise invoked by one or more processors (e.g., processor 110), and / or when the initial encoder function 113 receives one or more logs as input in block 402. The logs received as input in block 402 may be received in combination with one or more inputs 102, 202, and / or 302. One or more systems (e.g., systems 100, 200, and / or 300) may execute process 400, such as to jointly encode different types of data, such as text, numeric, categorical, and / or metadata. Process 400 may include encoding text, numeric log data, categorical log data, and metadata appended to one or more logs using feature tokenizers.
[0074] Upon receiving the log input in box 402, the initial encoder function 113 may attempt to identify the relevant information type (e.g., text, numeric, category, or metadata) in the log input. The initial encoder function 113 may then proceed to decision box 404. In decision box 404, the initial encoder function 113 determines whether the relevant type of information (e.g., text, numeric, category, or metadata) has been identified in the log input. If the relevant type of information (e.g., text, numeric, category, or metadata) is identified, the decision in decision box 404 may result in "yes"; otherwise, the decision in decision box 404 may result in "no". If the decision in decision box 404 is "yes", the initial encoder function 113 may generate one or more codes in box 406, such as codes corresponding to the identified information type. After generating the codes in box 406, the initial encoder function 113 may then proceed to decision box 404 to determine whether another relevant type of information (e.g., text, numeric, category, or metadata) can be identified in the log. As an example, text information is identified, and the processor performing the initial encoder function 113 in box 406 generates an encoding corresponding to the text information. Then, continuing from this example, the processor performing the initial encoder function 113 returns to decision box 404 to determine whether other relevant types of information, such as numerical information, have been identified. This can be repeated until encodings are generated for each relevant type of information included in the log. Boxes 404 and 406 can also be performed in parallel for each type of information to be identified.
[0075] If the decision in decision box 404 is "No", the initial encoder function 113 can proceed to decision box 408. If one or more results have been obtained, the decision in decision box 408 can be "Yes". For example, if at least one code has been generated in box 406, a result is obtained. If the decision in the decision box is "Yes", the initial encoder function 113 provides the result code (e.g., initial code EL1) in box 410. If more than one code has been generated (e.g., during multiple iterations in box 406), the initial encoder function 113 can combine the codes generated in box 406 to obtain the result code provided in box 410. For example, the initial encoder function 113 can combine one or more codes generated in box 406 by calculating the mean or weighted mean of the codes generated in box 406. In at least one embodiment, the result code is result code 216 and / or 314. If, during execution of process 400, the initial encoder function 113 executes box 410 by providing a result encoding (e.g., to the processor), then the initial encoder function 113 may proceed to the point where execution of one or more operations and / or process 400 as described herein may terminate. If the decision in decision box 408 is "No", then the initial encoder function 113 may proceed to the point where execution of one or more operations and / or process 400 as described herein may terminate.
[0076] Box 410 of process 400 may be performed by one or more attention layers that, when providing the resulting encoding, assign one or more weights to each feature (or an element of the encoding obtained in box 406) of each input vector (or the encoding obtained in box 406). Features may include one or more elements of the encoding, such that one or more attention layers may assign weights to one or more features. For example, an attention layer may compute one or more alignment scores between a query (e.g., a vector used to determine the corresponding similarity between key inputs by using a dot product or scaled dot product) and each input vector (e.g., a key), apply a softmax operation to the alignment scores to obtain attention weights, multiply each input vector by its corresponding attention weight, and sum the weighted input vectors to obtain the resulting vector. In at least one embodiment, a processor performing process 400 may execute an encoder to output one or more vector encodings in box 406, such as one or more vectors of equal length. For example, given k vectors of dimension d, which represent k encodings of different information extracted from the log (e.g., at or before decision box 404) and encoded by a dedicated encoder (e.g., at box 406), the k encodings can be aggregated or combined using a Transformer encoder with L layers (e.g., one, two or more layers) and / or one or more multi-head self-attention layers (e.g., at box 410).
[0077] Box 406 of process 400 may include encoding the input (e.g., numerical data), for example by applying a learned fully connected layer W to embed the input into a high-dimensional space (e.g., a high-dimensional space of dimension d) and / or applying random Fourier features to the input, followed by the application of the learned layer W (e.g., this may facilitate the use of neural networks to embed low-dimensional inputs). In at least one embodiment, process 400 may include feature tokenization (e.g., at box 406) in combination with a transformer model (e.g., an attention layer) (e.g., at box 410). The feature tokenizer may transform one or more input features into one or more embeddings. One or more architectures used in conjunction with process 400 may include MLP, ResNet, and / or one or more tabular data models.
[0078] Although process 400 has been described as being executed by initial encoder function 113, process 400 may be executed by different functions, one or more processes, one or more services, one or more processors, etc. In at least one embodiment, part or all of process 400 (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and is implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more application programs) executed jointly by hardware, software, or a combination thereof on one or more processors. In at least one embodiment, the code is stored in the form of a computer program on a computer-readable storage medium comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of the computer-readable instructions available for executing process 400 are not stored using only transient signals (e.g., propagating transient electrical or electromagnetic transmissions). In at least one embodiment, the non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver of transient signals. In at least one embodiment, process 400 is executed at least partially on a computer system (e.g., a computer system described elsewhere in this disclosure). In at least one embodiment, process 400 is executed by logic (e.g., hardware, software, or a combination of hardware and software).
[0079] In at least one embodiment, one or more processors use process 400, for example, to encode at least one log message at least partially by: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code at least partially by combining at least the first code and the second code; and / or otherwise performing the operations described herein. In at least one embodiment, as an example, a set of instructions is stored on a machine-readable medium (e.g., non-transitory) that, if executed by one or more processors, causes one or more processors to perform process 400, for example, to encode at least one log message at least partially by: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code at least partially by combining at least the first code and the second code; and / or otherwise performing the operations described herein. In at least one embodiment, process 400 is included Figures 1 to 17B and / or Figure 23 The process shown, and / or otherwise includes Figures 1 to 17B and / or Figure 23 The process shown is used to encode at least one log message by at least partially: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code by at least partially combining the first code and the second code; and / or otherwise performing the operations described herein. In at least one embodiment, Figures 1 to 17B and / or Figure 23 The system execution process 400 shown herein, for example, is used to encode at least one log message at least partially by: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code at least partially by combining at least the first code and the second code; and / or otherwise performing the operations described herein. In at least one embodiment, Figures 17A to 22 One or more hardware processes 400 shown are used, for example, to encode at least one log message at least by: encoding information of a first type in at least one log message to obtain a first encoding; encoding information of a second type in at least one log message to obtain a second encoding; obtaining a result encoding at least by combining at least the first encoding and the second encoding; and / or otherwise performing the operations described herein.
[0080] Figure 5 This is a block diagram illustrating a system 500 for training one or more neural networks (e.g., neural network NN1) to encode one or more log sequences 508, according to at least one embodiment. In at least one embodiment, system 500 may implement at least a portion of encoder functionality 115 (see [link to documentation]). Figure 1 The system 500 may include a neural network training module 504 and / or communicate with the neural network training module 504. The system 500 may include one or more processors 502 as described herein for executing instructions (e.g., included in the neural network training module 504) to train and / or execute one or more neural networks (e.g., neural network NN1). As an example, the system 500 trains model 510 at least in part based on using contrastive learning and / or reducing or minimizing triplet loss.
[0081] System 500 can perform pre-training of model 510 to encode one or more log sequences without task-specific labels. In at least one embodiment, model 510 is implemented as an encoder to be trained using a triplet loss 506 between one or more vector encodings produced by model 510. By pre-training model 510 without task-specific labels, system 500 makes model 510 easily trainable to encode logs for different types of tasks with minimal labeling (e.g., as input to other neural networks and / or machine learning processes). System 500 can train a neural network (e.g., neural network NN1) using a contrastive learning method to encode log sequences from a dataset without task-specific labels and minimize a triplet loss 506 computed at least in part based on the encoded log sequences using a triplet loss function.
[0082] System 500 (its executable process 900) can create a training dataset without task-specific labels from a query sequence or raw logs (e.g., which may be referred to as anchor sequence 508A). Anchor sequence 508A comprises one or more individual log messages, each referred to as anchor 518A. Anchor sequence 508A can be modified or augmented to generate semantically similar and semantically different log sequences. Each log message in semantically similar or positive sequence 508B is referred to as a positive 518B example, and each log message in semantically different or negative sequence 508C is referred to as a negative 518C example. As an example, different combinations of log messages in the dataset can be identified as anchor sequence 508A, which can be modified to create log sequences that are semantically similar to anchor sequence 508A and log sequences that are semantically different from anchor sequence 508A. Anchor sequence 508A, positive sequence 508B, and negative sequence 508C together may be referred to as a sequence triple. Labels can then be used to identify anchor sequence 508A, positive sequence 508B, and negative sequence 508C, but task-specific labels may not be used. For example, if a downstream process uses the output of model 510 to determine the priority of log messages or log sequences, the training dataset may include labels identifying anchor sequence 508A, positive sequence 508B, and negative sequence 508C as anchors, positive, and negative, respectively, but may not include labels identifying sequences as associated with any particular priority.
[0083] In at least one embodiment, each positive 518B example is more similar to (e.g., semantically similar to) anchor 518A than each negative 518C example. As an example, a log message may include a text descriptor (e.g., INFO or WARN) associated with an event that can be classified by an anomaly detector into a priority level (e.g., low priority or high priority) (e.g., "INFO" = low priority, "WARN" = high priority). A positive example 518B of a log message including the text descriptor "INFO" would be a variant of the log message where "INFO" is replaced by another low-priority descriptor. On the other hand, continuing with the above example, a negative example 518C of a log message having the text descriptor "INFO" would be a variant of the log message where "INFO" is replaced by a high-priority descriptor, such as "WARN". A training dataset can be created by selecting different anchors for different sets of characters indicating information in the log to include in anchor sequence 208A, and using at least a portion of the selected anchors to create positive 518B and negative 518C examples to include in positive sequence 208B and negative sequence 208C, respectively.
[0084] Before training model 510 using the training dataset, an encoding process (e.g., process 400) can encode each anchor 518A, positive 518B example, and negative 518C example as a vector for each sequence triple (e.g., corresponding to one or more events in the log). In at least some embodiments, the initial encoding of the anchors (e.g., initial encoding EL1) can be combined to form anchor sequence 508A, the initial encoding of the positive examples can be combined to form positive sequence 508B, and the initial encoding of the negative examples can be combined to form negative sequence 508C. In at least some embodiments, anchors can be combined to form anchor sequence 508A, and anchor sequence 508A can be encoded to create the initial encoding of anchor sequence 508A. Similarly, positive examples can be combined to form positive sequence 508B, and negative examples can be combined to form negative sequence 508C, and then positive sequence 508B and negative sequence 508C can be encoded to create the initial encodings of positive sequence 508B and negative sequence 508C, respectively. Model 510 receives the initial encoding (e.g., initial encoding EL1) of anchor sequence 508A, positive sequence 508B and negative sequence 508C and encodes them as vectors or encodings “A” 512, “P” 514 and “N” 516.
[0085] The codes “A” 512, “P” 514, and “N” 516 correspond to the latent space (e.g., vector space 802, see [link]). Figure 8 The three positions in the encoding of the response triplet are defined by encoding “A” 512, “P” 514, and “N” 516. The model is trained 510 by adjusting the model parameters (e.g., weights) to reduce or minimize the loss function (e.g., triplet loss 506) based on the distance between the three positions encoded in the response triplet.
[0086] During training, model 510 (e.g., one or more transducer encoders) may receive a vectorized dataset (e.g., initial encoding EL1) without task-specific labels as input. This dataset may include one or more vectorized sequence triples for each of at least a subset of events in the dataset. For each sequence triple, model 510 may encode anchor sequence 508A, positive sequence 508B, and negative sequence 508C to produce encodings “A” 512, “P” 514, and “N” 516, respectively. Therefore, the generated encodings may include encoding “A” 512 corresponding to anchor sequence 508A, encoding “P” 512 corresponding to positive sequence 508B, and encoding “N” 516 corresponding to negative sequence 508C. A triplet loss 506 may then be computed for each response triple, which includes encodings “A” 512, “P” 514, and “N” 516. For each model configuration (e.g., set of parameter values, weight values, etc.), the triplet loss can be aggregated (e.g., total, average, etc.) for all response triples, and the model configuration that results in the minimum total triplet loss for the response triples can be selected for model 510 used at deployment. For example, processor 502 can use backpropagation to update one or more neural network weights, and then use model 510 to perform one or more inference operations. The triplet loss 506 can encourage the encoding of vectorized log events (e.g., encoding "A" 512, "P" 514, and "N" 516), which results in the distance between the encoded "A" 512 of anchor sequence 508A and the encoded "P" 514 of positive sequence 508B being less than the distance between the encoded "A" 512 of anchor sequence 508A and the encoded "N" 516 of negative sequence 508C. Furthermore, a margin distance can be specified, and the triplet loss 506 can encourage the encoding "P" 514 of the positive sequence 508B to be separated from the encoding "N" 516 of the negative sequence 508C by at least a margin distance. In at least one embodiment, the loss function is L(a,p,n) = max{d(a i ,p i )-d(a i ,n i )+margin,0}, where d(x) i ,y i )=‖x i -y i || p .
[0087] After determining the weights of model 510, model 510 (e.g., neural network NN1) can be used to infer the encoding (e.g., encoding EL2) of logs (which are encoded as initial encoding EL1). These encodings (e.g., encoding EL2) can be provided to one or more other processes, such as one or more other neural networks (e.g., MLP). For example, the encodings can be provided to a neural network trained to detect anomalies, which can infer whether each encoding indicates that an anomaly was recorded in each log. For example, the encodings (e.g., encoding EL2) generated by model 510 (e.g., neural network NN1) can be provided to classification function 116.
[0088] In at least one embodiment, processor 502 uses neural network training module 504 (e.g., neural network training module 1724) to train one or more neural networks (e.g., model 510 trained using vector-encoded triplet loss 506). In at least one embodiment, processor 502 executes neural network training module 504 and processes such as those described herein by including, or otherwise encoding, instructions that cause (e.g., by processor 502) to execute said one or more processes or are otherwise available for execution of said one or more processes. In at least one embodiment, the processor using neural network training module 504 obtains or is otherwise provided with one or more neural networks (e.g., by one or more systems (e.g., in combination)). Figure 1 The system described herein). In at least one embodiment, the processor 502 using the neural network training module 504 performs one or more processes (e.g., at least in conjunction with the above). Figures 5 to 9 The processes described herein) train the one or more neural networks (e.g., neural network NN1) using a training dataset. In at least one embodiment, the processor 502 using the neural network training module 504 trains the one or more neural networks using any suitable training process (e.g., a process described by triplet loss in combination with one or more vector encodings trained by model 510 (e.g., an encoder)).
[0089] In at least one embodiment, system 500 includes a collection of one or more hardware and / or software computing resources having instructions that, when executed, perform one or more communication processes (e.g., the communication processes described herein). In at least one embodiment, system 500 is a software program, an application program, and / or variations thereof, executing on computer hardware. In at least one embodiment, one or more processes of system 500 are executed by any suitable processing system or unit (e.g., a graphics processing unit (GPU), a general-purpose GPU (GPGPU), a parallel processing unit (PPU), a central processing unit (CPU), a data processing unit (DPU) (as described below)) in any suitable manner, including sequentially, in parallel, and / or variations thereof. In at least one embodiment, system 500 uses machine learning training frameworks (e.g., PYTORCH, TENSORFLOW, BOOST, CAFFE, MICROSOFT COGNITIVE TOOLKIT / CNTK, MXNET, CHAINER, KERAS, DEEPLEARNING4J) and / or other training frameworks to implement and perform the operations described herein to train a neural network to encode at least one vector associated with at least one log sequence and / or otherwise perform the operations described herein. In at least one embodiment, as an example, training the neural network model includes using a server (e.g., an NVIDIA DGX server) that also includes at least a GPU (e.g., AMD MI200, VEGAL10, VEGO20, and ARCTURUS), an optimizer (e.g., ADAM OPTIMIZER), or a discriminator architecture (e.g., a discriminator architecture from face-vid2vid for training with GAN loss).
[0090] In at least one embodiment, system 500 includes modules (e.g., modules 1724-1730, see below). Figure 17BThe system trains a neural network to encode at least one vector associated with at least one log sequence. In at least one embodiment, the module includes any combination of any type of logic (e.g., software, hardware, firmware) and / or circuitry configured to perform the function. In at least one embodiment, the module includes one or more circuits that form part of a larger system (e.g., integrated circuit (IC), system-on-a-chip (SoC), central processing unit (CPU), graphics processing unit (GPU), data processing unit (DPU), etc.). In at least one embodiment, the controller includes any type of logic (e.g., software, hardware, firmware) and / or any combination of circuitry configured to perform the function. In at least one embodiment, the software includes software packages, code, programming languages, drivers, instructions, instruction sets, or some combination thereof. In at least one embodiment, the hardware includes hardwired circuitry, programmable circuitry, state machine circuitry, fixed-function circuitry, execution unit circuitry, firmware storing instructions executed by the programmable circuitry, or some combination thereof.
[0091] In at least one embodiment, system 500 includes one or more logic units. In at least one embodiment, the logic unit includes firmware logic, hardware logic, or some combination thereof, configured to provide any functionality as further described herein. In at least one embodiment, the logic unit includes circuitry that forms part of a larger system (e.g., IC, SoC, CPU, GPU, DPU). In at least one embodiment, the logic unit includes logic circuitry for implementing firmware and / or hardware for training a neural network to encode at least one vector associated with at least one log sequence.
[0092] In at least one embodiment, system 500 includes one or more engines. In at least one embodiment, an engine includes modules and / or logic units as further described herein. In at least one embodiment, a component includes modules and / or logic units as further described herein. In at least one embodiment, an engine includes software logic, firmware logic, hardware logic, or some combination thereof, configured to provide any functionality as further described herein. In at least one embodiment, a component includes software logic, firmware logic, hardware logic, or some combination thereof, configured to provide any functionality as further described herein. In at least one embodiment, operations performed by hardware and / or firmware may alternatively be implemented via software modules, which may be embodied as software packages, code, and / or instruction sets. In at least one embodiment, a logic unit may also utilize a portion of software to implement its functionality.
[0093] In at least one embodiment, system 500 includes one or more processors for encoding at least one vector associated with at least one log sequence using a neural network, the neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein. In at least one embodiment, system 500 is included in Figures 1 to 17B and / or Figure 23 The system shown, and / or otherwise includes Figures 1 to 17B and / or Figure 23 The system shown herein is used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein.
[0094] In at least one embodiment, system 500 executes Figures 1 to 17B and / or Figure 23 One or more processes illustrated herein, for example, for encoding at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein. In at least one embodiment, system 500 includes Figures 17A to 22One or more pieces of hardware, such as those used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein.
[0095] Figure 6 This is a block diagram of a system 600 for training one or more transformer encoders to encode one or more log sequences, according to at least one embodiment. Logs can provide a rich source of information about the lifecycle of systems and services. The large-scale generation of logs and their inherent characteristics (e.g., lack of standardization and use of domain-specific terminology) can make manually extracting meaningful insights challenging. Furthermore, log lines often result in print statements written by developers. They can typically include domain-specific terminology, function names, and specific identifiers that may not adhere to language syntax or uniform standards. Log analysis tasks, such as anomaly detection, can manipulate log sequences for prediction.
[0096] exist Figure 6 In this embodiment, one or more processors 602 implement a machine learning model 604 (e.g., model 510), which includes a neural network 608 (e.g., a transducer encoder) and a multilayer perceptron (“MLP”) 606. The output of the neural network 608 is provided as input to the MLP 606. The machine learning model 604 receives a log sequence 612 (e.g., including one or more initial codes EL1) and an associated position sequence 614 (e.g., provided by the position encoder function 111) as input and outputs an encoding of that input (e.g., encoding EL2).
[0097] In at least one embodiment, processor 602 may tokenize one or more log sequences into one or more tags 610 (e.g., to aggregate sequences). Tags 610 may include event codes 612A-612E aggregated to form log sequence 612 and / or defining log sequence 612, and associated position codes 614A-614E aggregated to form position sequence 614 and / or defining position sequence 614. The input to neural network 608 includes one or more log event codes 612A-612E defining log sequence 612 (e.g., initial code EL1) and one or more position codes 614A-614E defining position sequence 614 (e.g., provided by position encoder function 111). Each log may correspond to one or more pairs of log sequences 612 and position sequences 614. In at least one embodiment, log sequence 612 is a vector. In at least one embodiment, position sequence 614 is a vector. In at least one embodiment, generating one or more positive log sequences and one or more negative log sequences by enhancing the anchor sequence may include: obtaining the anchor sequence from one or more datasets (e.g., an HDFS dataset) and transforming (e.g., enhancing, flipping, etc.) one or more messages or portions of the anchor sequence to modify its meaning. For example, if the anchor sequence includes a specific log message containing a specific event, processor 602 (e.g., performing neural network training module 504) may transform the specific log message into a corresponding positive event (e.g., to create a positive example) or a negative event (e.g., to create a negative example). As an additional non-limiting example, processor 602 may transform the anchor sequence by truncating and / or parsing the log message (e.g., removing numbers, punctuation marks, and special characters). Positive and negative examples may be generated from the input anchor sequence by flipping one or more messages.
[0098] System 600 performs a process (e.g., process 900) for encoding one or more sequences of log messages that do not have explicit labels associated with downstream target tasks. System 600 may apply local augmentations to generate positive sequences 508B (semantically similar) and negative sequences 508C (semantically different) from a given anchor sequence 508A, and then optimize neural network 608 (e.g., neural network NN1) using contrastive learning (triple loss). Since labels for downstream tasks may be undefined, system 600 can leverage a large amount of available data and generate generic codes for one or more downstream machine learning programs or processes.
[0099] Then, the neural network 608 trained with triplet loss (e.g., a converter encoder) can be fine-tuned for a specific downstream task (e.g., anomaly detection), or the neural network 608 can be trained for a specific downstream task (e.g., anomaly detection). The neural network 608 (e.g., a converter encoder) can compute sequence encodings as input to the MLP 606, which does not operate on sequences and requires less labeled data (e.g., random forest, logistic regression, isolated forest). In at least one embodiment, one or more neural networks (e.g., neural network 608 and MLP 606) minimize the triplet loss, for example, by using the following equation: L(a,p,n)=max{d(a i ,p i )-d(a i ,n i )+margin,0} and d(x i ,y i )=‖x i -y i || p .
[0100] In at least one embodiment, system 600 includes one or more processors for encoding at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein. In at least one embodiment, system 600 is included Figures 1 to 17B and / or Figure 23 The system shown, and / or otherwise includes Figures 1 to 17B and / or Figure 23The system shown is used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein.
[0101] In at least one embodiment, system 600 executes Figures 1 to 17B and / or Figure 23 One or more processes illustrated herein, for example, for encoding at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein. In at least one embodiment, system 600 includes Figures 17A to 22 One or more pieces of hardware, as shown, are used, for example, to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein.
[0102] Figure 7 The diagram illustrates a system 700 that embeds vectors representing one or more logs according to at least one embodiment. In at least one embodiment, one or more logs are used to train a neural network based at least in part on a triplet loss: anchors 702A obtained from the logs (e.g., anchor sequence 508A, see...). Figure 5), at least in part based on the positive example 702B obtained from anchor 702A (e.g., positive sequence 508B, see Figure 5 ), and negative example 702C (e.g., negative sequence 508C, see below) obtained at least in part based on anchor 702A. Figure 5 ).
[0103] System 700 (its executable process 900) can begin by creating a training dataset without task-specific labels based on a query sequence or raw logs (e.g., which may be referred to as anchor 702A). Anchor 702A can be augmented to generate semantically similar and semantically distinct sequences (referred to as positive example 702B and negative example 702C, respectively). As another example, different combinations of logs in the dataset can be identified as anchor 702A, logs semantically similar to anchor 702A can be identified as positive example 702B, and logs semantically different from anchor 702A can be identified as negative example 702C. Anchor 702A, positive example 702B, and negative example 702C can be collectively referred to as sequence triples. Anchor 702A, positive example 702B, and negative example 702C can then be identified using labels, but task-specific labels cannot be used. In at least one embodiment, the positive example 702B is more similar to the anchor 702A (e.g., semantic similarity) than the negative example 702C is to the anchor 702A, making them closer in position in the embedding space.
[0104] During training, a machine learning process (e.g., a combination of neural network NN1, model 510, neural network 608, and MLP 606, one or more transducer encoders, etc.) may receive one or more sequence triplets as input from a vectorized dataset (e.g., including one or more initial encoders EL1) corresponding to one or more logs 702. The sequence triplets in the vectorized dataset do not have task-specific labels. The machine learning process generates vectorized response triplets, such as embedding vectors 704, for each sequence triplet of at least a portion of the dataset. For each sequence triplet of log 702, the machine learning process generates embedding vectors 704A-704C corresponding to anchor 702A, positive example 702B, and negative example 702C, respectively. Embedding vectors 704A-704C may each represent a position in the embedding space 708.
[0105] One or more processors (e.g., those executing neural network training module 504) can compute the triplet loss (as described herein) for each response triplet (embedding vector 704) of the machine learning process output and select settings (e.g., parameters, weights, etc.) of the machine learning process that result in the desired (e.g., minimum) triplet loss. This can be achieved by executing... Figures 5 to 9The location of the embedding vector is generated by one or more operations described herein. In at least one embodiment, one or more neural networks (e.g., neural network 608 and MLP 606) compute the triplet loss, for example by using one or more equations, such as L(a,p,n)=max{d(a i ,p i )-d(a i ,n i )+margin,0} and d(x i ,y i )=‖x i -y i || p .
[0106] In at least one embodiment, system 700 includes one or more processors for encoding at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein. In at least one embodiment, system 700 is included Figures 1 to 17B and / or Figure 23 The system shown, and / or otherwise includes Figures 1 to 17B and / or Figure 23 The system shown herein is used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein.
[0107] In at least one embodiment, system 700 performs Figures 1 to 17B and / or Figure 23One or more processes illustrated herein, for example, for encoding at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein. In at least one embodiment, system 700 includes Figures 17A to 22 One or more pieces of hardware, as shown, are used, for example, to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein.
[0108] Figure 8 This is a schematic diagram illustrating a system 800 that trains one or more neural networks at least partially based on triplet loss according to at least one embodiment. System 800 may perform training 804 of one or more neural networks at least partially based on location encodings in a vector space 802, such that one or more weights of the neural network are updated according to a triplet loss that minimizes the distance between points in the vector space 802.
[0109] In at least one embodiment, the processor of system 800 performs training 804 based at least in part on minimizing the triplet loss of one or more log encodings. A triplet loss can be computed for each response triplet obtained at least in part based on sequence triplets obtained at least in part based on one or more logs. The response triplet includes an embedding of the position vector of the sequence triplet containing anchors, positive examples, and negative examples. For each configuration (e.g., set of parameter values, weight values, etc.) of the machine learning process used to generate the response triplets, the triplet loss can be aggregated (e.g., total, average, etc.) over all response triplets, and a configuration that results in the minimum total triplet loss for the response triplets can be selected for use at deployment time. For example, the sum of the triplet losses of one or more logs can be calculated for all response triplets, and one or more model weights that result in the minimum total triplet loss can be selected for the machine learning process (e.g., model 510). The triplet loss encourages the encoding (e.g., encodings 810-814) of vectorized log events in vector space 802, such encoding results in a distance between encoding 810 of the obtained anchor 518A and encoding 812 of the obtained positive example 518B being less than the distance between encoding 810 of the obtained anchor 518A and encoding 814 of the obtained negative example 518C. Furthermore, a margin distance can be specified, and the triplet loss encourages the distance between encodings 812 and 814 of the obtained positive example 518B and negative example 518C, respectively, to be at least a margin distance. In at least one embodiment, one or more neural networks (e.g., neural network 608 and MLP 606) minimize the triplet loss, for example, by using one or more of the following equations: L(a,p,n)=max{d(a i ,p i )-d(a i ,n i )+margin,0} and d(x i ,y i )=‖x i -y i || p .
[0110] In at least one embodiment, system 800 includes one or more processors for encoding at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein. In at least one embodiment, system 800 is included Figures 1 to 17B and / or Figure 23 The system shown, and / or otherwise includes Figures 1 to 17B and / or Figure 23 The system shown herein is used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein.
[0111] In at least one embodiment, system 800 performs Figures 1 to 17B and / or Figure 23 One or more processes illustrated herein, for example, for encoding at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein. In at least one embodiment, system 800 includes Figures 17A to 22One or more pieces of hardware shown herein, for example, are used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein.
[0112] Figure 9 This is a flowchart illustrating a process 900 for training a neural network to encode at least one vector associated with a log sequence, according to at least one embodiment. In at least one embodiment, process 900 begins when invoked by a processor and / or when the processor receives one or more logs as input in block 902. In block 902, the processor may receive a combination of... Figures 1 to 16 The description refers to one or more logs and / or one or more inputs. For example, in box 902, the processor may receive an initial encoding EL1 representing one or more logs or one or more portions thereof. After receiving one or more logs as input in box 902, in box 904, the processor may identify at least one sequence triplet, each sequence triplet comprising a first log sequence, a similar second log sequence, and a dissimilar third log sequence. Identifying the first log sequence, the similar second log sequence, and the dissimilar third log sequence in box 904 may include: receiving a dataset, identifying, as about Figures 5 to 8 The and in Figures 5 to 8 The anchor shown, generating positive examples as similar second log sequences, and generating negative examples as dissimilar third log sequences. Identifying the first log sequence, the similar second log sequence, and the dissimilar third log sequence in box 904 may include: receiving a triplet of log sequences, wherein the first sequence is identified as the anchor, and selecting the most similar log sequence from the anchor as the similar second log sequence, and then identifying the log sequence least similar to the anchor as the dissimilar third log sequence.
[0113] Once sequence triples are identified in box 904, in box 906 the processor can use a model (e.g., model 510, neural network NN1, etc.) to encode the first, second, and third log sequences of each sequence triple into a vector. The encoded vectors in box 906 can each have the same length. To encode the first, second, and third log sequences into a vector in box 906, the processor can use... Figures 5 to 8The operation described in the document refers to one or more operations. The three encoded vectors obtained from the first, second, and third log sequences in each sequence triplet are response triples.
[0114] Then, at box 908, the processor of execution process 900 uses the response triples obtained in box 906 to compute the total triplet loss for the current configuration of the model used to generate the response triples in box 906. The processor of execution process 900 may compute the triplet loss for each response triplet and aggregate (e.g., sum, average, etc.) one or more triplet losses to obtain the total triplet loss. As an example, the triplet loss helps ensure that the position code of the first log sequence corresponding to the anchor is closer to the position code of the second similar sequence, rather than the position code of the anchor being closer to the position code of the dissimilar sequence, while still adhering to... Figure 8 The margin shown.
[0115] At decision box 910, the processor of execution process 900 determines whether to modify the model (e.g., change model parameters, weights, and / or other settings). If the processor determines that doing so will produce better results, the processor may decide to modify the model. When the processor decides to modify the model, the decision in decision box 910 is "yes". Otherwise, the decision in decision box 910 is "no". When the decision in decision box 910 is "yes", in box 912, the processor modifies the model and returns to box 906 to generate a new encoding of sequence triples. On the other hand, when the decision in decision box 910 is "no", the processor proceeds to box 914.
[0116] At box 914, the processor of execution process 900 can select the configuration of the model associated with the desired (e.g., minimum) total triplet loss.
[0117] After minimizing the triplet loss in box 914, the processor executing process 900 can output the selected model configuration (e.g., one or more model weights) in box 916. The selected model configuration (e.g., one or more output model weights) can then be used, for example, in neural network training module 504 (see [link to module 914]). Figure 5The model is used to update a neural network, such as neural network NN1. Updating the neural network using the model configuration (e.g., model weights) selected at box 916 can result in an encoder trained using a triplet loss obtained from one or more vector encodings, for example, once the desired performance is achieved after one or more training iterations using boxes 906-912 of process 900. Process 900 may include generating the output of one or more model weights in box 916, providing said weights to update the neural network, repeating the process one or more iterations, performing one or more operations described herein, and / or proceeding to the end. For example, process 900 may terminate after box 916. After updating the model (e.g., neural network NN1) according to the selected model configuration, the model can be used to encode log messages and / or sequences (e.g., as part of an anomaly detection pipeline).
[0118] In at least one embodiment, part or all of process 900 (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and is implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more application programs) executed jointly by hardware, software, or a combination thereof on one or more processors. In at least one embodiment, the code is stored in the form of a computer program on a computer-readable storage medium comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of the computer-readable instructions available for executing process 900 are not stored using only transient signals (e.g., propagating transient electrical or electromagnetic transmissions). In at least one embodiment, the non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver of transient signals. In at least one embodiment, process 900 is executed at least partially on a computer system (e.g., a computer system described elsewhere in this disclosure). In at least one embodiment, logic (e.g., hardware, software, or a combination of hardware and software) performs process 900.
[0119] In at least one embodiment, one or more processors use process 900, for example, to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein. In at least one embodiment, as an example, a set of instructions is stored on a machine-readable medium (e.g., non-transitory) that, if executed by one or more processors, causes one or more processors to perform process 900, for example, to encode at least one vector associated with at least one log sequence using a neural network, the neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein. In at least one embodiment, process 900 is included Figures 1 to 17B and / or Figure 23 The process shown, and / or otherwise includes Figures 1 to 17B and / or Figure 23 The process shown is used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein.
[0120] In at least one embodiment, Figures 1 to 17B and / or Figure 23The system execution process 900 shown herein, for example, is used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein. In at least one embodiment, Figures 17A to 22 One or more hardware execution processes 900 shown herein are used, for example, to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing the operations described herein.
[0121] Figure 10This is a block diagram illustrating a system 1000 that performs at least one neural network (e.g., encoder 114, neural network NN1, neural network NN2, classifier 122, and / or others) to classify one or more logs, according to at least one embodiment. In at least one embodiment, system 1000 includes one or more encoders and one or more neural networks as described herein, such that system 1000 performs classification (e.g., anomaly classification) indicated or present in one or more logs. System 1000 may include one or more processors, neural networks, encoder 1018, log event classification model 1006, anomaly detection model 1014, one or more models 1016 for performing root cause analysis, and / or combinations thereof. System 1000 may perform log preprocessing 1002, which may be implemented at least in part by preprocessing function 112, initial encoder function 113, one or more encoders 114, encoder function 115, and / or neural network NN1. System 1000 may perform classification operation 1004, which includes or involves encoder 1018 and / or log event classification model 1006. The classification operation 1004 can be implemented at least partially using the classification function 116, and the encoder 1018 and / or the log event classification model 1006 can be implemented at least partially using a neural network NN2. The system 1000 can perform one or more combined operations 1010, which can be implemented at least partially by the aggregation function 117. The system 1000 can execute an anomaly detection model 1014, a root cause analysis model 1016, and / or one or more other models, which can be implemented at least partially by the downstream function 118 and / or the classifier 122.
[0122] While logs may include information generated by software running on the computing system (e.g., error messages), telemetry data 1008 (e.g., provided by telemetry function 108) includes information about the computing system itself, such as bit error rate (BER), CPU utilization, memory utilization, disk I / O, temperature, etc., as the computing system executes its software. Collecting telemetry data 1008 may involve measuring data transmissions from remote sources, such as physical or electrical data. Telemetry data 1008 can be collected using sensors or other devices, such as temperature sensors, counters (e.g., for counting anomalous events over time), or other telemetry information as described herein.
[0123] Telemetry data 1008 and logs can both include information that can be used to evaluate computing systems (data centers). If at least some types of available data are not used when performing anomaly detection, this can negatively impact the capabilities of downstream processes (e.g., event prediction, root cause analysis, and / or observation generation) because useful information may be lost, thus failing to detect anomalies.
[0124] System 1000 may execute one or more processes 1100-1300, such as combining log data and telemetry data 1008 to detect anomalies within a computing system (e.g., a data center). Furthermore, system 1000 may encode network topology data 1012 (e.g., devices, physical connections, or locations) in combination with or separately from telemetry data 1008 and log data, and / or incorporate topology data 1012 into anomaly detection and / or other operations performed by system 1000.
[0125] System 1000 can perform log preprocessing 1002, which encodes textual, numerical, and categorical (e.g., metadata) information included and / or associated with log entries (e.g., log events). The processor performs log preprocessing 1002 (e.g., by performing preprocessing function 112) on one or more log entries (e.g., log line stream 1002A) to clean the content (e.g., remove irrelevant information) and / or extract useful information before encoding the log entries. For a specific log entry in log line stream 1002A, useful information may include time, timestamp, identification information, and / or one or more descriptions of the content of the specific log entry. Useful information can be extracted as parameter values (and separated to create...). Figure 1 The data SD shown is used. Parameter values can also be extracted from metadata, such as priority and / or message type associated with log entries (e.g., for classifying log data). System 1000 then uses the extracted parameter values (e.g., data SD) to encode each log entry of log line stream 1002A into a vector (e.g., using initial encoder function 113). In at least one embodiment, system 1000 uses process 1100 (e.g., by using one or more encoders 208 (see...)). Figure 2 Log preprocessing 1002 is performed on log line stream 1002A. Log preprocessing 1002 includes encoding log line stream 1002A with topology information and / or metadata obtained from topology and metadata information 1002C, such that the embedded vectors represent both the logs and the corresponding topology and / or metadata associated with the logs. Then, system 1000 generates processed and encoded event and node identifiers 1002B, which may include one or more vectors. At this point, system 1000 (e.g., its encoder function 115) may encode the processed and encoded event and node identifiers 1002B (e.g., a neural network NN1) to produce one or more vectors.
[0126] The system 1000, which performs classification operation 1004 (e.g., implements classification function 116), can use predefined event tags to classify processed events and node identifiers 1002B (e.g., whether an alarm or ignored log is triggered), for example, to detect anomalies. The predefined event tags can be associated with one or more characteristics of tags, character sets, or vectors representing logs, allowing the encoder 1018 to identify whether an alarm or ignored log message and / or log sequence is triggered from one or more predefined event tags for anomaly detection model 1014 to perform anomaly detection. The encoder 1018 may include... Figures 1 to 16 The description refers to one or more encoders or neural networks. However, there may be situations where one or more predefined event labels can lead to classification conflicts or unknown classifications in the logs.
[0127] The processor of system 1000 may (e.g., using encoder 1018) classify each encoded log entry (e.g., which encodes log events) or determine that the classification of an encoded log entry is unknown. If encoder 1018 is unable to classify the encoding, classification operation 1004 may use log event classification model 1006 to classify the encoding. In at least some embodiments, both encoder 1018 and log event classification model 1006 may be used to determine the classification of one or more encodings. In at least one embodiment, log event classification model 1006 is or otherwise includes encoder 1408 trained using a similarity loss on one or more vector encodings. By way of a non-limiting example, classification operation 1004 (e.g., using encoder 1018) may attempt to classify each encoded log entry into either the "alarm" or "ignore" class. Classification operation 1004 may use each event predefined by a domain expert as belonging to either the "alarm" or "ignore" class to classify the encoded log entry. For example, a predefined event associated with the text descriptor "WARN" can be associated with the "Alarm" class. Classification operation 1004 can use this predefined event to classify encoded log entries that encode the text descriptor "WARN" as belonging to the "Alarm" class. However, classification operation 1004 may encounter specific encoded log entries that, due to encoding new information (e.g., a new event), do not match any predefined event or match more than one predefined event, resulting in ambiguous classifications (e.g., conflicting classifications of "Alarm" or "Ignore"). For such encoded log entries, classification operation 1004 can use log event classification model 1006 to determine their classification.
[0128] Log event classification model 1006 may include encoder 1408 (e.g., one or more neural networks, such as neural network NN1, neural network NN2, and / or classifier 122) that encodes coded log entries (e.g., log events) using semantic similarity to produce classified log entries. The semantic similarity encoder (e.g., LLM) can be fine-tuned without using task-specific labels. Log event classification model 1006 can generate classifications associated with the encodings, which can be used to update predefined event labels (which include a set of predefined events or encodings associated with the determined classifications), and classification operation 1004 (e.g., encoder 1018) uses these predefined event labels to classify the encodings. In this way, encodings previously unassociated with a classification (and therefore unknown) can be added to the predefined event labels. Log event classification model 1006 can determine the classifications and update the predefined event labels to include the determined classifications (e.g., whether a particular log should be classified as an alarm or ignored). An encoder (e.g., a neural network) can be readily trained to encode one or more log entries and produce categorized log entries for different types of tasks (e.g., as input to an anomaly detection model 1014 and / or root cause analysis 1016), with minimal annotation for small event sets. The categorized log entries can be used as input to downstream processes (e.g., one or more neural networks, such as neural network NN1, neural network NN2, classifier 122), which can further categorize the categorized log entries and / or perform other inference operations.
[0129] System 1000 can fine-tune log event classification model 1006 (e.g., neural network NN1 and / or neural network NN2) by training the log event classification model 1006 using semantic similarity and cosine similarity losses associated with encoded log entry pairs. As an example, encoded log entries can encode text descriptors (e.g., INFO or WARN) associated with events that can be classified as priority levels (e.g., low priority or high priority) by an anomaly detector (e.g., "INFO" = low priority, "WARN" = high priority). A pair of encoded log entries associated with a high similarity score would include a first log entry containing the text descriptor "INFO" and a second log entry containing another low-priority descriptor. Conversely, two log entries with low similarity scores would include a first log entry with the text descriptor "INFO" (e.g., low priority) and a second log entry with the text descriptor "WARN" (e.g., high priority). System 1000 can use a loss function (e.g., a cosine similarity loss function) to generate a loss value (e.g., a cosine similarity loss value) for the model results obtained for two vectors about two encoded log entries. The loss values obtained for multiple pairs of encoded log entries in the training dataset can be aggregated (e.g., summing, averaging, etc.), and a model configuration (e.g., weight values, parameter values, and / or other settings) that produces the desired loss (e.g., minimum aggregated cosine similarity loss value) can be selected. After fine-tuning the log event classification model 1006, if the log event classification model 1006 determines that a previously known classification (e.g., the text descriptor “WARN” belongs to the “alarm” class) is semantically similar to a new encoded log entry, then the log event classification model 1006 will classify the new event similarly. Furthermore, this determination made by the log event classification model 1006 can be used to further update the predefined event labels as described above. Once log entries are encoded and categorized (e.g., categorized as “ignore” or “alarm”), combination operation 1010 combines log event information (e.g., categorization obtained through classification operation 1004) with telemetry data 1008 and topology data 1012.
[0130] Third, the combination operation 1010 combines (e.g., fuses) the classified log entries with telemetry data 1008 and / or topology data 1012, for example, using node-based fusion and aggregation. Node-based fusion and aggregation combines the classified log entries, node counters (and / or node identifiers), and telemetry data 1008 together. For example, features of the classified log entries, telemetry data 1008, and topology data 1012 can be combined into a joint table. As another example, these features can be combined by creating a joint vector representing the feature set (e.g., combining at least one log entry, telemetry data 1008, and / or topology data 1012), and the vector representation of the data can be used to train the anomaly detection model 1014. The anomaly detection model 1014 can extract and / or use the joint features.
[0131] System 1000 can provide one or more joint features as input to an anomaly detection model 1014, which can classify one or more categorized log entries as anomalies. Anomaly detection model 1014 can use topology data 1012 (e.g., node identifiers) to detect the location of anomalies. Topology data 1012 can be used for root cause analysis 1016 (RCA), such as determining when anomalies occur at clustered locations (e.g., combining information tables or creating joint vector representations of features). The output of anomaly detection model 1014 can include reports, such as for generating alarms requiring manual action or as input to root cause analysis 1016.
[0132] In at least one embodiment, system 1000 includes a collection of one or more hardware and / or software computing resources having instructions that, when executed, perform one or more communication processes, such as those described herein. In at least one embodiment, system 1000 is a software program, an application program, and / or variations thereof, executing on computer hardware. In at least one embodiment, one or more processes of system 1000 are executed by any suitable processing system or unit (e.g., graphics processing unit (GPU), general-purpose GPU (GPGPU), parallel processing unit (PPU), central processing unit (CPU), data processing unit (DPU) (described below)) in any suitable manner, including sequentially, in parallel, and / or variations thereof. In at least one embodiment, system 1000 uses a machine learning training framework (e.g., PYTORCH, TENSORFLOW, BOOST, CAFFE, MICROSOFT COGNITIVE TOOLKIT / CNTK, MXNET, CHAINER, KERAS, DEEPLEARNING4J, and / or other training frameworks) to implement and perform the operations described herein to execute a neural network to classify one or more logs and / or otherwise perform the operations described herein. In at least one embodiment, as an example, training the neural network model includes using a server (e.g., an NVIDIA DGX server) that also includes at least a GPU (e.g., AMD MI200, VEGAL10, VEGO20, and ARCTURUS), an optimizer (e.g., ADAMOPTIMIZER), or a discriminator architecture (e.g., a discriminator architecture from face-vid2vid for training with GAN loss).
[0133] In at least one embodiment, system 1000 includes modules (e.g., modules 1724-1730, see below). Figure 17BThe system executes a neural network to classify one or more logs. In at least one embodiment, the module includes any combination of logic (e.g., software, hardware, firmware) and / or circuitry configured to perform the function. In at least one embodiment, the module includes one or more circuitry that forms part of a larger system (e.g., an integrated circuit (IC), a system-on-a-chip (SoC), a central processing unit (CPU), a graphics processing unit (GPU), a data processing unit (DPU), etc.). In at least one embodiment, the controller includes any combination of logic (e.g., software, hardware, firmware) and / or circuitry configured to perform the function. In at least one embodiment, the software includes software packages, code, programming languages, drivers, instructions, instruction sets, or some combination thereof. In at least one embodiment, the hardware includes hardwired circuitry, programmable circuitry, state machine circuitry, fixed-function circuitry, execution unit circuitry, firmware storing instructions executed by the programmable circuitry, or some combination thereof.
[0134] In at least one embodiment, system 1000 includes one or more logic units. In at least one embodiment, the logic unit includes firmware logic, hardware logic, or some combination thereof, configured to provide any functionality as further described herein. In at least one embodiment, the logic unit includes circuitry that forms part of a larger system (e.g., IC, SoC, CPU, GPU, DPU). In at least one embodiment, the logic unit includes logic circuitry for implementing firmware and / or hardware to execute a neural network to classify one or more logs.
[0135] In at least one embodiment, system 1000 includes one or more engines. In at least one embodiment, an engine includes modules and / or logic units as further described herein. In at least one embodiment, a component includes modules and / or logic units as further described herein. In at least one embodiment, an engine includes software logic, firmware logic, hardware logic, or some combination thereof, configured to provide any functionality further described herein. In at least one embodiment, a component includes software logic, firmware logic, hardware logic, or some combination thereof, configured to provide any functionality as further described herein. In at least one embodiment, operations performed by hardware and / or firmware may alternatively be implemented via software modules, which may be embodied as software packages, code, and / or instruction sets. In at least one embodiment, a logic unit may also utilize a portion of software to implement its functionality.
[0136] In at least one embodiment, system 1000 includes one or more processors configured to: classify one or more log entries to obtain one or more classified log entries; obtain combined information at least in part by combining at least one or more classified log entries and telemetry data 1008; classify the combined information using at least one machine learning process; and / or otherwise perform the operations described herein. In at least one embodiment, system 1000 is included Figures 1 to 17B and / or Figure 23 The system shown and / or otherwise includes Figures 1 to 17B and / or Figure 23 The system shown is used to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least in part by combining at least one or more classified log entries and telemetry data 1008; to classify the combined information using at least one machine learning process; and / or otherwise perform the operations described herein. In at least one embodiment, system 1000 performs... Figures 1 to 17B and / or Figure 23 The system 1000 includes one or more processes, such as classifying one or more log entries to obtain one or more classified log entries; obtaining combined information at least in part by combining at least one or more classified log entries and telemetry data 1008; classifying the combined information using at least one machine learning process; and / or otherwise performing the operations described herein. In at least one embodiment, the system 1000 includes Figures 17A to 22 The hardware shown herein may be used for classifying one or more log entries to obtain one or more classified log entries; obtaining combined information at least in part by combining at least one or more classified log entries and telemetry data 1008; classifying the combined information using at least one machine learning process; and / or otherwise performing the operations described herein.
[0137] Figure 11An exemplary process 1100 for preprocessing at least one log according to at least one embodiment is illustrated. Process 1100 is an exemplary process, and the processor may otherwise perform content cleanup (e.g., removal of numbers, special characters, and / or separated camelCase words). In at least one embodiment, preprocessing function 112 and / or initial encoder function 113 execute process 1100. As an example, the processor executes process 1100 on one or more logs generated by a subnet manager (SM) for performing computer networking (e.g., InfiniBand (IB) networking). One or more processors may execute process 1100, for example, for performing log preprocessing 1002. In at least one embodiment, log preprocessing may include using encoder 208 and / or 308 (see...). Figure 2 and / or Figure 3 Process 1100 may include obtaining (as input) one or more raw logs 1102 (e.g., log entries for an InfiniBand (IB) network), cleaning the contents and extracting general fields 1104 (e.g., characteristics), extracting subnet manager (SM) parameters 1106 (e.g., OpenSM parameters), extracting topology information 1108, extracting metadata 1110, and outputting preprocessed logs 1112 or a combination thereof.
[0138] The input raw log 1102 may otherwise be a log prior to preprocessing, allowing or preventing the removal of certain portions of the log message during preprocessing. Process 1100 may also include cleanup (e.g., removal) of punctuation, numbers, and / or special characters during content cleanup and extraction of common fields 1104. For example, log preprocessing process 1100 may include extracting one or more parameters (where one or more parameters may be ignored), for example, to create a fixed vocabulary. The processor executing process 1100 may then obtain or extract topological information related to one or more logs, such as from topological and metadata information 1002C, and continue extracting metadata 1110, such as metadata from topological and metadata information 1002C. Figures 1 to 16 One or more other steps of encoding and / or preprocessing the log information described herein may be included in process 1100 to generate the output of preprocessed log 1112, for example, for use by system 1000.
[0139] In at least one embodiment, one or more processors use process 1100, for example, to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least partially by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or otherwise perform the operations described herein. In at least one embodiment, as an example, a set of instructions is stored on a machine-readable medium (e.g., non-transitory) that, if executed by one or more processors, causes one or more processors to perform process 1100, for example, to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least partially by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or otherwise perform the operations described herein. In at least one embodiment, process 1100 is included... Figures 1 to 17B and / or Figure 23 The process shown, and / or otherwise includes Figures 1 to 17B and / or Figure 23 The process shown herein is used to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least in part by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or otherwise perform the operations described herein. In at least one embodiment, Figures 1 to 17B and / or Figure 23 The system execution process 1100 shown herein includes, for example, classifying one or more log entries to obtain one or more classified log entries; obtaining combined information by at least partially combining one or more classified log entries and telemetry information; classifying the combined information using at least one machine learning process; and / or otherwise performing the operations described herein. In at least one embodiment, Figures 17A to 22 The one or more hardware execution processes 1100 shown herein are used, for example, to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least in part by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or to otherwise perform the operations described herein.
[0140] Figure 12The flowchart illustrates a process 1200 for classifying one or more log messages according to at least one embodiment. In at least one embodiment, process 1200 begins when invoked by one or more processors and / or when one or more processors receive log entry input at block 1204.
[0141] Anomalies in communication networks occur at different levels and in different modalities. For example, anomalies can occur through log data generated by network devices or through telemetry streams generated from counters that measure different properties, such as temperature and bit error rate (BER). Furthermore, the underlying network topology plays a role in correlating detected anomalies with network behavior and assessing their impact. While each modality can generate a large amount of data, each modality provides an incomplete view of the system, thus providing incomplete input for anomaly detection. Reasoning and integrating large amounts of data from multiple modalities is crucial for accurately detecting anomalies and identifying those that actually affect network behavior.
[0142] The processor executing process 1200 integrates log, telemetry, and topology data into anomaly detection. Process 1200 may include fusing log information, telemetry information, and network topology to detect anomalies in the communication network. Process 1200 may include processing one or more log lines, then extracting one or more classifications at least partially based on the one or more log lines and mapping the one or more classifications to at least one unique node identifier. Process 1200 may include associating logs with node telemetry. In at least one embodiment, the processor executing process 1200 classifies one or more log entries at box 1206 (e.g., using classification operation 1004) at least partially based on classifying log entry input received at box 1204 using one or more encoders 1018 at least partially based on one or more predefined event labels and / or using a log event classification model 1006 to classify log entry input received at box 1204. As an example, if an event has not been previously labeled (due to ambiguity or the presence of new, unseen events), the log event classification model 1006 predicts its label in box 1206. In box 1206, the classification of an event may rely on predefined labels and / or the log event classification model 1006. In box 1206, the classification output may be further used to update (e.g., offline after domain expert review) the predefined labels, which may be stored in a database.
[0143] Then, one or more processors in execution process 1200 can combine telemetry information, topology information, and log information in box 1208. In box 1208, the processors can fuse classified log events and node counters and perform joint feature extraction and anomaly detection. The information can be combined into a table in box 1208 and / or vectors corresponding to the combined telemetry information, topology information, and log information in box 1208 can be generated, such that the anomaly detection model 1014 is trained from the combined vectors. As an example, the vectors described herein can be N-dimensional tensors.
[0144] After combining telemetry information, topology information, and log information in block 1208, the processor executing process 1200 may continue to classify the combined information in block 1210. In block 1208, the processor may classify one or more log events as important or unimportant (alarm or ignore, respectively). In at least one embodiment, the combined information may be classified to determine anomaly classification, one or more event predictions, identified root causes, generated observations, and / or one or more information indications. Using network topology, the processor executing process 1200 may identify clusters of anomalies and classify the anomalies based on their topological attributes (e.g., anomalies involving physically close nodes). Output may include providing one or more classifications and / or reports in block 1212, which may generate alarms for manual operators or serve as input for root cause analysis. As an example, process 1200 may be executed for anomaly detection in InfiniBand (IB) networks, Ethernet networks, and / or to generate input for root cause analysis in communication networks.
[0145] The processor can measure the performance of the execution process 1000. One or more measurements of the execution of said process 1000 may include precision, recall, or measurements of one or more generated scores. For example, precision may be measured as the score of how many of the predicted events that are positive (e.g., outliers) are actually positive (e.g., the number of true positives divided by the sum of true positives and false positives). For example, a score of recall may measure how many actual positive cases are correctly predicted by the model (e.g., the number of true positives divided by the sum of true positives and false negatives). Measurements of performance may also include the F1 score, which is the harmonic mean of recall and precision. Performance metrics may include term frequency (TF), which is the frequency of a particular term relative to documents. Examples of frequency measurements may include raw counts, normalized counts, log-scale counts, or other frequency measurements. Inverse document frequency (IDF) may include how common (or uncommon) a term t is in a corpus D with N documents. As an example, TF-IDF comprises the product of TF and IDF values (e.g., the importance of a term is inversely correlated with its frequency in a document). As an example, possible measures of the performance of one or more models used in association with process 1200 may include...
[0146] As an example, process 1200 may include using one or more tokenizers, such as a wordpiece tokenizer. A wordpiece tokenizer may include first setting characters and symbols into its base vocabulary. WordPiece may include selecting pairs that maximize the likelihood of the training data, rather than relying on the frequency of the pairs. As an example, the rare word "datablockscanner" is broken down into more frequent subwords: {"data", "block", "scan", "ner"}. In this way, the number of OOV words can be reduced and their meanings can be captured. WordPiece can handle OOV words and potentially reduce the size of the vocabulary.
[0147] In at least one embodiment, part or all of process 1200 (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and is implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more application programs) jointly executed on one or more processors by hardware, software, or a combination thereof. In at least one embodiment, the code is stored in the form of a computer program on a computer-readable storage medium comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of the computer-readable instructions available for executing process 1200 are not stored using only transient signals (e.g., propagating transient electrical or electromagnetic transmissions). In at least one embodiment, the non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver of transient signals. In at least one embodiment, process 1200 is executed at least partially on a computer system (e.g., a computer system described elsewhere in this disclosure). In at least one embodiment, logic (e.g., hardware, software, or a combination of hardware and software) executes process 1200.
[0148] In at least one embodiment, one or more processors use process 1200, for example, to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least partially by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or otherwise perform the operations described herein. In at least one embodiment, as an example, a set of instructions is stored on a machine-readable medium (e.g., non-transitory) that, if executed by one or more processors, causes one or more processors to perform process 1200, for example, to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least partially by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or otherwise perform the operations described herein. In at least one embodiment, process 1200 is included... Figures 1 to 17B and / or Figure 23 The process shown, and / or otherwise includes Figures 1 to 17B and / or Figure 23The process shown is used to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least in part by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or otherwise perform the operations described herein. In at least one embodiment, Figures 1 to 17B and / or Figure 23 The system execution process 1200 shown herein includes, for example, classifying one or more log entries to obtain one or more classified log entries; obtaining combined information by at least partially combining one or more classified log entries and telemetry information; classifying the combined information using at least one machine learning process; and / or otherwise performing the operations described herein. In at least one embodiment, Figures 17A to 22 The hardware used in the example 1200 is for, for example, classifying one or more log entries to obtain one or more classified log entries; obtaining combined information at least in part by combining at least one or more classified log entries and telemetry information; classifying the combined information using at least one machine learning process; and / or otherwise performing the operations described herein.
[0149] Figure 13 This is a flowchart illustrating a process 1300 for classifying one or more logs according to at least one embodiment. A processor performing process 1300 may receive encoded logs in block 1304, determine classifications using a classifier in block 1308, update one or more known classifications (e.g., predefined event labels) (e.g., used by an encoder) in block 1310, encode one or more logs with the classifications in block 1312, provide the encoding in block 1314, and / or perform one or more operations described herein, or a combination thereof. In at least one embodiment, the processor begins process 1300 upon being invoked and / or upon receiving encoded logs in block 1304. In at least one embodiment, process 1300 is performed by encoder 1018 and / or log event classification model 1006.
[0150] Upon receiving the encoded log in box 1304, the processor proceeds to decision box 1306. In at least one embodiment, the decision in decision box 1306 is "yes" if the log's classification is known, and "no" otherwise. If the decision in decision box 1306 is "yes," the processor executing process 1300 encodes one or more logs with the known classification in box 1312. As an example, the classification is known if it is included in one or more predefined event labels. For example, the classification is unknown if the encoded log does not match or correspond to a predefined event label, or if the predefined event label includes a conflicting classification of the encoded log. If the decision in decision box 1306 is "no," the processor executing process 1300 proceeds to box 1308 to determine the classification using one or more classifiers trained at least in part on a similarity loss determined using model results obtained for encoding two or more vectors (e.g., encoder 1408 and / or log event classification model 1006). For example, the processor executing process 1300 may use a classifier trained to determine the classification based at least in part on a similarity loss (e.g., cosine similarity loss) calculated about the model results obtained for encoding two or more vectors. For example, the model may be trained using process 1600. For example, system 1400 and / or system 1500 may use a classifier trained at least in part on a similarity loss (e.g., cosine similarity loss) calculated about the model results obtained for encoding two or more vectors in box 1308 to determine the classification.
[0151] After the processor in execution process 1300 determines the classification using the classifier in block 1308, it can proceed to block 1310 to update the classification operation 1004 (e.g., using encoder 1018, see [link]). Figure 10 The processor may use one or more known categories (e.g., predefined event labels). As an example, the processor may update one or more known categories of the encoder offline in box 1310 and / or update one or more known categories of the encoder by submitting one or more categories obtained in box 1308 to the domain supervisor for review, which in turn may update the known categories. Following box 1310, the processor executing process 1300 may proceed to box 1312 to encode the log (e.g., the encoded log received at box 1304) using the categories determined in box 1308. In box 1312, the log (e.g., the encoded log received at box 1304) may be encoded to include indications of alarms, ignore, review categories, and / or other information included in, associated with, and / or inferred from the log.
[0152] After encoding one or more logs with one or more categories (e.g., codes) in box 1312, the processor executing process 1300 may proceed to box 1314, where the processor provides the encoding (e.g., to one or more processors, processes, services, etc.) as the output of process 1300. At box 1314, the processor executing process 1300 may provide the encoding including the categories obtained in box 1308, perform one or more operations described herein, iterate one or more steps of process 1300, combinations thereof, otherwise perform the operations described herein, and / or terminate. In at least one embodiment, process 1300 terminates after box 1314.
[0153] In at least one embodiment, part or all of process 1300 (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and is implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more application programs) jointly executed on one or more processors by hardware, software, or a combination thereof. In at least one embodiment, the code is stored in the form of a computer program on a computer-readable storage medium comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of the computer-readable instructions available for executing process 1300 are not stored using only transient signals (e.g., propagating transient electrical or electromagnetic transmissions). In at least one embodiment, the non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver of transient signals. In at least one embodiment, process 1300 is executed at least partially on a computer system (e.g., a computer system described elsewhere in this disclosure). In at least one embodiment, logic (e.g., hardware, software, or a combination of hardware and software) executes process 1300.
[0154] In at least one embodiment, one or more processors use process 1300, for example, to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least partially by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or otherwise perform the operations described herein. In at least one embodiment, as an example, a set of instructions is stored on a machine-readable medium (e.g., non-transitory) that, if executed by one or more processors, causes one or more processors to perform process 1300, for example, to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least partially by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or otherwise perform the operations described herein. In at least one embodiment, process 1300 is included... Figures 1 to 17B and / or Figure 23 The process shown, and / or otherwise includes Figures 1 to 17B and / or Figure 23 The process shown herein is used to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least in part by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or to otherwise perform the operations described herein.
[0155] In at least one embodiment, Figures 1 to 17B and / or Figure 23 The system execution process 1300 shown herein includes, for example, classifying one or more log entries to obtain one or more classified log entries; obtaining combined information by at least partially combining one or more classified log entries and telemetry information; classifying the combined information using at least one machine learning process; and / or otherwise performing the operations described herein. In at least one embodiment, Figures 17A to 22 The hardware used in the example 1300 is for, for example, classifying one or more log entries to obtain one or more classified log entries; obtaining combined information at least in part by combining at least one or more classified log entries and telemetry information; classifying the combined information using at least one machine learning process; and / or otherwise performing the operations described herein.
[0156] Figure 14This is a block diagram illustrating a system 1400 including an encoder 1408 according to at least one embodiment, the encoder 1408 being trained to generate encodings of one or more logs at least in part based on a similarity loss. In at least one embodiment, the system 1400 includes one or more processors 1406, an encoder 1408 trained using a similarity loss, and / or one or more downstream applications 1414. One or more downstream applications 1414 may include anomaly detection 1414A (e.g., by executing anomaly detection model 1014, see [link]). Figure 10 Event prediction 1414B, root cause analysis 1414C (e.g., root cause analysis 1016, see...) Figure 10 The processor 1406 generates observations 1414D and / or executes one or more classifiers 122. In at least one embodiment, the processor 1406 includes a processor 1722 (see [link to processor 1414D]). Figure 17B The encoder 1408 may include one or more neural networks NN1, one or more neural networks NN2, and / or one or more classifiers 122.
[0157] Input 1402 may include one or more inputs 102, 202, 302 and / or 1502 (see...) Figure 1 , Figure 2 , Figure 3 and / or Figure 15 Input 1402 may include one or more logs, log lines, log sequences, tags representing one or more log event codes 612 and position codes 614, logs 702 and / or embedding vectors 704, log line stream 1002A, input to the original log 1102, log lines 1404, log line pairs, and / or other inputs described herein (see [link to documentation]). Figures 1 to 16 One or more logs 104 may include information such as text data 104A (e.g., text data 206A and / or 306A), numerical data 104B (e.g., numerical data 206B and / or 306B), and / or categorical data 104C (e.g., categorical data 206C and / or 306C). Input 1402 may include topology information, telemetry information, and / or metadata.
[0158] System 1400 can execute process 1600, for example, to fine-tune a neural network to encode logs using similarity scores. Since the fine-tuned neural network can be trained without task-specific labels, it can be easily trained to encode logs for different types of tasks (e.g., as input to other neural networks and / or machine learning processes), with minimal labeling for small sets of events. System 1400 can use cosine similarity loss to fine-tune encoder 1408 using semantic similarity with respect to multiple pairs of log entries. After encoder 1408 is trained, the results can be used as input to downstream processes (e.g., one or more neural networks) that classify the encoded log entries. For example, encoder 1408 could be a log event classification model 1006 that classifies input as “ignored” or “alarm.” If previously unseen log entries are encoded, log event classification model 1006 can classify the encoded unseen log entries into the same class as previously seen and encoded similar log entries. Therefore, unlike self-supervised learning, a fixed vocabulary is not required. System 1400 includes an encoder 1408 trained using a similarity loss, such as the encoder trained by system 1500. The encoder 1408 trained using the similarity loss (e.g., encoder 1512) can generate one or more outputs 1410, such as generated semantic codes 1412. Semantic codes 1412 can include vector encodings (e.g., tensors). The generated semantic codes 1412 can be associated with the similarity of one or more logs (e.g., represented as similarity scores), or information indicating whether a log line is alerted or ignored for anomaly detection (e.g., using log event classification model 1006). The generated semantic codes 1412 can be used in conjunction with one or more of the downstream applications 1414.
[0159] In at least one embodiment, system 1400 includes a collection of one or more hardware and / or software computing resources having instructions that, when executed, perform one or more communication processes, such as those described herein. In at least one embodiment, system 1400 is a software program, an application program, and / or variations thereof, executing on computer hardware. In at least one embodiment, one or more processes of system 1400 are executed by any suitable processing system or unit (e.g., a graphics processing unit (GPU), a general-purpose GPU (GPGPU), a parallel processing unit (PPU), a central processing unit (CPU), a data processing unit (DPU) (described below)) in any suitable manner, including sequentially, in parallel, and / or variations thereof. In at least one embodiment, system 1400 uses machine learning training frameworks (e.g., PYTORCH, TENSORFLOW, BOOST, CAFFE, MICROSOFT COGNITIVE TOOLKIT / CNTK, MXNET, CHAINER, KERAS, DEEPLEARNING4J) and / or other training frameworks to implement and perform the operations described herein to train a neural network to encode at least one vector associated with a log and / or otherwise perform the operations described herein. In at least one embodiment, as an example, training the neural network model includes using a server (e.g., an NVIDIA DGX server) that also includes at least a GPU (e.g., AMD MI200, VEGAL10, VEGO20, and ARCTURUS), an optimizer (e.g., ADAM OPTIMIZER), or a discriminator architecture (e.g., a discriminator architecture from face-vid2vid for training with GAN loss).
[0160] In at least one embodiment, system 1400 includes modules (e.g., modules 1724-1730, see below). Figure 17BThe system executes a neural network to train the neural network to encode at least one vector associated with a log. In at least one embodiment, the module includes any combination of logic (e.g., software, hardware, firmware) and / or circuitry configured to perform the function. In at least one embodiment, the module includes one or more circuitry that form part of a larger system (e.g., an integrated circuit (IC), a system-on-a-chip (SoC), a central processing unit (CPU), a graphics processing unit (GPU), a data processing unit (DPU), etc.). In at least one embodiment, the controller includes any combination of logic (e.g., software, hardware, firmware) and / or circuitry configured to perform the function. In at least one embodiment, the software includes software packages, code, programming languages, drivers, instructions, instruction sets, or some combination thereof. In at least one embodiment, the hardware includes hardwired circuitry, programmable circuitry, state machine circuitry, fixed-function circuitry, execution unit circuitry, firmware storing instructions executed by the programmable circuitry, or some combination thereof.
[0161] In at least one embodiment, system 1400 includes one or more logic units. In at least one embodiment, the logic unit includes firmware logic, hardware logic, or some combination thereof configured to provide any functionality as further described herein. In at least one embodiment, the logic unit includes circuitry that forms part of a larger system (e.g., IC, SoC, CPU, GPU, DPU). In at least one embodiment, the logic unit includes logic circuitry for implementing firmware and / or hardware for executing a neural network to train the neural network to encode at least one vector associated with a log.
[0162] In at least one embodiment, system 1400 includes one or more engines. In at least one embodiment, an engine includes modules and / or logic units as further described herein. In at least one embodiment, a component includes modules and / or logic units as further described herein. In at least one embodiment, an engine includes software logic, firmware logic, hardware logic, or some combination thereof configured to provide any functionality as further described herein. In at least one embodiment, a component includes software logic, firmware logic, hardware logic, or some combination thereof configured to provide any functionality as further described herein. In at least one embodiment, operations performed by hardware and / or firmware may alternatively be implemented via software modules, which may be embodied as software packages, code, and / or instruction sets. In at least one embodiment, a logic unit may also utilize a portion of software to implement its functionality.
[0163] In at least one embodiment, system 1400 includes one or more processors for encoding at least one log message using at least one neural network, said neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining a metric indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. In at least one embodiment, system 1400 is included Figures 1 to 17B and / or Figure 23 The system shown, and / or otherwise includes Figures 1 to 17B and / or Figure 23 The system shown is used to encode at least one log message using at least one neural network, the neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. In at least one embodiment, system 1400 performs... Figures 1 to 17B and / or Figure 23 The process illustrated herein includes one or more procedures, such as encoding at least one log message using at least one neural network, said neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. In at least one embodiment, system 1400 includes Figures 17A to 22One or more pieces of hardware, as shown, are used, for example, to encode at least one log message using at least one neural network, which is trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein.
[0164] Figure 15 The diagram illustrates a block diagram of a system 1500 that trains one or more encoders, at least in part, based on cosine similarity loss, according to at least one embodiment. In at least one embodiment, system 1500 uses similarity loss to train encoder 1512, such as log event classification model 1006, and / or one or more encoders 1408. In at least one embodiment, system 1500 implements classification function 116 (see [link to documentation]). Figure 1 ).
[0165] System 1500 can create a training dataset 1504 of log line pairs 1504A and associated similarity scores 1504B. One of the similarity scores 1504B can be assigned to each pair of log entries (e.g., log line pair 1504A) using, for example, domain knowledge and / or terms included in the log entries. As an example, before assigning similarity scores, log line pair 1510, which includes log lines 1510A and 1510B, is preprocessed (e.g., using log preprocessing 1002), such as to clean the content (e.g., remove irrelevant information) and / or extract useful information. Useful information may include time, timestamps, identification information, and / or a description of the log entry content. Useful information is extracted as parameter values. Parameter values can also be extracted from metadata, such as event priority or message type. Each log entry is then encoded into a vector using the extracted parameter values. At this point, the encoded log entries and associated similarity scores can be used to fine-tune the neural network.
[0166] For each log line pair 1510, the associated similarity score 1504B indicates the level of similarity between the two log entries (log lines 1510A and 1510B, e.g., log events). As an example, the similarity score 1504B can be a value within a range (e.g., 1-5) and can be used to rank the similarity between different log entry pairs (e.g., log lines 1510A and 1510B). Continuing this example, a similarity score of 1504B of 5 could indicate the most similar (e.g., identical) log line 1510, while a similarity score of 1 could indicate a less similar (e.g., opposite) log line 1510. As an example, log line 1510 (e.g., log entry) may include a text descriptor associated with an event (e.g., INFO or WARN) that can be categorized by an anomaly detector into a priority level (e.g., low priority or high priority) (e.g., "INFO" = low priority, "WARN" = high priority). As an example, two log entries 1510A and 1510B with a high similarity score may include a first log entry containing the text descriptor "INFO" and a second log entry containing another low-priority descriptor. As another example, two log entries 1510A and 1510B with a low similarity score of 1504B would include a first log entry 1510A with the text descriptor "INFO" (e.g., low priority) and a second log entry 1510B with the text descriptor "WARN" (e.g., high priority).
[0167] During training and / or fine-tuning, the neural network (e.g., encoder 1512) may receive a vectorized training dataset 1504 as input 1502, which includes log entry pairs (e.g., log line pairs 1504A) with their associated similarity scores 1504B. Encoder 1512 may encode the log line pairs 1510 to obtain two vector codes 1516A and 1516B for log lines 1510A and 1510B, respectively. Processor 1506 may use the vector codes 1516A and 1516B to perform a loss function (cosine similarity loss function) that generates loss values (e.g., cosine similarity loss). Encoder 1512 may encode each log line pair 1504A in the training dataset 1504 for various different configurations of encoder 1512 (e.g., different sets of parameter values, weight values, etc.). For each of these configurations of encoder 1512, processor 1506 obtains an aggregated loss value by aggregating (e.g., calculating the total, averaging, etc.) the loss values obtained for log line pair 1504A, obtains a aggregated similarity score by aggregating (e.g., calculating the total, averaging, etc.) the similarity scores associated with log line pair 1504A, compares the aggregated loss value with the aggregated similarity score, and selects the configuration that results in the smallest difference between the aggregated loss value and the aggregated similarity score. The selected configuration can be used by encoder 1512 at deployment time, for example, by performing backpropagation to update one or more neural network weights, and then using encoder 1512 to perform one or more inference operations.
[0168] As an example, a cosine similarity loss 1518 is calculated using the cosine (θ) of two vectors. A neural network encoder 1512 can encode vectorized log line pairs into vector codes 1516A and 1516B, and compute a cosine similarity loss 1518 for both vector codes 1516A and 1516B. Next, the cosine similarity (e.g., the cosine similarity loss 1518) can be compared to a similarity score 1504B associated with the log line pair 1510. A processor 1506 can generate an output 1520, such as a fine-tuned encoder 1522 with a selected configuration determined using the similarity loss. The configuration of the fine-tuned encoder 1522 (e.g., one or more model weight values) can be selected by identifying configurations that result in smaller differences between the cosine similarity and the similarity score associated with the log line pair 1510.
[0169] After determining the configuration of encoder 1512 (e.g., model weights), encoder 1512 can be deployed and used to infer the encoding of vectorized log entries. These encodings can be provided to one or more other processes, such as one or more other neural networks (e.g., anomaly detection model 1014). For example, the encodings can be provided to a neural network trained to detect anomalies, which can infer whether each encoding indicates whether an anomaly was recorded in each log.
[0170] As an example, encoder 1512 can provide encoded log entries to an anomaly detection model 1014 to classify whether a log entry (e.g., a log event) is sufficiently dissimilar to other log entries (e.g., log events) to meet the criteria for being an anomaly. The information included in log events can change and evolve over time. Therefore, the model may encounter new log entries or variations of existing log entries on which it has not previously trained (referred to as unseen log entries). Encoder 1512, trained using the similarity loss described herein, can generate vector codes for new log entries, and processor 1506 can compute a cosine similarity loss between the vector code generated for the new log entry and the vector codes associated with log entries having known classifications. Processor 1506 can assign the classification associated with the new log entry to the log entry with which it has the minimum cosine similarity loss. The value of the cosine similarity loss can be used to determine whether to alert a domain expert or ignore a domain expert (e.g., if the cosine similarity loss is greater than a threshold). For example, if a domain expert instructs that a first code (e.g., classification) be assigned to a log entry, and encoder 1512 generates a second code, the magnitude of the loss value (e.g., cosine similarity loss) between the first and second codes can be used to determine whether to alert the domain expert or ignore the domain expert's code. For example, if the loss value exceeds a first threshold, the domain expert can be ignored, and / or if the loss value exceeds a second threshold, an alert can be issued to the domain expert.
[0171] The classification returned by the domain expert for previously unseen log entries (e.g., log events) can be further used to refine (e.g., through backpropagation) the encoder 1512 and / or the anomaly detection model (e.g., by updating the weight values and / or other configuration settings of the encoder 1512 and / or the anomaly detection model), for example, by using similarity scores between the new log entries and other log entries. For example, the classification provided by the domain expert can be paired with one or more log lines, with a similarity score added to the similarity score 1504B of each newly added pair, and the new pair added to log line pair 1504A in the training dataset 1504. The classification provided by the domain expert can be added to predefined event labels and used by classification operation 1004 (e.g., encoder 1018).
[0172] In at least one embodiment, system 1500 includes one or more processors for encoding at least one log message using at least one neural network, said neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. In at least one embodiment, system 1500 is included Figures 1 to 17B and / or Figure 23 The system shown, and / or otherwise includes Figures 1 to 17B and / or Figure 23 The system shown is used to encode at least one log message using at least one neural network, said neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. In at least one embodiment, system 1500 performs... Figures 1 to 17B and / or Figure 23The process illustrated herein includes one or more procedures, for example, for encoding at least one log message using at least one neural network, said neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. In at least one embodiment, system 1500 includes Figures 17A to 22 One or more pieces of hardware, as shown, are used, for example, to encode at least one log message using at least one neural network, which is trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein.
[0173] Figure 16 This is a flowchart illustrating a process 1600 for training and / or fine-tuning a model (e.g., neural network NN1, neural network NN2, encoder 1512, log event classification model 1006, and / or the like) according to at least one embodiment. A processor performing process 1600 may receive a training dataset including encoded log pairs associated with similarity score inputs at box 1604, obtain similarity scores and encoded log pairs from the training dataset at box 1606, generate a first vector encoding and a second vector encoding using the model at least partially based on the encoded log pairs at box 1607, generate a similarity loss between the first and second vector encodings at box 1608, determine one or more metrics for indicating similarity between the individual pairs at box 1614 after all pairs in the training set have been encoded, select a model configuration based on the metrics determined for one or more different model configurations at box 1620, and / or perform one or more operations described herein, or a combination thereof. In at least one embodiment, the processor invokes procedure 1600 and / or receives training dataset input (e.g., training dataset 1504) in box 1604. In at least one embodiment, procedure 1600 is executed by system 1500 to train encoder 1512. The training dataset received as input in box 1604 may include one or more log line pairs 1054A and one of the similarity scores 1504B corresponding to each pair.
[0174] After receiving the training dataset input in box 1604, the processor executing process 1600 (e.g., processor 1506) may obtain a similarity score (from similarity score 1504B) and associated encoded log pairs (e.g., log line pairs 1510) from the training dataset in box 1606. Then, in box 1607, the processor may generate an inference result pair (e.g., a first vector encoding 1516A and a second vector encoding 1516B) using the model, at least in part, based on the encoded log pairs obtained in box 1606. Next, in box 1608, the processor may generate a similarity loss (e.g., cosine similarity loss 1518) between the first and second inference results in the inference result pair. Then, at box 1610, the processor determines a metric (e.g., a performance measurement) to indicate the similarity between the similarity score and the similarity loss, such as a score, vector, integer, and / or other indication of the metric. In at least one embodiment, process 1600 may include providing the metric to a processor, process, and / or service performing one or more of the operations described herein.
[0175] Then, at decision box 1612, the processor determines whether the training dataset includes more encoded log pairs. If the training dataset includes more encoded log pairs, the decision at decision box 1612 is "yes". Otherwise, the decision at decision box 1612 is "no". When the decision at decision box 1612 is "yes", the processor returns to box 1606 to obtain another encoded log pair and its associated similarity score from the training dataset. On the other hand, when the decision at decision box 1612 is "no", at box 1614, the processor aggregates (e.g., calculates the total, calculates the average, etc.) the metric determined in box 1610 for the encoded log pairs in the training dataset. This metric is associated with the current configuration of the model that generates inference result pairs for each encoded log pair included in the training dataset, as shown in box 1607.
[0176] Then, at decision box 1616, the processor determines whether to modify the model. If the processor determines to modify the model, the decision at decision box 1616 is "yes". Otherwise, the decision at decision box 1616 is "no". When the decision at decision box 1616 is "yes", the processor modifies the model at box 1618 and then returns to box 1606 to begin processing each encoded log pair included in the training dataset with the modified model. When the decision at decision box 1616 is "no", at box 1620, the processor selects a model configuration (e.g., weight values, parameter values, and / or other settings) for which the aggregate metric determined at box 1614 indicates the expected amount of similarity between the similarity score and the similarity loss (e.g., the maximum amount of similarity). In at least one embodiment, process 1600 terminates after box 1620. After execution process 1600, a model (e.g., as neural network NN1, neural network NN2, encoder 1512, log event classification model 1006, and / or the like) can be deployed and used to encode one or more encoded logs. For example, the processor executing process 1600 can use backpropagation to update the weights of one or more neural networks in the model, and then use the model to perform one or more inference operations.
[0177] The processor of process 1600 fine-tunes the encoder to encode log lines. This model can be a language model pre-trained on a semantic similarity task (e.g., LLM), and process 1600 can be used to fine-tune this model using log line pairs, where each log line pair is at least partially based on a similarity score assigned using domain knowledge, language, or through explicit annotation. Process 1600 can leverage domain knowledge with relatively little manual effort (e.g., partial annotation is sufficient to generate many pairs) and can better capture the semantic meaning of log messages. Because the encoding is trained in a general manner, it can be used for a variety of downstream log analysis tasks.
[0178] Process 1600 includes training to fine-tune one or more language models (e.g., encoders) to capture the semantic meaning of one or more log events, which can be done in a task-general manner (which may include minimal annotation effort). Process 1600 can generate results encoded for multiple log analysis tasks and generated with a model that does not involve large amounts of data. Process 1600 can train one or more log analysis models to learn to encode log lines as part of a specific downstream task (e.g., anomaly detection), for example in a task-agnostic pre-trained framework that assigns similarity scores (based on domain expert assumptions or explicitly) to log line pairs and fine-tunes the language model with a semantic similarity task.
[0179] In at least one embodiment, part or all of process 1600 (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and is implemented as code (e.g., computer-executable instructions, one or more computer programs, or one or more application programs) jointly executed on one or more processors by hardware, software, or a combination thereof. In at least one embodiment, the code is stored in the form of a computer program on a computer-readable storage medium comprising a plurality of computer-readable instructions executable by one or more processors. In at least one embodiment, the computer-readable storage medium is a non-transitory computer-readable medium. In at least one embodiment, at least some of the computer-readable instructions available for executing process 1600 are not stored using only transient signals (e.g., propagating transient electrical or electromagnetic transmissions). In at least one embodiment, the non-transitory computer-readable medium does not necessarily include non-transitory data storage circuitry (e.g., buffers, caches, and queues) within a transceiver of transient signals. In at least one embodiment, process 1600 is executed at least partially on a computer system (e.g., a computer system described elsewhere in this disclosure). In at least one embodiment, logic (e.g., hardware, software, or a combination of hardware and software) executes process 1600.
[0180] In at least one embodiment, one or more processors use process 1600, for example, to encode at least one log message using at least one neural network, the neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. In at least one embodiment, as an example, a set of instructions is stored on a machine-readable medium (e.g., non-transitory) that, if executed by one or more processors, causes one or more processors to perform process 1600, for example, to encode at least one log message using at least one neural network, the neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein.
[0181] In at least one embodiment, process 1600 is included Figures 1 to 17B and / or Figure 23 The process shown, and / or otherwise includes Figures 1 to 17B and / or Figure 23 The process shown is used to encode at least one log message using at least one neural network, which is trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. In at least one embodiment, Figures 1 to 17B and / or Figure 23 The system execution process 1600 shown herein, for example, is used to encode at least one log message using at least one neural network, which is trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. In at least one embodiment, Figures 17A to 22 One or more hardware usage processes 1600 shown herein are used, for example, to encode at least one log message using at least one neural network, which is trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein.
[0182] In the following description, numerous specific details are set forth to provide a more thorough understanding of at least one embodiment. However, it will be apparent to those skilled in the art that the inventive concept can be practiced without one or more of these specific details.
[0183] Figure 17AAn example of a system 1700 according to at least one embodiment is shown. The system 1700 includes one or more drivers and / or one or more runtimes (shown as reference numeral 1704), which include one or more libraries 1706 for providing one or more application programming interfaces (“APIs”) 1710. In at least one embodiment, the system 1700 includes driver 1704 and / or runtime 1704, which includes library 1706 for providing API 1710. In at least one embodiment, API 1710 is a software instruction set that, if executed, causes one or more processors (e.g., ...) to... Figure 17B The processor 1722 shown performs one or more computational operations. In at least one embodiment, one or more APIs 1710 are distributed or otherwise provided as part of one or more components of one or more libraries 1706, one or more runtimes 1704, one or more drivers 1704, and / or any other grouping of software and / or executable code further described herein. In at least one embodiment, one or more APIs 1710 perform one or more computational operations in response to a call from one or more software programs 1702.
[0184] In at least one embodiment, one or more software programs 1702 are software modules and / or include one or more software modules. In at least one embodiment, the software modules are... Figure 17B The following are further shown non-exclusively as one or more modules 1724-1730 and described therewith. In at least one embodiment, one or more software programs 1702 are collections of software code, commands, instructions, and / or other text sequences for instructing a computing device (e.g., executing a neural network to encode and / or classify log messages) to perform one or more computational operations and / or invoke one or more other sets of instructions (e.g., API 1710 or API function 1712) for execution by the computing device. In at least one embodiment, the functionality provided by one or more APIs 1710 includes API functions 1712, such as those functions that can be used to accelerate one or more portions of software program 1702 using one or more parallel processing units (PPUs) (e.g., graphics processing units (GPUs)).
[0185] In at least one embodiment, one or more APIs 1710 are one or more hardware interfaces of one or more circuits for performing one or more computational operations. In at least one embodiment, one or more APIs 1710 described herein are implemented for performing combinations Figures 1 to 16The description refers to one or more circuits employing one or more techniques. In at least one embodiment, one or more software programs 1702 include instructions that, if executed, cause one or more hardware devices and / or circuits to perform a combination. Figures 1 to 16 One or more techniques are further described. In at least one embodiment, system 1700 includes information regarding... Figure 1 The system 100 described includes one or more or all of its components, and the system 1700 can perform one or more or all of the processes and / or operations performed by the system and its components. In at least one embodiment, the system 1700 includes... Figure 2 The system 200 described includes one or more or all of its components, and the system 1700 can perform one or more or all of the processes and / or operations performed by the system and its components. In at least one embodiment, the system 1700 includes... Figure 5 The system 500 is described as having one or more or all of its components, and the system 1700 is capable of performing one or more or all of the processes and / or operations performed by the system and its components. In at least one embodiment, the system 1700 includes... Figure 10 The system 1000 is described as having one or more or all of its components, and the system 1700 is capable of performing one or more or all of the processes and / or operations performed by the system and its components. In at least one embodiment, the system 1700 includes... Figure 14 The system 1400 is described as having one or more or all of its components, and the system 1700 can perform one or more or all of the processes and / or operations performed by the system and its components.
[0186] In at least one embodiment, software program 1702 (e.g., a user-implemented software program) utilizes one or more APIs of API 1710 to perform various computational operations, such as memory reservation, matrix multiplication, arithmetic operations, and / or any computational operations performed by a PPU (e.g., a GPU), as further described herein. In at least one embodiment, function 1712 includes a set of callable functions provided by one or more APIs 1710, referred herein as APIs, API functions, software functions, and / or functions, each performing one or more computational operations, such as computational operations related to parallel computing. In at least one embodiment, one or more APIs 1710 cause a neural network to encode and / or classify log messages, and / or perform the operations described herein (e.g., in conjunction with...). Figures 1 to 16 Other operations described.
[0187] In at least one embodiment, one or more software programs 1702 interact with or otherwise communicate with one or more APIs 1710 to use one or more processors (e.g., Figure 17B The processor 1722 shown herein (e.g., one or more PPUs, such as a GPU) performs one or more computational operations. In at least one embodiment, the one or more computational operations using one or more PPUs include at least one or more sets of computational operations that will be accelerated by being executed at least partially by said one or more PPUs. In at least one embodiment, one or more software programs 1702 interact with one or more APIs 1710 to enable a neural network to encode and / or classify log messages, and / or perform the operations described herein (e.g., in conjunction with...). Figures 1 to 16 Other operations described.
[0188] In at least one embodiment, the interface is software instructions that, when executed, provide access to one or more functions 1712 provided by one or more APIs 1710. In at least one embodiment, when a software developer compiles one or more software programs 1702 in conjunction with one or more libraries 1706, the one or more software programs 1702 use a native interface, the one or more libraries 1706 including or otherwise providing access to one or more APIs 1710. In at least one embodiment, one or more software programs 1702 are statically compiled in conjunction with one or more pre-compiled libraries of library 1706 and / or uncompiled source code, the uncompiled source code including instructions for executing one or more APIs 1710. In at least one embodiment, one or more software programs 1702 are dynamically compiled, and the dynamically compiled software programs are linked to one or more pre-compiled libraries of library 1706 using a linker, library 1706 including one or more APIs 1710.
[0189] In at least one embodiment, when a software developer executes a software program that utilizes at least one library 1706 (which includes one or more APIs 1710) or otherwise communicates with the at least one library 1706 via a network or other remote communication medium, one or more software programs 1702 use a remote interface. In at least one embodiment, one or more libraries 1706 (which include one or more APIs 1710) will be executed by a remote computing service (e.g., a computing resource service provider). In at least one embodiment, one or more libraries 1706 including one or more specific APIs (APIs in API 1710) will be executed by any other computing host that provides the specific APIs to one or more software programs 1702.
[0190] In at least one embodiment, a processor that executes or uses one or more specific software programs of software program 1702 (e.g., Figure 17B The processor 1722 shown calls, uses, executes, and / or otherwise implements one or more APIs 1710 to allocate and otherwise manage memory 1714 to be used by a particular software program. In at least one embodiment, one or more particular software programs of software program 1702 utilize one or more APIs 1710 to allocate and otherwise manage memory 1714 to be used by one or more portions of a particular software program that will be accelerated using one or more PPUs (e.g., GPUs) or any other accelerators or processors further described herein. In at least one embodiment, one or more software programs 1702 request one or more neural networks to perform signal processing using one or more functions 1712 provided by one or more APIs 1710. In at least one embodiment, memory is implemented to perform one or more operations to combine Figures 1 to 16 The processor that encodes and / or classifies one or more log messages includes memory 1714.
[0191] In at least one embodiment, one or more APIs of API 1710 are APIs for facilitating parallel computing. In at least one embodiment, one or more APIs of API 1710 are any other APIs further described herein. In at least one embodiment, one or more APIs of API 1710 are provided by one or more drivers of driver 1704 and / or one or more runtimes of runtime 1704. In at least one embodiment, one or more APIs of API 1710 are provided by a CUDA user-mode driver. In at least one embodiment, one or more APIs of API 1710 are provided by a CUDA runtime. In at least one embodiment, one or more drivers 1704 are data values and software instructions that, if executed, perform and / or otherwise facilitate the operation of one or more functions 1712 of one or more APIs 1710 during the loading and execution of one or more portions of at least one software program 1702. In at least one embodiment, one or more runtimes 1704 are data values and / or software instructions that, if executed, perform or otherwise facilitate the operation of one or more functions 1712 of one or more APIs 1710 during the execution of at least one software program 1702. In at least one embodiment, one or more specific software programs of software program 1702 utilize one or more APIs 1710 implemented and / or otherwise provided by one or more drivers 1704 and / or one or more runtimes 1704 to perform combinatorial arithmetic operations by the specific software programs during execution by one or more PPUs (e.g., GPUs).
[0192] In at least one embodiment, one or more software programs 1702 utilize one or more APIs 1710 provided by one or more drivers 1704 and / or one or more runtimes 1704 to perform combinatorial arithmetic operations on one or more PPUs (e.g., GPUs). In at least one embodiment, one or more APIs 1710 provide combinatorial arithmetic operations via one or more drivers 1704 and / or one or more runtimes 1704, as described above. In at least one embodiment, one or more software programs 1702 utilize one or more APIs 1710 provided by one or more drivers 1704 and / or one or more runtimes 1704 to allocate or otherwise reserve one or more blocks of memory 1714 for one or more PPUs (e.g., GPUs). In at least one embodiment, one or more software programs 1702 utilize one or more APIs 1710 provided by one or more drivers 1704 and / or one or more runtimes 1704 to allocate or otherwise reserve blocks of memory 1714.
[0193] In at least one embodiment, to improve the usability and / or performance of one or more specific software programs in software program 1702, one or more portions of the specific software programs will be accelerated by one or more PPUs (e.g., GPUs). In at least one embodiment, one or more functions 1712 receive one or more input parameters indicating one or more inputs to one or more neural networks and / or other data to be utilized by the neural networks, such as one or more hyperparameters of the neural networks. In at least one embodiment, the input parameters include one or more inputs and / or other data. In at least one embodiment, the input parameters include one or more pointers to one or more memory locations storing the inputs and / or other data.
[0194] In at least one embodiment, system 1700 includes at least one processor (e.g., Figure 17B The processor 1722 shown includes one or more circuitry for executing one or more software programs to combine two or more APIs 1710 into a single API. In at least one embodiment, system 1700 includes at least one processor (e.g., Figure 17B The processors shown (1722) use one or more APIs 1710 to enable a neural network to encode and / or classify one or more log messages, and / or otherwise perform the operations described herein. In at least one embodiment, system 1700 includes at least one processor (e.g., Figure 17BThe processor 1722 shown uses one or more APIs 1710 to execute... Figures 1 to 16 One or more of the operations shown and / or described therein, for example Figure 4 , Figure 9 , Figure 11-13 and / or Figure 16 Or any one or more of the processes shown in the above-described portion. In at least one embodiment, system 1700 includes at least one processor (e.g., Figure 17B The processor 1722 shown is used to execute one or more functions 1712, for example, in combination with Figures 1 to 16 The function described. In at least one embodiment, one or more API 1710s will be combined with Figures 18A to 22 The aforementioned hardware execution.
[0195] Figure 17B Block diagram 1720 illustrates example processor 1722 and modules 1724-1730 according to at least one embodiment. Reference Figure 17B In at least one embodiment, processor 1722 may be processor 110, 502, 602, 1406 and / or 1506 (see...) Figure 1 , Figure 5 , Figure 6 , Figure 14 and / or Figure 15 ) Implementation. In at least one embodiment, processor 1722 may execute one or more processes, such as the processes described herein concerning executing a neural network to encode and / or classify one or more log messages, and / or may otherwise perform the operations described herein. In at least one embodiment, processor 1722 executes one or more processes, such as those in conjunction with the figures Figure 4 , Figure 9 , Figure 1-13 and / or Figure 16 The process described.
[0196] In at least one embodiment, processor 1722 includes one or more processors, such as combined Figures 18A to 22The processor described herein. In at least one embodiment, processor 1722 may be any suitable processing unit and / or combination of processing units, such as one or more CPUs, GPUs, DPUs, GPGPUs, PPUs, and / or variations thereof. Processor 1722 includes modules 1724-1730, which may include a neural network training module 1724 (e.g., neural network training module 504); a triplet loss module 1726; a similarity loss module 1726; a log, telemetry, and topology classification module 1728; and an anomaly detection module 1730. Modules 1724-1730 may be distributed across multiple processors that communicate via a bus, network, write-to-shared memory, and / or any suitable communication procedure (e.g., the communication procedure described herein). In at least one embodiment, modules 1724-1730 may include processor-executable instructions that implement training of a neural network to encode and / or classify one or more log messages and / or otherwise perform the operations described herein.
[0197] As used in any implementation described herein, unless the context explicitly states otherwise or expressly to the contrary, a module refers to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functionality described herein. Software may be embodied as a software package, code, and / or instruction set or instructions, while “hardware” as used in any implementation described herein may, for example, individually or in any combination, include hardwired circuitry, programmable circuitry, state machine circuitry, fixed-function circuitry, execution unit circuitry, and / or firmware storing instructions executed by programmable circuitry. Modules may be embodied collectively or individually as circuitry forming part of a larger system (e.g., integrated circuit (IC), system-on-a-chip (SoC), etc.). Modules, in conjunction with any suitable processing unit and / or combination of processing units (e.g., one or more CPUs, GPUs, GPGPUs, DPUs, PPUs, and / or variants thereof), perform one or more processes.
[0198] In at least one embodiment, as used in any implementation described herein, unless the context explicitly specifies otherwise or explicitly to the contrary, terms such as “module” and nominalized verbs (e.g., image manager, image analyzer, analysis engine, controller, and / or other terms) refer to any combination of software logic, firmware logic, hardware logic, and / or circuitry configured to provide the functionality described herein. In at least one embodiment, software may be embodied as a software package, code, and / or instruction set or instructions, and “hardware” as used in any implementation described herein may, for example, individually or in any combination, include hardwired circuitry, programmable circuitry, state machine circuitry, fixed-function circuitry, execution unit circuitry, and / or firmware storing instructions executed by programmable circuitry. In at least one embodiment, a module may be embodied collectively or individually as circuitry forming part of a larger system (e.g., integrated circuit (IC), system-on-a-chip (SoC), etc.).
[0199] In at least one embodiment, Figures 17A to 17B The system described herein is used to employ various algorithms, formulas, and processes (e.g., combining...) Figure 1 The described methods encode and / or classify one or more logs and / or otherwise perform the operations described herein. In at least one embodiment, Figures 17A to 17B The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The ones described herein are used to encode and / or classify one or more logs and / or otherwise perform the operations described herein. As an example, Figures 17A to 17B The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing one or more of the operations described herein.
[0200] As an example, Figures 17A to 17B The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message using at least one neural network, said at least one neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. As an example, Figures 17A to 17B The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least in part by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or to otherwise perform the operations described herein. As an example, Figures 17A to 17B The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message by at least partially: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code by at least partially combining the first code and the second code; and / or otherwise performing the operations described herein.
[0201] logic
[0202] Figure 18A A logic 1815 according to at least one embodiment is illustrated. As described elsewhere herein, this logic can be used in one or more devices to perform the operations discussed herein. In at least one embodiment, logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. In at least one embodiment, logic 1815 is inference and / or training logic. The following is in conjunction with... Figure 18A and / or Figure 18BDetails regarding logic 1815 are provided. In at least one embodiment, logic refers to any combination of software logic, hardware logic, and / or firmware logic for providing the functions or operations described herein, wherein the logic may collectively or individually be embodied as a circuit system forming part of a larger system (e.g., an integrated circuit (IC), a system-on-a-chip (SoC), or one or more processors (e.g., a CPU, a GPU)).
[0203] In at least one embodiment, logic 1815 may include, but is not limited to, code and / or data storage 1801 for storing forward and / or output weights and / or input / output data, and / or other parameters for configuring neurons or layers of a neural network trained and / or used for inference in one or more embodiments. In at least one embodiment, logic 1815 may include or be coupled to code and / or data storage 1801 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)). In at least one embodiment, code (such as graph code) loads weight or other parameter information into the processor ALU based on the architecture of the neural network to which the code corresponds. In at least one embodiment, code and / or data storage 1801 stores weight parameters and / or input / output data of each layer of a neural network trained or used in one or more embodiments during forward propagation of input / output data and / or weight parameters using aspects of one or more embodiments. In at least one embodiment, any portion of the code and / or data storage 1801 may be included within other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0204] In at least one embodiment, any portion of the code and / or data storage 1801 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 1801 may be a cache memory, dynamic random access memory (“DRAM”), static random access memory (“SRAM”), non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 1801 is internal or external to the processor, for example, or including DRAM, SRAM, flash memory, or some other storage type, may depend on the available on-chip versus off-chip storage, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in the inference and / or training of the neural network, or some combination of these factors.
[0205] In at least one embodiment, logic 1815 may include, but is not limited to, code and / or data storage 1805 for storing backpropagation and / or output weights and / or input / output data corresponding to neurons or layers of a neural network trained and / or used for inference in one or more embodiments. In at least one embodiment, during training and / or inference using one or more embodiments, code and / or data storage 1805 stores weight parameters and / or input / output data for each layer of the neural network trained or used in one or more embodiments during backpropagation of input / output data and / or weight parameters. In at least one embodiment, logic 1815 may include or be coupled to code and / or data storage 1805 for storing graph code or other software to control timing and / or sequence, wherein weight and / or other parameter information is loaded to configure logic including integer and / or floating-point units (collectively, arithmetic logic units (ALUs)).
[0206] In at least one embodiment, code (such as graph code) causes the architecture of the neural network corresponding to that code to load weights or other parameter information into the processor ALU. In at least one embodiment, any portion of the code and / or data storage 1805 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. In at least one embodiment, any portion of the code and / or data storage 1805 may be internal or external to one or more processors or other hardware logic devices or circuits. In at least one embodiment, the code and / or data storage 1805 may be cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other storage. In at least one embodiment, the choice of whether the code and / or data storage 1805 is internal or external to the processor, for example, including DRAM, SRAM, flash memory, or some other type of storage, may depend on the available on-chip versus off-chip storage, the latency requirements of the training and / or inference functions being performed, the batch size of the data used in the inference and / or training of the neural network, or some combination of these factors.
[0207] In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be separate storage structures. In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be the same storage structure. In at least one embodiment, code and / or data storage 1801 and code and / or data storage 1805 may be partially combined and partially separated. In at least one embodiment, any portion of code and / or data storage 1801 and code and / or data storage 1805 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory.
[0208] In at least one embodiment, logic 1815 may include, but is not limited to, one or more arithmetic logic units (“ALUs”) 1810 (including integer and / or floating-point units) for performing logical and / or mathematical operations at least in part based on training and / or inference code (e.g., graph code) or as instructed thereto, the results of which may produce activations (e.g., output values from layers or neurons within a neural network) stored in activation storage 1820, which are functions of input / output and / or weight parameter data stored in code and / or data storage 1801 and / or code and / or data storage 1805. In at least one embodiment, activations stored in activation storage 1820 are generated based on linear algebra and / or matrix-based mathematics performed by ALU 1810 in response to execution instructions or other code, wherein weight values stored in code and / or data storage 1805 and / or code and / or data storage 1801 are used as operands, and other values, such as bias values, gradient information, momentum values, or other parameters or hyperparameters, may be stored in code and / or data storage 1805 or code and / or data storage 1801 or other on-chip or off-chip storage.
[0209] In at least one embodiment, one or more processors or other hardware logic devices or circuits include one or more ALUs 1810, while in another embodiment, one or more ALUs 1810 may be external to the processor or other hardware logic device or the circuitry that uses them (e.g., a coprocessor). In at least one embodiment, ALUs 1810 may be included within an execution unit of a processor, or otherwise included in an ALU bank accessible by the execution unit of the processor, which may be within the same processor or distributed among different processors of different types (e.g., a central processing unit, a graphics processing unit, a fixed-function unit, etc.). In at least one embodiment, code and / or data storage 1801, code and / or data storage 1805, and activation storage 1820 may share a processor or other hardware logic device or circuitry, while in another embodiment, they may be in different processors or other hardware logic devices or circuitry, or in some combination of the same and different processors or other hardware logic devices or circuitry. In at least one embodiment, any portion of activation storage 1820 may be included together with other on-chip or off-chip data storage, including the processor's L1, L2, or L3 cache or system memory. Furthermore, inference and / or training code may be stored together with other code accessible to the processor or other hardware logic or circuitry, and may be retrieved and / or processed using the processor’s fetch, decode, schedule, execute, exit, and / or other logic circuitry.
[0210] In at least one embodiment, the active memory 1820 may be a cache memory, DRAM, SRAM, non-volatile memory (e.g., flash memory), or other memory. In at least one embodiment, the active memory 1820 may be entirely or partially located within or outside one or more processors or other logic circuits. In at least one embodiment, the choice of whether the active memory 1820 is internal to or external to the processor, for example, or including DRAM, SRAM, flash memory, or certain other memory types, may depend on the available on-chip versus off-chip memory, the latency requirements for performing training and / or inference functions, the batch size of the data used in inference and / or training the neural network, or some combination of these factors.
[0211] In at least one embodiment, Figure 18A The logic 1815 shown can be used in conjunction with an application-specific integrated circuit (“ASIC”), such as those from Google. Processing unit, from Graphcore TM The inference processing unit (IPU) or from Intel. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 18AThe logic 1815 shown can be used in conjunction with central processing unit (“CPU”) hardware, graphics processing unit (“GPU”) hardware, or other hardware such as field programmable gate array (“FPGA”)
[0212] In at least one embodiment, Figure 18A The system described herein is used to employ various algorithms, formulas, and processes (e.g., combining...) Figure 1 The described methods encode and / or classify one or more logs and / or otherwise perform the operations described herein. In at least one embodiment, Figure 18A The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The ones described herein are used to encode and / or classify one or more logs and / or otherwise perform the operations described herein. As an example, Figure 18A The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing one or more of the operations described herein.
[0213] As an example, Figure 18A The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message using at least one neural network, said at least one neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. As an example, Figure 18AThe description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least in part by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or to otherwise perform the operations described herein. As an example, Figure 18A The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message by at least partially: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code by at least partially combining the first code and the second code; and / or otherwise performing the operations described herein.
[0214] Figure 18B A logic 1815 according to at least one embodiment is illustrated. In at least one embodiment, the logic 1815 is inference and / or training logic. In at least one embodiment, the logic 1815 may include, but is not limited to, hardware logic, wherein computational resources, along with weight values or other information corresponding to one or more layers of neurons within a neural network, are used dedicatedly or otherwise exclusively. In at least one embodiment, Figure 18B The logic 1815 shown can be used in conjunction with application-specific integrated circuits (ASICs), such as those from Google. Processing unit, from Graphcore TM The inference processing unit (IPU) or from Intel. (e.g., "Lake Crest") processor. In at least one embodiment, Figure 18B The logic 1815 shown can be used in conjunction with central processing unit (CPU) hardware, graphics processing unit (GPU) hardware, or other hardware (e.g., field-programmable gate array (FPGA)). In at least one embodiment, logic 1815 includes, but is not limited to, code and / or data storage 1801 and code and / or data storage 1805, which can be used to store code (e.g., graph code), weight values, and / or other information, including bias values, gradient information, momentum values, and / or other parameter or hyperparameter information. Figure 18BIn at least one embodiment shown, each of code and / or data storage 1801 and code and / or data storage 1805 is associated with dedicated computing resources (e.g., computing hardware 1802 and computing hardware 1806), respectively. In at least one embodiment, each of computing hardware 1802 and computing hardware 1806 includes one or more ALUs that perform mathematical functions (e.g., linear algebraic functions) on the information stored in code and / or data storage 1801 and code and / or data storage 1805, respectively, with the results stored in active memory 1820.
[0215] In at least one embodiment, each of the code and / or data storage 1801 and 1805 and the corresponding computing hardware 1802 and 1806 corresponds to a different layer of the neural network, such that an activation obtained from one storage / computation pair 1801 / 1802 of the code and / or data storage 1801 and computing hardware 1802 is provided as input to the next storage / computation pair 1805 / 1806 of the code and / or data storage 1805 and computing hardware 1806, in order to reflect the conceptual organization of the neural network. In at least one embodiment, each storage / computation pair 1801 / 1802 and 1805 / 1806 may correspond to more than one neural network layer. In at least one embodiment, additional storage / computation pairs (not shown) may be included in logic 1815 following or paralleling the storage / computation pairs 1801 / 1802 and 1805 / 1806.
[0216] In at least one embodiment, Figure 18B The system described herein is used to employ various algorithms, formulas, and processes (e.g., combining...) Figure 1 The described methods encode and / or classify one or more logs and / or otherwise perform the operations described herein. In at least one embodiment, Figure 18B The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The ones described herein are used to encode and / or classify one or more logs and / or otherwise perform the operations described herein. As an example, Figure 18B The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23The methods described herein are used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing one or more of the operations described herein.
[0217] As an example, Figure 18B The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message using at least one neural network, said at least one neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. As an example, Figure 18B The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least in part by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or to otherwise perform the operations described herein. As an example, Figure 18B The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message by at least partially: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code by at least partially combining the first code and the second code; and / or otherwise performing the operations described herein.
[0218] Data Center
[0219] Figure 19 An example data center 1900 that can be used in at least one embodiment is shown. In at least one embodiment, the data center 1900 includes a data center infrastructure layer 1910, a framework layer 1920, a software layer 1930, and an application layer 1940.
[0220] In at least one embodiment, such as Figure 19 As shown, the data center infrastructure layer 1910 may include a resource coordinator 1912, grouped computing resources 1914, and node computing resources (“nodes CR”) 1916(1)-1916(N), where “N” represents a positive integer (which may be an integer “N” different from the integers used in other diagrams). In at least one embodiment, nodes CR 1916(1)-1916(N) may include, but are not limited to, any number of central processing units (“CPUs”) or other processors (including accelerators, field-programmable gate arrays (FPGAs), graphics processors, etc.), memory storage devices 1918(1)-1918(N) (e.g., dynamic read-only memory, solid-state storage, or disk drives), network input / output (“NW I / O”) devices, network switches, virtual machines (“VMs”), power modules, and cooling modules, etc. In at least one embodiment, one or more nodes CR 1916(1)-1916(N) may be servers having one or more of the aforementioned computing resources.
[0221] In at least one embodiment, the grouped computing resources 1914 may include individual groups of node CRs housed within one or more racks (not shown), or a plurality of racks housed within data centers (also not shown) in various geographical locations. In at least one embodiment, the individual groups of node CRs within the grouped computing resources 1914 may include computing, networking, memory, or storage resources that can be configured or allocated to support groups of one or more workloads. In at least one embodiment, several node CRs, including CPUs or processors, may be grouped within one or more racks to provide computing resources to support one or more workloads. In at least one embodiment, the one or more racks may also include any number of power modules, cooling modules, and network switches in any combination.
[0222] In at least one embodiment, resource coordinator 1912 may be configured or otherwise control one or more nodes CR1916(1)-1916(N) and / or grouped computing resources 1914. In at least one embodiment, resource coordinator 1912 may include a Software Design Infrastructure (“SDI”) management entity for data center 1900. In at least one embodiment, resource coordinator 1912 may include hardware, software, or some combination thereof.
[0223] In at least one embodiment, such as Figure 19 As shown, framework layer 1920 includes job scheduler 1922, configuration manager 1924, resource manager 1926, and distributed file system 1928. In at least one embodiment, framework layer 1920 may include a framework of software 1932 supporting software layer 1930 and / or one or more applications 1942 supporting application layer 1940. In at least one embodiment, software 1932 or application 1942 may respectively include web-based service software or applications, such as service software or applications provided by Amazon Web Services, Google Cloud, and Microsoft Azure. In at least one embodiment, framework layer 1920 may be, but is not limited to, a type of free and open-source software web application framework, such as Apache Spark™ (hereinafter referred to as "Spark") which can utilize distributed file system 1928 for large-scale data processing (e.g., "big data"). In at least one embodiment, job scheduler 1922 may include Spark drivers for facilitating the scheduling of workloads supported by the various layers of data center 1900. In at least one embodiment, configuration manager 1924 may be able to configure different layers, such as software layer 1930 and framework layer 1920 including Spark and distributed file system 1928 for supporting large-scale data processing. In at least one embodiment, resource manager 1926 may be able to manage clustered or grouped computing resources mapped to or allocated to support distributed file system 1928 and job scheduler 1922. In at least one embodiment, clustered or grouped computing resources may include grouped computing resources 1914 at data center infrastructure layer 1910. In at least one embodiment, resource manager 1926 may coordinate with resource coordinator 1912 to manage these mapped or allocated computing resources.
[0224] In at least one embodiment, the software 1932 included in software layer 1930 may include software used by at least portions of nodes CR1916(1)-1916(N), grouped computing resources 1914, and / or the distributed file system 1928 of framework layer 1920. In at least one embodiment, one or more types of software may include, but are not limited to, Internet web page search software, email virus scanning software, database software, and streaming video content software.
[0225] In at least one embodiment, one or more applications 1942 included in application layer 1940 may include one or more types of applications used by at least portions of nodes CR1916(1)-1916(N), grouped computing resources 1914, and / or the distributed file system 1928 of framework layer 1920. In at least one embodiment, one or more types of applications may include, but are not limited to, any number of genomics applications, cognitive computing, applications, and machine learning applications, including training or inference software, machine learning framework software (e.g., PyTorch, TensorFlow, Caffe, etc.), or other machine learning applications used in conjunction with one or more embodiments.
[0226] In at least one embodiment, any of the configuration manager 1924, resource manager 1926, and resource coordinator 1912 can implement any number and type of self-modification actions based on any amount and type of data acquired in any technically feasible manner. In at least one embodiment, self-modification actions can mitigate potentially poor configuration decisions by data center operators of data center 1900 and can prevent underutilization and / or poor performance of the data center.
[0227] In at least one embodiment, data center 1900 may include tools, services, software, or other resources for training one or more machine learning models or using one or more machine learning models to predict or infer information according to one or more embodiments described herein. For example, in at least one embodiment, a machine learning model can be trained by calculating weight parameters based on a neural network architecture using the software and computing resources described above with respect to data center 1900. In at least one embodiment, information can be inferred or predicted using trained machine learning models corresponding to one or more neural networks using the resources described above with respect to data center 1900 by using weight parameters calculated through one or more training techniques described herein.
[0228] In at least one embodiment, the data center may use a CPU, application-specific integrated circuit (ASIC), GPU, FPGA, or other hardware to utilize the aforementioned resources to perform training and / or inference. Furthermore, one or more of the aforementioned software and / or hardware resources may be configured as a service to allow a user to train or perform information inference, such as image recognition, speech recognition, or other artificial intelligence services.
[0229] The logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 18A and / or Figure 18B Details regarding logic 1815 are provided. In at least one embodiment, logic 1815 may be used in data center 1900 for inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0230] In at least one embodiment, Figure 19 The system described herein is used to employ various algorithms, formulas, and processes (e.g., combining...) Figure 1 The described methods encode and / or classify one or more logs and / or otherwise perform the operations described herein. In at least one embodiment, Figure 19 The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The ones described herein are used to encode and / or classify one or more logs and / or otherwise perform the operations described herein. As an example, Figure 19 The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing one or more of the operations described herein.
[0231] As an example, Figure 19 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message using at least one neural network, said at least one neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. As an example, Figure 19 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least in part by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or to otherwise perform the operations described herein. As an example, Figure 19 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message by at least partially: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code by at least partially combining the first code and the second code; and / or otherwise performing the operations described herein.
[0232] Computer System
[0233] Figure 20This is a block diagram illustrating an exemplary computer system according to at least one embodiment. The exemplary computer system may be a system of interconnected devices and components, a system-on-a-chip (SoC), or some combination thereof formed with a processor, which may include an execution unit for executing instructions. In at least one embodiment, according to this disclosure, such as in the embodiments described herein, computer system 2000 may include, but is not limited to, components such as processor 2002 for employing execution units (including logic) to execute algorithms for process data. In at least one embodiment, computer system 2000 may include a processor, such as those available from Intel Corporation of Santa Clara, California. Processor family, Xeon TM , XScale TM and / or StrongARM TM , Core TM or Nervana TM The microprocessor can be used, although other systems (including PCs, engineering workstations, set-top boxes, etc.) with other microprocessors can also be used. In at least one embodiment, the computer system 2000 can execute a version of the Windows operating system available from Microsoft Corporation of Redmond, Washington, although other operating systems (such as UNIX and Linux), embedded software, and / or graphical user interfaces can also be used.
[0234] The embodiments can be used in other devices, such as handheld devices and embedded applications. Some examples of handheld devices include cellular phones, Internet Protocol (IP) devices, digital cameras, personal digital assistants (“PDAs”), and handheld PCs. In at least one embodiment, the embedded application may include a microcontroller, a digital signal processor (“DSP”), a system-on-a-chip (SoC), a network computer (“NetPC”), a set-top box, a network hub, a wide area network (“WAN”) switch, or any other system capable of executing one or more instructions according to at least one embodiment.
[0235] In at least one embodiment, the computer system 2000 may include, but is not limited to, a processor 2002, which may include, but is not limited to, one or more execution units 2008 for performing machine learning model training and / or inference according to the techniques described herein. In at least one embodiment, the computer system 2000 is a single-processor desktop or server system, but in another embodiment, the computer system 2000 may be a multiprocessor system. In at least one embodiment, the processor 2002 may include, but is not limited to, for example, a Complex Instruction Set Computer (“CISC”) microprocessor, a Reduced Instruction Set Computing (“RISC”) microprocessor, a Very Long Instruction Word (“VLIW”) microprocessor, a processor implementing instruction set combination, or any other processor device, such as a digital signal processor. In at least one embodiment, the processor 2002 may be coupled to a processor bus 2010, which can transmit data signals between the processor 2002 and other components in the computer system 2000.
[0236] In at least one embodiment, processor 2002 may include, but is not limited to, a Level 1 (“L1”) internal cache memory (“cache”) 2004. In at least one embodiment, processor 2002 may have a single internal cache or multiple levels of internal caches. In at least one embodiment, the cache memory may reside external to processor 2002. Depending on specific implementation and requirements, other embodiments may also include a combination of internal and external caches. In at least one embodiment, register file 2006 may store different types of data in various registers, including but not limited to integer registers, floating-point registers, status registers, and instruction pointer registers.
[0237] In at least one embodiment, an execution unit 2008, including but not limited to logic for performing integer and floating-point operations, is also located in the processor 2002. In at least one embodiment, the processor 2002 may also include a microcode (“ucode”) read-only memory (“ROM”) storing the microcode of certain macro instructions. In at least one embodiment, the execution unit 2008 may include logic for processing a packaged instruction set 2009. In at least one embodiment, by including the packaged instruction set 2009 in the instruction set of the general-purpose processor and the associated circuitry to be executed, operations used by numerous multimedia applications can be performed using packaged data in the processor 2002. In at least one embodiment, numerous multimedia applications can be accelerated and executed more efficiently by performing operations on packaged data using the full width of the processor's data bus, eliminating the need to transfer smaller data units on the processor's data bus to perform one or more operations on one data element at a time.
[0238] In at least one embodiment, the execution unit 2008 may also be used in a microcontroller, embedded processor, graphics device, DSP, and other types of logic circuitry. In at least one embodiment, the computer system 2000 may include, but is not limited to, a memory 2020. In at least one embodiment, the memory 2020 may be a dynamic random access memory (“DRAM”) device, a static random access memory (“SRAM”) device, a flash memory device, or other memory device. In at least one embodiment, the memory 2020 may store one or more instructions 2019 and / or data 2021 represented by data signals executable by the processor 2002.
[0239] In at least one embodiment, the system logic chip may be coupled to processor bus 2010 and memory 2020. In at least one embodiment, the system logic chip may include, but is not limited to, a memory controller hub (“MCH”) 2016, and processor 2002 may communicate with MCH 2016 via processor bus 2010. In at least one embodiment, MCH 2016 may provide a high-bandwidth memory path 2018 to memory 2020 for instruction and data storage, as well as for storage of graphics commands, data, and textures. In at least one embodiment, MCH 2016 may direct data signals between processor 2002, memory 2020, and other components in computer system 2000, and bridge data signals between processor bus 2010, memory 2020, and system I / O interface 2022. In at least one embodiment, the system logic chip may provide a graphics port for coupling to a graphics controller. In at least one embodiment, MCH 2016 may be coupled to memory 2020 via high-bandwidth memory path 2018, and graphics / video card 2012 may be coupled to MCH 2016 via accelerated graphics port (“AGP”) interconnect 2014.
[0240] In at least one embodiment, the computer system 2000 may use the system I / O interface 2022 as a proprietary hub interface bus to couple the MCH 2016 to the I / O controller hub (“ICH”) 2030. In at least one embodiment, the ICH 2030 may provide direct connectivity to certain I / O devices via a local I / O bus. In at least one embodiment, the local I / O bus may include, but is not limited to, a high-speed I / O bus for connecting peripheral devices to the memory 2020, chipset, and processor 2002. Examples may include, but are not limited to, an audio controller 2029, a firmware hub (“Flash BIOS”) 2028, a wireless transceiver 2026, a data storage 2024, a conventional I / O controller 2023 including a user input and keyboard interface 2025, a serial expansion port 2027 (such as a Universal Serial Bus (“USB”) port), and a network controller 2034. In at least one embodiment, the data storage 2024 may include a hard disk drive, a floppy disk drive, a CD-ROM device, a flash memory device, or other mass storage device.
[0241] In at least one embodiment, Figure 20 A system including interconnected hardware devices or "chips" is shown, while in other embodiments, Figure 20 An exemplary SoC can be shown. In at least one embodiment, Figure 20 The devices shown can be interconnected using proprietary interconnects, standardized interconnects (e.g., PCIe), or some combination thereof. In at least one embodiment, one or more components of the computer system 2000 are interconnected using a Computational Fast Link (CXL) interconnect.
[0242] The logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. This document combines... Figure 18A and / or Figure 18B Details regarding logic 1815 are provided. In at least one embodiment, logic 1815 can be used in a computer system for performing inference or prediction operations based at least in part on weight parameters computed using neural network training operations, neural network functions and / or architectures, or neural network use cases described herein.
[0243] In at least one embodiment, Figure 20 The system described herein is used to employ various algorithms, formulas, and processes (e.g., combining...) Figure 1 The described methods encode and / or classify one or more logs and / or otherwise perform the operations described herein. In at least one embodiment, Figure 20 The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 16The ones described herein are used to encode and / or classify one or more logs and / or otherwise perform the operations described herein. As an example, Figure 20 The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 16 The methods described herein are used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing one or more of the operations described herein. As an example, Figure 20 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 16 The methods described herein are used to encode at least one log message using at least one neural network, said at least one neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the similarity loss; and / or otherwise performing the operations described herein. As an example, Figure 20 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 16 The methods described herein are used to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least in part by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or to otherwise perform the operations described herein. As an example, Figure 20 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 16 The methods described herein are used to encode at least one log message by at least partially: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code by at least partially combining the first code and the second code; and / or otherwise performing the operations described herein.
[0244] In at least one embodiment, Figure 20 The system described herein is used to employ various algorithms, formulas, and processes (e.g., combining...) Figure 1 The described methods encode and / or classify one or more logs and / or otherwise perform the operations described herein. In at least one embodiment, Figure 20 The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The ones described herein are used to encode and / or classify one or more logs and / or otherwise perform the operations described herein. As an example, Figure 20 The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing one or more of the operations described herein.
[0245] As an example, Figure 20 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message using at least one neural network, said at least one neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. As an example, Figure 20 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23The methods described herein are used to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least in part by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or to otherwise perform the operations described herein. As an example, Figure 20 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message by at least partially: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code by at least partially combining the first code and the second code; and / or otherwise performing the operations described herein.
[0246] Neural network training and deployment
[0247] Figure 21 Training and deployment of a deep neural network according to at least one embodiment are illustrated. In at least one embodiment, an untrained neural network 2106 is trained using a training dataset 2102. In at least one embodiment, the training framework 2104 is the PyTorch framework, while in other embodiments, the training framework 2104 is TensorFlow, Boost, Caffe, Microsoft Cognitive Toolkit / CNTK, MXNet, Chainer, Keras, Deeplearning4j, or other training frameworks. In at least one embodiment, the training framework 2104 trains the untrained neural network 2106 and enables it to be trained using the processing resources described herein to generate a trained neural network 2108. In at least one embodiment, weights may be randomly selected or selected by pre-training using a deep belief network. In at least one embodiment, training may be performed in a supervised, partially supervised, or unsupervised manner.
[0248] In at least one embodiment, the untrained neural network 2106 is trained using supervised learning, wherein the training dataset 2102 includes inputs paired with desired outputs of the inputs, or wherein the training dataset 2102 includes inputs with known outputs, and the outputs of the neural network 2106 are manually graded. In at least one embodiment, the untrained neural network 2106 is trained in a supervised manner and processes inputs from the training dataset 2102, comparing the resulting outputs with a set of expected or desired outputs. In at least one embodiment, errors are then backpropagated through the untrained neural network 2106. In at least one embodiment, a training framework 2104 adjusts the weights controlling the untrained neural network 2106. In at least one embodiment, the training framework 2104 includes tools for monitoring how well the untrained neural network 2106 converges to a model (e.g., a trained neural network 2108) suitable for generating correct answers (e.g., in result 2114) based on input data (e.g., a new dataset 2112). In at least one embodiment, the training framework 2104 repeatedly trains the untrained neural network 2106 while adjusting the weights to refine the output of the untrained neural network 2106 using a loss function and a tuning algorithm (e.g., stochastic gradient descent). In at least one embodiment, the training framework 2104 trains the untrained neural network 2106 until the untrained neural network 2106 reaches the desired accuracy. In at least one embodiment, the trained neural network 2108 can then be deployed to perform any number of machine learning operations.
[0249] In at least one embodiment, the untrained neural network 2106 is trained using unsupervised learning, wherein the untrained neural network 2106 attempts to self-train using unlabeled data. In at least one embodiment, the unsupervised learning training dataset 2102 will include input data without any associated output data or "ground truth" data. In at least one embodiment, the untrained neural network 2106 can learn groupings within the training dataset 2102 and can determine the relationship between each input and the untrained dataset 2102. In at least one embodiment, unsupervised training can be used to generate a self-organizing map in the trained neural network 2108, which is capable of performing operations that help reduce the dimensionality of the new dataset 2112. In at least one embodiment, unsupervised training can also be used to perform anomaly detection, which allows the identification of data points in the new dataset 2112 that deviate from the normal patterns of the new dataset 2112.
[0250] In at least one embodiment, semi-supervised learning can be used, which is a technique that includes a mixture of labeled and unlabeled data in the training dataset 2102. In at least one embodiment, the training framework 2104 can be used to perform incremental learning, such as through transfer learning techniques. In at least one embodiment, incremental learning enables the trained neural network 2108 to adapt to a new dataset 2112 without forgetting the knowledge injected into the trained neural network 2108 during initial training.
[0251] In at least one embodiment, the training framework 2104 is a framework that is processed in conjunction with a software development kit (e.g., the OpenVINO (Open Visual Inference and Neural Network Optimization) toolkit). In at least one embodiment, the OpenVINO toolkit is a toolkit such as those developed by Intel Corporation of Santa Clara, California. In at least one embodiment, OpenVINO includes or uses logic 1815 to perform the operations described herein. In at least one embodiment, a SoC, integrated circuit, or processor uses OpenVINO to perform the operations described herein.
[0252] In at least one embodiment, OpenVINO is a toolkit for facilitating the development of applications (particularly neural network applications) for various tasks and operations, such as human visual simulation, speech recognition, natural language processing, recommender systems, and / or variations thereof. In at least one embodiment, OpenVINO supports neural networks, such as convolutional neural networks (CNNs), recurrent neural networks, and / or attention-based neural networks, and / or various other neural network models. In at least one embodiment, OpenVINO supports various software libraries, such as OpenCV, OpenCL, and / or variations thereof.
[0253] In at least one embodiment, OpenVINO supports neural network models for a variety of tasks and operations, such as classification, segmentation, object detection, face recognition, speech recognition, pose estimation (e.g., people and / or objects), monocular depth estimation, image inpainting, style transfer, action recognition, colorization, and / or variations thereof.
[0254] In at least one embodiment, OpenVINO includes one or more software tools and / or modules for model optimization, also referred to as a model optimizer. In at least one embodiment, the model optimizer is a command-line tool that facilitates the transition between training and deployment of a neural network model. In at least one embodiment, the model optimizer optimizes the neural network model for execution on various devices and / or processing units such as GPUs, CPUs, PPUs, GPGPUs, and / or variants thereof. In at least one embodiment, the model optimizer generates an internal representation of the model and optimizes the model to generate an intermediate representation. In at least one embodiment, the model optimizer reduces the number of layers in the model. In at least one embodiment, the model optimizer removes layers from the model used for training. In at least one embodiment, the model optimizer performs various neural network operations, such as modifying the model's input (e.g., adjusting the size of the model's input), modifying the size of the model's input (e.g., modifying the model's batch size), modifying the model's structure (e.g., modifying the model's layers), normalizing, standardizing, quantizing (e.g., converting the model's weights from a first representation such as floating-point to a second representation such as integer), and / or variants thereof.
[0255] In at least one embodiment, OpenVINO includes one or more software libraries for inference, also referred to as an inference engine. In at least one embodiment, the inference engine is a C++ library or any suitable programming language library. In at least one embodiment, the inference engine is used to infer input data. In at least one embodiment, the inference engine implements various classes to infer input data and generate one or more results. In at least one embodiment, the inference engine implements one or more API functions to process intermediate representations, set input and / or output formats, and / or execute models on one or more devices.
[0256] In at least one embodiment, OpenVINO provides various capabilities for heterogeneous execution of one or more neural network models. In at least one embodiment, heterogeneous execution or heterogeneous computing refers to one or more computational processes and / or systems utilizing one or more types of processors and / or cores. In at least one embodiment, OpenVINO provides various software functions to execute programs on one or more devices. In at least one embodiment, OpenVINO provides various software functions to execute programs and / or portions of programs on different devices. In at least one embodiment, OpenVINO provides various software functions, for example, to run a first code portion on a CPU and a second code portion on a GPU and / or FPGA. In at least one embodiment, OpenVINO provides various software functions to execute one or more layers of a neural network on one or more devices (e.g., executing a first set of layers on a first device (e.g., a GPU) and a second set of layers on a second device (e.g., a CPU)).
[0257] In at least one embodiment, OpenVINO includes various functionalities similar to those associated with CUDA programming models, such as various neural network model operations associated with frameworks such as TensorFlow, PyTorch, and / or their variants. In at least one embodiment, one or more CUDA programming model operations are performed using OpenVINO. In at least one embodiment, the various systems, methods, and / or techniques described herein are implemented using OpenVINO.
[0258] In at least one embodiment, Figure 21 The system described herein is used to employ various algorithms, formulas, and processes (e.g., combining...) Figure 1 The described methods encode and / or classify one or more logs and / or otherwise perform the operations described herein. In at least one embodiment, Figure 21 The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The ones described herein are used to encode and / or classify one or more logs and / or otherwise perform the operations described herein. As an example, Figure 21 The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23The methods described herein are used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing one or more of the operations described herein.
[0259] As an example, Figure 21 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message using at least one neural network, said at least one neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. As an example, Figure 21 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least in part by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or to otherwise perform the operations described herein. As an example, Figure 21 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message by at least partially: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code by at least partially combining the first code and the second code; and / or otherwise performing the operations described herein.
[0260] Figure 22 This is a system diagram illustrating a system 2200 for interfacing with an application 2202 to process data, according to at least one embodiment. In at least one embodiment, the application 2202 uses a Large Language Model (LLM) 2212 to generate output data 2220 based at least in part on input data 2210. In at least one embodiment, the input data 2210 is a text prompt. In at least one embodiment, the input data 2210 includes unstructured text. In at least one embodiment, the input data 2210 includes a series of tokens. In at least one embodiment, the tokens are part of the input data. In at least one embodiment, the tokens are words. In at least one embodiment, the tokens are characters. In at least one embodiment, the tokens are subwords. In at least one embodiment, the input data 2210 is formatted in Chat Markup Language (ChatML). In at least one embodiment, the input data 2210 is an image. In at least one embodiment, the input data 2210 is one or more video frames. In at least one embodiment, the input data 2210 is any other expressive medium.
[0261] In at least one embodiment, the large language model 2212 includes a deep neural network. In at least one embodiment, the deep neural network is a neural network having two or more layers. In at least one embodiment, the large language model 2212 includes a converter model. In at least one embodiment, the large language model 2212 includes a neural network configured to perform natural language processing. In at least one embodiment, the large language model 2212 is configured to process one or more data sequences. In at least one embodiment, the large language model 2212 is configured to process text. In at least one embodiment, the weights and biases of the large language model 2212 are configured to process text. In at least one embodiment, the large language model 2212 is configured to determine patterns in data to perform one or more natural language processing tasks. In at least one embodiment, the natural language processing task includes text generation. In at least one embodiment, the natural language processing task includes question answering. In at least one embodiment, performing the natural language processing task produces output data 2220.
[0262] In at least one embodiment, the processor uses input data 2210 to query retrieval database 2214. In at least one embodiment, retrieval database 2214 is a key-value store. In at least one embodiment, retrieval database 2214 is a corpus used to train large language model 2212. In at least one embodiment, the processor uses retrieval database 2214 to provide updated information to large language model 2212. In at least one embodiment, retrieval database 2214 includes data from Internet sources. In at least one embodiment, large language model 2212 does not use retrieval database 2214 to perform inference.
[0263] In at least one embodiment, the encoder encodes input data 2210 into one or more feature vectors. In at least one embodiment, the encoder encodes input data 2210 into sentence embedding vectors. In at least one embodiment, the processor uses the sentence embedding vectors to perform nearest neighbor search to generate one or more neighbors 2216. In at least one embodiment, one or more neighbors 2216 are values retrieved from database 2214 corresponding to a key including input data 2210. In at least one embodiment, one or more neighbors 2216 include text data. In at least one embodiment, the encoder 2218 encodes one or more neighbors 2216. In at least one embodiment, the encoder 2218 encodes one or more neighbors 2216 into text embedding vectors. In at least one embodiment, the encoder 2218 encodes one or more neighbors 2216 into sentence embedding vectors. In at least one embodiment, the large language model 2212 uses input data 2210 and data generated by encoder 2218 to generate output data 2220. In at least one embodiment, processor 2206 uses Large Language Model (LLM) application programming interface (API) 2204 to interact with application 2202. In at least one embodiment, processor 2206 uses Large Language Model (LLM) application programming interface (API) 2204 to access Large Language Model 2212.
[0264] In at least one embodiment, output data 2220 includes computer instructions. In at least one embodiment, output data 2220 includes instructions written in the CUDA programming language. In at least one embodiment, output data 2220 includes instructions to be executed by processor 2206. In at least one embodiment, output data 2220 includes instructions for controlling the execution of one or more algorithm modules 2208. In at least one embodiment, one or more algorithm modules 2208 include, for example, one or more neural networks for performing pattern recognition. In at least one embodiment, one or more algorithm modules 2208 include, for example, one or more neural networks for performing frame generation. In at least one embodiment, one or more algorithm modules 2208 include, for example, one or more neural networks for generating driving paths. In at least one embodiment, one or more algorithm modules 2208 include, for example, one or more neural networks for generating 5G signals. In at least one embodiment, processor 2206 interacts with application 2202 using Large Language Model (LLM) application programming interface (API) 2204. In at least one embodiment, the processor 2206 may use one or more parallel computing platforms and / or programming models (e.g., NVIDIA's CUDA model).
[0265] In at least one embodiment, the text herein combines Figure 22 The aspects of the described system and technology are incorporated into the aspects of the preceding figures. For example, in at least one embodiment, the apparatus described in the preceding figures includes a processor 2206.
[0266] For example, in at least one embodiment, system 2200 uses ChatGPT to write CUDA code. For example, in at least one embodiment, system 2200 uses ChatGPT to train an object classification neural network. For example, in at least one embodiment, system 2200 uses ChatGPT and a neural network to identify driving paths. For example, in at least one embodiment, system 2200 uses ChatGPT and a neural network to generate 5G signals.
[0267] Logic 1815 is used to perform inference and / or training operations associated with one or more embodiments. Details about Logic 1815 are combined in this document. Figure 18A and / or Figure 18B Provided. In at least one embodiment, logic 1815 may be used in system 2200 to infer or predict operations based at least in part on weight parameters, neural network functions and / or architectures computed using neural network training operations or neural network use cases described herein.
[0268] In at least one embodiment, Figure 22 The system described herein is used to employ various algorithms, formulas, and processes (e.g., combining...) Figure 1 The described methods encode and / or classify one or more logs and / or otherwise perform the operations described herein. In at least one embodiment, Figure 22 The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The ones described herein are used to encode and / or classify one or more logs and / or otherwise perform the operations described herein. As an example, Figure 22 The one or more systems described herein are used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23The methods described herein are used to encode at least one vector associated with at least one log sequence using a neural network, said neural network being trained at least in part by: obtaining a first encoded vector, a second encoded vector, and a third encoded vector by encoding a first vector associated with a first log sequence, a second vector associated with a second log sequence similar to the first log sequence, and a third vector associated with a third log sequence dissimilar to the first log sequence; selecting at least one model weight that increases the probability that the first encoded vector is closer to the second encoded vector than the third encoded vector; and / or otherwise performing one or more of the operations described herein.
[0269] As an example, Figure 22 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message using at least one neural network, said at least one neural network being trained at least in part by: obtaining similarity scores associated with a first vector and a second vector, the first vector being associated with one or more first log messages and the second vector being associated with one or more second log messages; generating at least one similarity value indicating a similarity loss between the first vector and the second vector; determining an index indicating the similarity between the similarity scores and the at least one similarity loss value; and / or otherwise performing the operations described herein. As an example, Figure 22 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to classify one or more log entries to obtain one or more classified log entries; to obtain combined information at least in part by combining at least one or more classified log entries and telemetry information; to classify the combined information using at least one machine learning process; and / or to otherwise perform the operations described herein. As an example, Figure 22 The description of one or more systems is used to implement one or more systems and / or processes (e.g., in combination). Figures 1 to 17B and / or Figure 23 The methods described herein are used to encode at least one log message by at least partially: encoding information of a first type in at least one log message to obtain a first code; encoding information of a second type in at least one log message to obtain a second code; obtaining a result code by at least partially combining the first code and the second code; and / or otherwise performing the operations described herein.
[0270] Figure 23 This is a flowchart illustrating a process 2300 of training a second neural network, at least in part, based on a first neural network to encode a log sequence, according to at least one embodiment. Processor 110 may execute process 2300 to train a converter encoder (e.g., one or more neural networks NN1, one or more neural networks NN2, and / or one or more classifiers 122). For example, encoder function 115 and / or classification function 116 may execute process 2300. In at least one embodiment, process 2300 is performed by system 600 (see...). Figure 6 ) is executed, at least in part, based on a neural network trained using similarity loss (e.g., encoder 1408, see Figure 14 (For example, a first (teacher) neural network) to train a second (student) neural network (e.g., neural network 608).
[0271] In at least one embodiment, the processor executing process 2300 is at least partially based on one or more first neural networks (e.g., a trainer and / or encoder 1408, see below). Figure 14 To train one or more second (e.g., student) neural networks (e.g., machine learning model 604, see...) Figure 6 This can be used to perform anomaly detection, for example. One or more second neural networks can be called one or more student neural networks; as another example, one or more first neural networks can be called one or more teacher neural networks. (Reference) Figure 23 The processor executing process 2300 may, at box 2302, cause one or more first (e.g., instructor) neural networks to receive or obtain a first training dataset comprising one or more first training sets including one or more log sequences associated with one or more labels; at box 2304, cause the first (e.g., instructor) neural networks to generate a second training dataset comprising one or more second training sets by generating one or more similarity scores for one or more log sequence pairs in the first training set; at box 2305, cause the similarity scores to be augmented according to one or more labels; and at box 2306, cause one or more second (e.g., student) neural networks to execute process 1600 using the second training set (see...). Figure 16 To select a model configuration, and if appropriate, output adjustments to one or more model configurations (e.g., weights) of a second (e.g., student) neural network in box 2308, and / or perform one or more operations described herein, or a combination thereof.
[0272] In at least one embodiment, firstly, one or more processors (e.g., processor 602, see...) Figure 6The process 2300 calls procedure 2300 and / or receives or obtains one or more log sequences and associated labels as input to a first training set (e.g., a dataset). Procedure 2300 includes receiving one or more log sequences and associated labels as input in box 2302, which may include an anomaly classification label (e.g., 1 or 0) for each log sequence. The anomaly classification label may include a value indicating whether the log sequence includes an anomaly, such as a value of "1" if the log sequence includes an anomaly, and a value of "0" otherwise. The first training dataset received or obtained as input in box 2302 may include one or more log sequences, one or more log line pairs 1504A (see...). Figure 15 The processor of execution process 2300 may receive one or more log sequences (e.g., generated or otherwise ordered as log sequence pairs) and / or one or more associated (e.g., corresponding) labels, such as one or more truth labels (e.g., exception indicators) as input in block 2302.
[0273] In at least one embodiment, in box 2302, a first (e.g., instructor) neural network receives one or more log sequences (e.g., to generate a second training set of similarity scores associated with one or more log sequence pairs, in box 2304) and associated labels, such as ground truth labels. One or more first (e.g., instructor) neural networks may include one or more encoders 1408 trained using similarity loss (see...). Figure 14 As an example, in box 2304, one or more first (e.g., instructor) neural networks may generate one or more second training sets to include one or more similarity scores for one or more log sequence pairs from the first training sets. As an example, each similarity score generated in box 2304 may be a similarity label, such as a value indicating the similarity between the log sequences in that pair.
[0274] The first training set may include information identifying training pairs within the first training set and / or two or more log sequences received as input in box 2302 may be formed or arranged as training pairs in box 2304. The processor executing process 2300 may perform a selection process that selects one or more training pairs. In at least one embodiment, one or more first (e.g., instructor) neural networks generate similarity labels (e.g., similarity scores) for each identified log sequence pair. As an example, for each pair, the first (e.g., instructor) neural network may compute a similarity value, such as cosine similarity. The processor executing process 2300 may then enhance (e.g., adjust) the similarity value in box 2305 to obtain a similarity score using the hyperparameter alpha and a ground truth anomaly label (e.g., "1" indicates an anomaly, otherwise "0"). The processor executing process 2300 and / or the first (e.g., instructor) neural network may enhance one or more similarity scores based on one or more labels in box 2305.
[0275] For example, a first (e.g., instructor) neural network may include a language model (e.g., LLM), and each sequence may be transformed into raw text (e.g., using the preprocessing described herein) and passed through the language model to produce encoding pairs. The language model may be pre-trained to determine sentence and / or paragraph similarity. The processor executing process 2300 may then use the encoding pairs to generate an initial similarity score, which the processor may adjust using alpha. As an example, if two log sequences are associated with the same classification (e.g., truth labels indicating anomaly classifications are both "1" or both are "0"), the processor may use the initial similarity score as a similarity score in a second training set (e.g., as a label) if the initial similarity score is greater than alpha; otherwise, the processor may set the similarity score in the second training set to be equal to alpha. If two log sequences are associated with different classifications indicated by truth labels (e.g., one is anomalous and the other is not), then if the initial similarity score is less than one minus alpha (e.g., initial similarity score < (1-alpha)), the processor executing process 2300 can use the initial similarity score as the similarity score in the second training set (e.g., as a label); otherwise, the processor can set the similarity score in the second training set to be equal to alpha. In an exemplary implementation, alpha is selected at least in part based on ablation, for example, alpha = 0.7.
[0276] In at least one embodiment, process 2300 may include converting one or more labels of one or more log sequences into similarity scores (e.g., a score of 1 if the label is "1" and a score of -1 if the label is "0"). As an example, converting one or more labels of one or more log sequences into similarity scores may be used in conjunction with or in lieu of boxes 2302, 2304, and 2305. For example, instead of executing boxes 2302-2305, the processor executing process 2300 may receive a first training dataset comprising one or more log sequences associated with one or more labels. The processor executing process 2300 may then select one or more log sequence pairs as described herein and generate a second training dataset comprising one or more log sequence pairs, each associated with a similarity score. For each pair, the similarity score can be generated by mapping each label to a similarity score. For example, a processor can assign a similarity score of "1" (e.g., assign label "1") to a log sequence associated with a label indicating the presence of an anomaly, and can assign a similarity score of "-1" (e.g., assign label "0") to a log sequence associated with a label indicating the absence of an anomaly. The similarity scores of each pair can then be combined (e.g., averaged). For example, if both log sequences contain an anomaly ((1+1) / 2 = 1), the sum of the two log sequences is "1"; if only one log sequence contains an anomaly ((1+-1) / 2 = 0), the sum of the two log sequences is zero; and if neither log sequence contains an anomaly ((-1+-1) / 2 = -1), the sum of the two log sequences is "-1". Therefore, this mapping can be used to convert labels into cosine similarity scores.
[0277] In box 2306, the second (e.g., student) neural network may use the second training set to perform process 1600 to select one or more model configurations. As an example, selecting a model configuration in box 2306 may include: using the second (e.g., student) neural network to infer the encoding of each pair of log sequences in the second training set, determining a similarity value (e.g., cosine similarity) between the encodings inferred for each pair, using a loss function (e.g., mean squared error) to calculate a loss value between the similarity value and the similarity score (in the second training set) for each pair, aggregating the calculated loss for each pair, and selecting a model configuration (e.g., model weights) that reduces or minimizes the total loss. For example, the processor of execution process 2300 may use the forward pass of a student model (e.g., a second neural network) to compute cosine similarity between corresponding encodings of log sequence pairs, and compare this cosine similarity with a cosine similarity label (or similarity score) generated by a teacher model (e.g., a first neural network) based at least in part on the mean squared error (MSE) loss between the cosine similarity determined by the student model and the cosine similarity labels determined by the teacher model (and included in the second training set). The processor of execution process 2300 may aggregate (e.g., calculate the total, average, etc.) the MSE losses computed for each pair in the second training set to obtain the total MSE loss for the current configuration of the second (student) neural network. The processor of execution process 2300 may cause the first (student) neural network to process the second training set multiple times using different model configurations (e.g., different weights). The processor of execution process 2300 may then select the model configuration that produces the minimum total MSE loss for the log sequence pairs in the second training set.
[0278] As an example, the second (e.g., student) neural network may receive one or more second training sets (e.g., log sequence pairs and one or more similarity scores) as input, such as one or more similarity scores generated by the first neural network. In at least one embodiment, once the processor executing process 2300 in block 2306 executes process 1600 to select a model configuration using the similarity scores generated by the first neural network in block 2304, the processor may output an adjustment to the model configuration (e.g., weights) of the second (student) neural network in block 2308. For example, if the model configuration determined by process 1600 differs from the current model configuration of the first (student) neural network, the processor may determine and output the adjustment to the model configuration in block 2308. In other words, outputting this information is appropriate. On the other hand, if the model configuration determined by process 1600 is not different from the current model configuration of the first (student) neural network, block 2308 may be omitted. In at least one embodiment, in block 2308, the processor may use the model configuration selected in block 2306 and / or the adjustments to the output in block 2308 to backpropagate updates to one or more model weights of a second (e.g., student) neural network. Following block 2308, the processor executing process 2300 may perform one or more operations described herein, and / or terminate.
[0279] As an example, the second (e.g., student) neural network may include one or more neural networks 608 (see...). Figure 6 At least in part, a second (e.g., student) neural network trained using process 2300 can be used to perform one or more inference operations. As an example, given a query sequence (e.g., a log sequence), the second (e.g., student) neural network trained using process 2300 can generate one or more codes (e.g., vector codes). A processor executing the second (e.g., student) neural network trained using process 2300 can compute a similarity value (e.g., cosine similarity "s") about the mean codes of the training sequence (e.g., determined using codes generated by the first (teacher) neural network, codes generated by the second (student) neural network, and / or truth labels).
[0280] For example, mean encoding can include a vector v μ The vector is calculated by first encoding one or more “normal” training sequences using a trained model (e.g., a first (teacher) neural network and / or a second (student) neural network) and then calculating the mean vector (e.g., summing all vectors into a single vector and dividing each element by the number of vectors). In at least one embodiment, the mean encoding can be calculated using the following equation: As an example, in this equation, v i,μ It can be the mean vector vμ The i-th element, v i,j It is the encoded vector v j The i-th element (which is the encoding of the j-th normal sequence in the training set computed by the trained model), and N is the total number of normal sequences in the generated training dataset (e.g., the second training dataset).
[0281] For example, a processor executing a second (student) neural network can classify a sequence as an anomalous when (1-RELU) > alpha and as non-anomaly when (1-RELU) ≤ alpha, where RELU refers to a rectified linear unit. In at least one embodiment, the second (e.g., student) neural network, trained at least in part using process 2300, can learn even with a small number of anomalous units (e.g., 100) in the training set, e.g., without pattern collapse and achieving the desired performance (e.g., measured by F1 score). Furthermore, the second (e.g., student) neural network may not assume a fixed vocabulary (e.g., although embodiments may include the use of a vocabulary) or rely on template extraction, which could introduce errors and / or affect model performance.
[0282] In at least one embodiment, part or all of process 2300 (or any other process described herein, or variations and / or combinations thereof) is executed under the control of one or more computer systems configured with computer-executable instructions and is i...
Claims
1. A method comprising: At least one log message is encoded, in part, through the following operations: Encode the information of a first type in the at least one log message to obtain a first code; Encode the second type of information in the at least one log message to obtain a second encoding; as well as The resulting encoding is obtained at least partially by combining the first encoding and the second encoding.
2. The method according to claim 1, wherein the first type of information and the second type of information include character information and category information.
3. The method according to claim 2, wherein the character information includes at least one of text information or numerical information.
4. The method of claim 2, wherein the category information includes a priority associated with the at least one log message.
5. The method of claim 1, wherein the result encoding includes vector encoding.
6. The method according to claim 1, further comprising: The information of a third type in the at least one log message is encoded to obtain a third encoding, the resulting encoding being obtained at least in part by combining the first encoding, the second encoding, and the third encoding.
7. The method of claim 6, wherein the attention layer is used to at least combine the first encoding, the second encoding, and the third encoding.
8. The method of claim 6, wherein at least one neural network is used to encode at least one of the first type of information, the second type of information, or the third type of information.
9. The method according to claim 1, wherein the first type of information and the second type of information are respectively text information and category information. A first neural network, including a text encoder, is used to encode the text information, and A second neural network, including a category encoder, is used to encode the category information.
10. The method according to claim 1, further comprising: The resulting encoding is used to perform anomaly detection.
11. A processor, comprising: One or more circuits, said one or more circuits being used to encode at least one log message by at least part of the following operations: Encode the information of a first type in the at least one log message to obtain a first code; Encode the second type of information in the at least one log message to obtain a second encoding; as well as The resulting encoding is obtained at least partially by combining the first encoding and the second encoding.
12. The processor of claim 11, wherein the first type of information includes text information, the second type of information includes category information, the third type of information includes numerical information, and the one or more circuits are configured to encode the at least one log message at least in part by: Encode the third type of information in the at least one log message to obtain a third encoding; and The resulting encoding is obtained at least in part by combining at least the first encoding, the second encoding, and the third encoding.
13. The processor of claim 12, wherein one or more circuits are configured to use an attention layer to at least combine the first encoding, the second encoding, and the third encoding.
14. The processor of claim 12, wherein the one or more circuits are configured to use at least one neural network to encode at least one of the first type of information, the second type of information, or the third type of information.
15. The processor of claim 11, wherein the result encoding includes vector encoding.
16. A system comprising: One or more processors, said one or more processors being used to encode at least one log message at least in part by: Encode the information of a first type in the at least one log message to obtain a first code; Encode the second type of information in the at least one log message to obtain a second encoding; as well as The resulting encoding is obtained at least partially by combining the first encoding and the second encoding.
17. The system of claim 16, wherein the one or more processors are configured to: encode information of a third type in the at least one log message to obtain a third encoding, wherein the resulting encoding is obtained at least in part by combining the first encoding, the second encoding, and the third encoding.
18. The system of claim 17, wherein the one or more processors are configured to use an attention layer to at least combine the first encoding, the second encoding, and the third encoding.
19. The system of claim 17, wherein the one or more processors are configured to use at least one neural network to encode at least one of the first type of information, the second type of information, or the third type of information.
20. The system of claim 16, wherein the one or more processors are configured to perform anomaly detection using the result encoding.
Citation Information
Patent Citations
Using neural networks to classify logs
US20250335549A1
Using contrastive learning to train neural networks
US20250335761A1
Using similarity loss to train neural networks
US20250335762A1