A method, system and medium for semantic embedding and event serialization of distributed unstructured logs
Patent Information
- Application Number
- CN202610610367.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-06
- Publication Date
- 2026-09-01
AI Technical Summary
[0004]1、语义提取精度不足,传统方法主要依赖正则表达式规则匹配或浅层词袋模型,如Drain、Spell或AEL等日志解析器,虽然能够提取模板,但对日志文本的深层语义关系捕捉能力有限,例如在多源异构日志中难以区分上下文相似的不同事件,导致事件识别准确率通常低于80%,无法有效支持复杂的攻击分析或故障溯源;
[0028] This invention constructs an enhanced Transformer model to semantically embed preprocessed log entries, generating high-dimensional embedding vectors. The model includes a multi-layer hybrid encoder, with each layer alternately stacking an adaptive multi-head self-attention sub-layer, a temporal autoregressive convolutional sub-layer, and a knowledge injection sub-layer. It can capture deep semantic relationships and contextual differences in log text, effectively distinguish similar events in multi-source heterogeneous logs, and support complex attack analysis and fault tracing.
Smart Images

Figure FT_1 
Figure FT_2
Abstract
Description
Technical Field
[0001] This invention relates to the field of distributed big data processing technology, and in particular to a method, system, and medium for semantic embedding and event serialization of distributed unstructured logs. Background Technology
[0002] With the rapid expansion and increasing complexity of distributed systems, unstructured log data generated by cloud computing, IoT, big data platforms, and enterprise-level microservice architectures is experiencing explosive growth. These logs are scattered across heterogeneous nodes, with diverse formats and complex semantics, containing rich information on system operating status, user behavior trajectories, potential security events, and fault clues. This places higher demands on real-time anomaly detection, attack chain tracing, fault diagnosis, and system operation and maintenance optimization. Particularly in areas such as network security monitoring, industrial IoT operation and maintenance, financial transaction auditing, and intelligent operation and maintenance, there is a need to quickly and accurately extract deep semantic information from massive amounts of complex unstructured logs and construct complete and coherent event sequences to support efficient threat response, root cause analysis, and decision optimization. Currently, log parsing and analysis technologies are widely used in commercial systems such as the ELK stack, Splunk, and Graylog, as well as open-source log processing frameworks.
[0003] However, existing distributed unstructured log processing technologies face the following key technical challenges in meeting these requirements:
[0004] 1. Insufficient semantic extraction accuracy: Traditional methods mainly rely on regular expression rule matching or shallow bag-of-words models, such as log parsers like Drain, Spell, or AEL. Although they can extract templates, they have limited ability to capture deep semantic relationships in log text. For example, in multi-source heterogeneous logs, it is difficult to distinguish different events with similar contexts, resulting in an event recognition accuracy that is usually below 80%, which cannot effectively support complex attack analysis or fault tracing.
[0005] 2. Limited event serialization capabilities: Existing technologies mostly use statistical clustering or fixed rule templates for event association, such as K-means-based clustering combined with manually defined time-series rules. This lacks adaptive modeling of dynamic time-series dependencies and causal relationships, resulting in fragmented or missing long-distance associations in the generated sequences. The sequence integrity is generally no more than 85%, and key event chains are easily missed in real-time scenarios.
[0006] 3. Insufficient efficiency and privacy protection in distributed processing: Although existing distributed frameworks such as batch processing based on Hadoop or Spark can handle large-scale logs, their sharding strategies are static, their load balancing mechanisms are simple, and they lack effective privacy protection measures when synchronizing parameters between nodes. This results in high processing latency (often on the order of several minutes) and the risk of sensitive information leakage. The probability of privacy leakage can reach 10%-15%, which limits its application in privacy-sensitive scenarios.
[0007] Therefore, a method for semantic embedding and event serialization of distributed unstructured logs is needed to solve the above problems. Summary of the Invention
[0008] The purpose of this application is to overcome the shortcomings of the existing technology and provide a method, system and medium for semantic embedding and event serialization of distributed unstructured logs.
[0009] To achieve the above objectives, this application provides the following technical solution:
[0010] In a first aspect, embodiments of this application provide a method for semantic embedding and event serialization of distributed unstructured logs, comprising the following steps:
[0011] Acquire multi-source unstructured log data generated by a distributed system. This log data includes discrete log entries from heterogeneous nodes, and each log entry contains a text description, timestamp, source identifier, variable fields, and metadata tags.
[0012] Adaptive preprocessing is performed on log entries. A neural network log pattern discoverer is used in combination with finite state automata and regular expressions to parse text descriptions, identify and extract variable fields, replace variable fields with semantically enhanced typed placeholders, and generate a standardized token sequence containing contextual hints.
[0013] An enhanced Transformer model is constructed to perform semantic embedding on preprocessed log entries and generate high-dimensional embedding vectors. This enhanced Transformer model includes a multi-layer hybrid encoder. Each multi-layer hybrid encoder alternately stacks an adaptive multi-head self-attention sub-layer, a temporal autoregressive convolutional sub-layer, and a knowledge injection sub-layer. The data is input into the multi-layer hybrid encoder for processing through data interaction.
[0014] A dynamic weighted similarity hypergraph is constructed based on high-dimensional embedding vectors. The hyperedges of this dynamic weighted similarity hypergraph are jointly calculated by the multi-dimensional similarity index of the high-dimensional embedding vectors. Log entries are clustered by spectral clustering combined with community evolution algorithm through instruction transmission to form event clusters of adaptive size, supporting incremental clustering processing of real-time log streams.
[0015] Advanced event serialization processing is performed on each event cluster to construct a multi-layer event causal dependency graph within the cluster. A graph traversal algorithm combined with reinforcement learning is applied through communication connection to optimize the path and generate a structured event sequence. The parameter aggregation adopts a variational attention mechanism to fuse the variable fields within the cluster into a probability distribution vector.
[0016] The above steps are executed in parallel in a distributed computing environment. Log data is distributed to computing nodes using an intelligent sharding strategy based on log semantic density, source heterogeneity, and real-time load. A federated learning mechanism combined with asynchronous differential privacy is used between nodes to synchronize high-dimensional embedding vectors, clustering results, and serialization parameters through data transmission paths, ensuring data privacy protection and global model consistency.
[0017] The input layer of the multi-layer hybrid encoder fuses adaptive triangular position encoding based on relative time difference, learnable source embedding matrix, and external knowledge graph embedding. The adaptive triangular position encoding, learnable source embedding matrix, and external knowledge graph embedding are added element-wise with the standardized token sequence through data interaction and then input into the multi-layer hybrid encoder. The adaptive multi-head self-attention sub-layer dynamically adjusts the number and dimension of attention heads according to the sequence length of log entries, the diversity of variable fields, and computing resources. Each attention head uses an independent query, key, and value projection matrix combined with a sparse attention mask to calculate semantic relevance and introduces a cross-attention mechanism to fuse the contextual information of adjacent log entries. The calculation results are passed to the temporal autoregressive convolutional sub-layer through data interaction.
[0018] The temporal autoregressive convolutional sublayer uses an autoregressive dilated convolutional kernel with an adaptively increasing dilation rate based on the log time span. This kernel is alternately placed after the adaptive multi-head self-attention sublayer and captures long-distance temporal dependencies through instruction transmission. At the same time, a multilayer perceptron, layer normalization, and variable residual connections are incorporated into the feedforward network to optimize gradient flow and model convergence speed.
[0019] The knowledge injection sublayer links log-related entities to a pre-built domain knowledge base through an external knowledge graph embedding module, generates knowledge-enhanced vectors, and fuses them layer by layer with adaptive triangular position encoding and learnable source embedding matrix. The learnable source embedding matrix is pre-trained through contrastive learning to distinguish semantic deviations between nodes, and the fusion result is input into the next layer of multi-layer hybrid encoder through communication connection.
[0020] The multidimensional similarity metrics include cosine similarity, Euclidean distance, and temporal correlation coefficient. The dynamic weighted similarity hypergraph retains only high-confidence hyperedges through a threshold filtering mechanism and introduces a time window sliding mechanism to track the evolution of historical event clusters, supporting cluster merging, splitting, and extinction operations. At the same time, an abnormal cluster detection submodule is applied to remove noisy clusters based on the high-dimensional embedding vector density distribution, and the stability and accuracy of the clustering results are ensured through instruction transmission.
[0021] The nodes of the multi-layer event causal dependency graph correspond to a subset of log entries. The edges are jointly determined by timestamp sequences, high-dimensional embedding vector matching, and predefined causal rule templates. The rule inference engine is integrated to extract potential causal patterns, and the node features are propagated through a graph attention network to enhance the edge weights. The predefined causal rule templates are dynamically updated from self-supervised learning and adapt to the log patterns of different distributed systems through communication connection methods.
[0022] The graph traversal algorithm combines reinforcement learning with a state-action-reward model to optimize the sequence generation path. The state represents the current multi-layer event causal dependency graph node, the action corresponds to the edge selection, and the reward is based on sequence coherence and information coverage. Robust structured event sequences are generated through data transmission paths.
[0023] The intelligent sharding strategy is based on a deep reinforcement learning agent that monitors the semantic density of incoming logs, source distribution, and node computational load in real time, and performs dynamic partitioning adjustments. It uses ring consistent hashing combined with virtual nodes to achieve load balancing, and integrates a fault prediction module to pre-migrate high-risk shards. It also improves system fault tolerance through instruction transmission.
[0024] The federated learning mechanism trains sub-models locally on each node, only synchronizing gradient updates and parameter increments after noise injection. It also uses a secure multi-party computation protocol to distribute global thresholds, knowledge graph fragments, and predefined causal rule templates, ensuring the security and privacy of cross-node collaboration through data interaction.
[0025] Secondly, embodiments of this application provide a semantic embedding and event serialization system for distributed unstructured logs. The system includes a memory and a processor. The memory includes a program for a method of semantic embedding and event serialization of distributed unstructured logs. When the program for semantic embedding and event serialization of distributed unstructured logs is executed by the processor, it implements the steps of the method of semantic embedding and event serialization of distributed unstructured logs as described above.
[0026] Thirdly, embodiments of this application provide a computer-readable storage medium storing program code, which, when executed by a processor, implements the steps of a distributed unstructured log semantic embedding and event serialization method as described above.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] This invention constructs an enhanced Transformer model to semantically embed preprocessed log entries, generating high-dimensional embedding vectors. The model includes a multi-layer hybrid encoder, with each layer alternately stacking an adaptive multi-head self-attention sub-layer, a temporal autoregressive convolutional sub-layer, and a knowledge injection sub-layer. It can capture deep semantic relationships and contextual differences in log text, effectively distinguish similar events in multi-source heterogeneous logs, and support complex attack analysis and fault tracing.
[0029] This invention constructs a dynamic weighted similarity hypergraph based on high-dimensional embedding vectors, and applies spectral clustering combined with community evolution algorithm to cluster events into clusters. Then, a multi-layer event causal dependency graph is constructed within each event cluster. A structured event sequence is generated by optimizing the path through graph traversal algorithm combined with reinforcement learning. The parameter aggregation adopts a variational attention mechanism, which can adaptively model dynamic temporal dependencies and causal relationships, avoid sequence fragmentation, and capture the event chain completely in real-time scenarios.
[0030] This invention executes the above steps in parallel in a distributed computing environment, uses an intelligent sharding strategy based on log semantic density, source heterogeneity, and real-time load to distribute log data, and employs a federated learning mechanism combined with asynchronous differential privacy synchronization of relevant parameters among nodes to ensure processing efficiency and load balancing while protecting data privacy, making it suitable for privacy-sensitive scenarios. Attached Figure Description
[0031] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 This is a flowchart of the method of the present invention;
[0033] Figure 2 This is a system architecture diagram of the present invention. Detailed Implementation
[0034] The technical solutions of the embodiments of this application will now be described with reference to the accompanying drawings. It should be noted that similar reference numerals and letters in the following drawings indicate similar items; therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0035] The terms “comprising,” “including,” or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase “comprising one…” does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0036] The terms “first,” “second,” etc., are used only to distinguish one entity or operation from another, and should not be construed as indicating or implying relative importance, nor as requiring or implying any such actual relationship or order between these entities or operations. Specific Implementation Example 1:
[0038] like Figures 1 to 2 As shown, a method for semantic embedding and event serialization of distributed unstructured logs includes the following steps:
[0039] Sp1: Obtain multi-source unstructured log data generated by the distributed system. This log data includes discrete log entries from heterogeneous nodes. Each log entry contains a text description, timestamp, source identifier, variable fields, and metadata tags.
[0040] Sp2. Adaptive preprocessing is performed on log entries. A neural network log pattern discoverer is used in combination with finite state automata and regular expressions to parse the text description, identify and extract variable fields, replace the variable fields with semantically enhanced typed placeholders, and generate a standardized token sequence containing contextual hints.
[0041] Sp3. Construct an enhanced Transformer model to perform semantic embedding on the preprocessed log entries and generate high-dimensional embedding vectors. This enhanced Transformer model includes a multi-layer hybrid encoder. Each multi-layer hybrid encoder alternately stacks an adaptive multi-head self-attention sub-layer, a temporal autoregressive convolutional sub-layer, and a knowledge injection sub-layer. The data is input into the multi-layer hybrid encoder for processing through data interaction.
[0042] Sp4. A dynamic weighted similarity hypergraph is constructed based on high-dimensional embedding vectors. The hyperedges of this dynamic weighted similarity hypergraph are jointly calculated by the multi-dimensional similarity index of the high-dimensional embedding vectors. Log entries are clustered by spectral clustering combined with community evolution algorithm through instruction transmission to form event clusters of adaptive size, supporting incremental clustering processing of real-time log streams.
[0043] Sp5. Perform advanced event serialization processing on each event cluster, construct a multi-layer event causal dependency graph within the cluster, and generate structured event sequences by applying a graph traversal algorithm combined with reinforcement learning to optimize the path through communication connection. The parameter aggregation adopts a variational attention mechanism to fuse intra-cluster variable fields into probability distribution vectors.
[0044] Sp6. The above steps are executed in parallel in a distributed computing environment. Log data is distributed to computing nodes using an intelligent sharding strategy based on log semantic density, source heterogeneity, and real-time load. A federated learning mechanism combined with asynchronous differential privacy is used between nodes to synchronize high-dimensional embedding vectors, clustering results, and serialization parameters through data transmission paths, ensuring data privacy protection and global model consistency.
[0045] In step Sp3, the input layer of the multilayer hybrid encoder is fused based on the adaptive triangular position coding with relative time difference, the learnable source embedding matrix, and the external knowledge graph embedding. The adaptive triangular position coding, the learnable source embedding matrix, and the external knowledge graph embedding are added element-wise to the standardized token sequence through data interaction and then input into the multilayer hybrid encoder.
[0046] In step Sp3, the adaptive multi-head self-attention sublayer dynamically adjusts the number and dimension of attention heads based on the sequence length of log entries, the diversity of variable fields, and computational resources. Each attention head uses an independent query, key, and value projection matrix combined with a sparse attention mask to calculate semantic relevance. A cross-attention mechanism is introduced to fuse the contextual information of adjacent log entries. The calculation results are passed to the temporal autoregressive convolutional sublayer through data interaction.
[0047] In step Sp3, the temporal autoregressive convolutional sublayer uses an autoregressive dilated convolutional kernel with an adaptively increasing dilation rate based on the log time span. This kernel is alternately placed after the adaptive multi-head self-attention sublayer and captures long-distance temporal dependencies through instruction transmission. At the same time, a multilayer perceptron, layer normalization, and variable residual connections are incorporated into the feedforward network to optimize gradient flow and model convergence speed.
[0048] In step Sp3, the knowledge injection sublayer links log-related entities to a pre-built domain knowledge base through an external knowledge graph embedding module, generates knowledge-enhanced vectors, and fuses them layer by layer with adaptive triangular position encoding and learnable source embedding matrix. The learnable source embedding matrix is pre-trained through contrastive learning to distinguish semantic deviations between nodes, and the fusion result is input into the next layer of multi-layer hybrid encoder through communication connection.
[0049] In step Sp4, the multidimensional similarity metrics include cosine similarity, Euclidean distance, and temporal correlation coefficient. The dynamic weighted similarity hypergraph retains only high-confidence hyperedges through a threshold filtering mechanism and introduces a time window sliding mechanism to track the evolution of historical event clusters, supporting cluster merging, splitting, and extinction operations. At the same time, an abnormal cluster detection submodule is applied to remove noisy clusters based on the high-dimensional embedding vector density distribution. The stability and accuracy of the clustering results are ensured through instruction transmission.
[0050] In step Sp5, the nodes of the multi-layer event causal dependency graph correspond to a subset of log entries. The edges are jointly determined by timestamp sequences, high-dimensional embedding vector matching, and predefined causal rule templates. The rule inference engine is integrated to extract potential causal patterns, and the node features are propagated through a graph attention network to enhance the edge weights. The predefined causal rule templates are dynamically updated from self-supervised learning and adapt to the log patterns of different distributed systems through communication connections.
[0051] In step Sp5, the graph traversal algorithm combines reinforcement learning with a state-action-reward model to optimize the sequence generation path. Here, the state represents the node of the current multi-layer event causal dependency graph, the action corresponds to the edge selection, and the reward is based on the sequence coherence and information coverage. Robust structured event sequences are generated through data transmission paths.
[0052] In step Sp6, the intelligent sharding strategy dynamically adjusts partitions based on the real-time monitoring of log inflow semantic density, source distribution, and node computational load by a deep reinforcement learning agent. It uses ring consistent hashing combined with virtual nodes to achieve load balancing and integrates a fault prediction module to pre-migrate high-risk shards, thereby improving system fault tolerance through instruction transmission.
[0053] In step Sp6, the federated learning mechanism trains sub-models locally on each node, synchronizing only gradient updates and parameter increments after noise injection. At the same time, it uses a secure multi-party computation protocol to distribute global thresholds, knowledge graph fragments, and predefined causal rule templates, ensuring the security and privacy of cross-node collaboration through data interaction. Specific Implementation Example 2:
[0055] like Figures 1 to 2As shown, based on Example 1, this example further refines and specifically parameterizes each step to adapt to large-scale distributed log processing scenarios.
[0056] In step Sp1, the multi-source unstructured log data generated by the distributed system is collected in real time through the Kafka message queue at a collection frequency of 100,000 log entries per second. The heterogeneous nodes include application server nodes, database nodes, and network device nodes. Each log entry contains a text description, timestamp, source identifier, variable fields, and metadata tags. The timestamp is accurate to the millisecond level, the source identifier is a combination of node IP and service name, and the metadata tags include log level and thread ID.
[0057] In step Sp2, adaptive preprocessing employs a two-stage process: the first stage utilizes a pre-trained BERT model as a neural network log pattern discoverer to perform preliminary pattern matching on log entries, identifying common log templates; the second stage combines finite state automata and regular expressions to perform fine-grained parsing of unmatched entries, identifying and extracting variable fields, including IP addresses, user IDs, error codes, and numerical parameters, and replacing these variable fields with semantically enhanced typed placeholders, such as... <ip>,<USER_ID> ,<ERROR_CODE> and <num>Meanwhile, context hint tokens [CLS] and [LOG_TYPE] are added to the beginning of the standardized token sequence to generate a standardized token sequence of length 512.
[0058] In step Sp3, the enhanced Transformer model has a total of 12 layers, with each layer having a multi-layer hybrid encoder with a hidden dimension of 768. The adaptive triangular position encoding uses a combination of sine and cosine functions with a frequency base of 10,000. The learnable source embedding matrix has a dimension of 128. It is pre-trained on a historical log dataset through contrastive learning. The external knowledge graph embeddings are derived from domain-specific knowledge bases, such as network security event knowledge bases or system operation and maintenance knowledge bases.
[0059] In the adaptive multi-head self-attention sub-layer of step Sp3, the number of attention heads is dynamically adjusted according to the sequence length. When the sequence length is less than 128, the number of heads is 8, and when the sequence length is between 128 and 512, the number of heads is 16. Each attention head uses an independent query, key, and value projection matrix. The sparse attention mask only retains the attention weights of the 32 positions before and after the current token. The cross-attention mechanism uses the current log entry as the query and the five adjacent log entries before and after as the key and value for fusion.
[0060] In the temporal autoregressive convolutional sublayer of step Sp3, the size of the autoregressive dilated convolutional kernel is 7, and the dilation rate is 1, 2, 4, 8, 16, and 32 from layer 1 to layer 12. The residual connections use a variable scaling factor, which is dynamically adjusted from 0.9 to 1.1 according to the layer depth. The feedforward network includes two linear layers with an intermediate dimension of 3072 and uses the GELU activation function.
[0061] In the knowledge injection sub-layer of step Sp3, the external knowledge graph embedding module first extracts key entities from log entries, such as event names and operation types, through entity recognition. Then, it queries the pre-built domain knowledge base to obtain the triplet relationship embedding of the corresponding entity, generates a knowledge enhancement vector with a dimension of 256, and performs layer normalization after adding it element by element with the position encoding and source embedding.
[0062] In step Sp4, the specific calculation formula for the multidimensional similarity index is joint weight w = 0.5 × cosine similarity + 0.3 × (1 - normalized Euclidean distance) + 0.2 × temporal correlation coefficient. The adaptive threshold of the threshold filtering mechanism is the 95th quantile of the embedded vector space. The window size of the time window sliding mechanism is 5 minutes. It supports processing 1000 new log entries per batch when performing incremental clustering. The abnormal cluster detection submodule adopts the local anomaly factor algorithm with a threshold of 3.
[0063] In step Sp5, the multi-layer event causal dependency graph is divided into two layers: the bottom layer is time series connection and the upper layer is semantic causal connection. The predefined causal rule templates include "login failure followed by login success" and "query operation followed by response return". The rule inference engine adopts forward chain inference, the graph attention network has 3 layers, the reinforcement learning optimization adopts the PPO algorithm, the reward function is 0.7×sequence coherence score + 0.3×information coverage score, the variational distribution of the variational attention mechanism is Gaussian distribution, and the mean and variance are predicted by two feedforward networks respectively.
[0064] In step Sp6, the state space of the deep reinforcement learning agent of the intelligent sharding strategy includes the current log semantic density, source distribution vector, and node load vector. The action space is a combination of shard number and node allocation. The reward is the sum of the negative average processing latency and load variance. The number of virtual nodes for the ring consistent hash is 100 per physical node. The fault prediction module predicts based on node heartbeat and resource utilization. The threshold is that the CPU utilization is 90% for more than 30 seconds.
[0065] In the federated learning mechanism of step Sp6, each computing node is trained locally for 3 epochs and then synchronized. The differential privacy noise scale is 1.0. The secure multi-party computation protocol uses a secret sharing scheme to distribute global thresholds, knowledge graph fragments and predefined causal rule templates. The synchronization frequency is once every 1 million log entries processed. Specific Implementation Example 3:
[0067] like Figures 1 to 2 As shown in Example 1, this example provides a more detailed explanation of the core algorithms involved in the method:
[0068] The enhanced Transformer model takes a preprocessed, normalized token sequence as input, including text descriptions, timestamps, and source identifiers for log entries. The output is a high-dimensional embedding vector representing the semantic features of the log entries. In the application logic, the algorithm first receives the normalized token sequence as basic input. Then, the input layer transforms the sequence into an initial embedding representation, including adding positional encoding to capture the sequence order. Next, the data enters a multi-layer hybrid encoder for layer-by-layer processing. Each encoder layer first applies an adaptive multi-head self-attention sublayer to compute global dependencies between tokens, generating attention-enhanced representations through weighted summation. Subsequently, a temporal autoregressive convolutional sublayer processes these representations, applying convolution operations position-by-position to extract temporal patterns. Then, a knowledge injection sublayer incorporates external information to further adjust the representation. At the end of each layer, layer normalization and residual connections are applied, summing the current layer's output with the input to stabilize gradient flow. After all layers, the average or positional output of the final layer is taken as the high-dimensional embedding vector, supporting subsequent clustering and serialization. This computation process emphasizes parallel processing of multiple tokens to avoid sequential dependencies and improve efficiency.
[0069] The input data for the adaptive multi-head self-attention sublayer is an intermediate representation of a standardized token sequence, including the sequence length and variable field diversity of the current log entry; the output is an attention-weighted feature vector. In the application logic, the algorithm first dynamically determines the number and dimensions of attention heads based on the sequence length and variable diversity, for example, using fewer heads for shorter sequences to save computation. Then, it generates independent query, key, and value matrices for each head, calculating the similarity score of each token to other tokens through matrix multiplication to form an attention weight matrix. Combined with a sparse attention mask, only the weights at key positions are retained, avoiding the computation of irrelevant terms. Next, a cross-attention mechanism is introduced, using the current log entry as the query and the representations of adjacent log entries as keys and values, performing additional similarity calculations and fusing the results. Finally, the key-value vector is multiplied by the attention weights through weighted summation to generate an enhanced feature vector, which is then concatenated with the multi-head results and projected back to the original dimensions, thereby improving the semantic accuracy of the embedded vector. This process allows the model to simultaneously focus on local and global semantic relationships in the logs.
[0070] The input data for the temporal autoregressive convolutional sublayer is the feature vector output from the self-attention sublayer, including log time span information; the output is a temporally enhanced feature vector. In the application logic, the algorithm first adaptively sets the dilation rate according to the log time span, for example, using a larger dilation for longer spans to cover long-distance dependencies; then, it employs an autoregressive dilated convolutional kernel, applying convolution operations position-by-position to the input vector, considering only information from the current and previous positions to maintain causality; after convolution, the result is processed through a feedforward network, including expanding the dimension in the first linear layer, calculating nonlinear transformation element-by-element in a multilayer perceptron, compressing back to the original dimension in the second linear layer, and applying the GELU activation function to enhance nonlinearity; simultaneously, layer normalization is incorporated to standardize the output, and variable residual connections are used to add the convolution result to the input, dynamically optimizing the connection weights based on layer depth; these are alternately placed after the self-attention sublayer to ensure continuous injection of temporal information. This computation process captures the temporal patterns of log events, such as event sequence and intervals, improving the model's understanding of dynamic logs.
[0071] The input data for the knowledge injection sublayer consists of location encoding, source embedding matrix, and log-related entities; the output is a knowledge-enhanced vector. In the application logic, the algorithm first identifies key entities in the logs, such as event types or operation objects, through an external knowledge graph embedding module. Then, it queries a pre-built domain knowledge base to retrieve the relationships and attributes of relevant entities, generating corresponding embedding vectors. Next, these knowledge vectors are added element-wise with adaptive triangular location encoding (calculated using trigonometric functions based on relative time differences) and a learnable source embedding matrix (pre-trained through contrastive learning, comparing the similarity of logs from different sources to learn biases) to form a fused representation. After fusion, layer-by-layer injection is performed, with layer normalization applied at the end of each layer to balance the scale. The source embedding matrix is optimized through contrastive learning to distinguish semantic differences between nodes. This process integrates external domain knowledge into the embedding, improving the semantic understanding depth of unstructured logs, such as identifying implicit log associations.
[0072] The input data for spectral clustering combined with the community evolution algorithm consists of high-dimensional embedding vectors and a dynamically weighted similarity hypergraph; the output is event clusters. In the application logic, the algorithm first calculates the joint weights between log entries based on multi-dimensional similarity metrics (such as cosine similarity, Euclidean distance, and temporal correlation coefficient), constructing a dynamically weighted similarity hypergraph where hyperedges connect multiple similar entries. Threshold filtering is used to retain only hyperedges with weights above the average level to reduce noise. Then, spectral methods are applied to calculate the Laplacian matrix eigenvectors of the hypergraph, and partitions are initialized using the top k smallest eigenvectors. Next, a time window sliding mechanism is introduced in conjunction with the community evolution algorithm to track historical event clusters, such as comparing the cluster overlap between the current and previous windows, performing merging (if the overlap is high), splitting (if the intra-cluster variance is large), or elimination (if the cluster size is small). Simultaneously, an anomaly cluster detection submodule calculates local anomaly factors based on the embedding vector density distribution, eliminating noisy clusters with densities below a threshold. Partition optimization is iteratively performed until the modularity stabilizes, forming stable semantic event clusters. This process handles the dynamic evolution of logs, ensuring the temporal consistency of clusters.
[0073] The graph traversal algorithm, combined with reinforcement learning, takes a multi-layered event causal dependency graph as input data, including nodes (a subset of log entries) and edges (timestamp sequences, embedding matching, and causal rules). The output is an optimized structured event sequence. In its application logic, the algorithm first constructs a state-action-reward model, where the state represents the current graph node and its neighborhood features, the action corresponds to selecting the next edge or node, and the reward is based on sequence coherence (sum of edge weights) and information coverage (log coverage ratio). Then, a reinforcement learning agent initializes the policy network. The agent traverses from the starting node, predicting action probabilities and sampling paths. During traversal, the cumulative reward is calculated, and the network parameters are updated using policy gradients, iterating multiple times until convergence. Combining graph traversal such as depth-first search as a baseline, reinforcement learning optimizes the path to maximize the reward and avoid local optima. Finally, a coherent event sequence is generated, supporting log analysis such as attack chain detection. This process learns to adapt to complex graph structures, improving the accuracy of sequence generation.
[0074] The variational attention mechanism takes in-cluster variable fields as input and outputs a probability distribution vector. In its application logic, the algorithm first represents the variable fields as a sequence of vectors. Then, it estimates the posterior distribution through variational inference, using two feedforward networks to predict the mean and variance, respectively, forming a Gaussian distribution sample. Next, it calculates attention weights, weights each field vector, and fuses them to generate an aggregate vector. Simultaneously, KL divergence regularization is introduced to balance the variational and prior distributions. Iterative optimization handles uncertainties such as variable noise. This process generates a robust parameter vector that integrates uncertain information during event serialization.
[0075] The federated learning mechanism, combined with asynchronous differential privacy, takes local sub-model parameters and high-dimensional embedding vectors as input data; the output is globally consistent model parameters. In the application logic, the algorithm first trains a sub-model locally on each node, updating parameters using local log data. Then, it asynchronously collects updates, immediately uploading gradients after noise injection (Gaussian noise is added to achieve differential privacy) when a node completes its work. A central server aggregates these gradients, calculates the average increment, and distributes them back to the nodes. A secure multi-party computation protocol is used to distribute global thresholds, knowledge graph fragments, and rule templates, ensuring no leakage of the original data. This process iterates multiple times until the model converges. This process ensures privacy protection and cross-node collaboration.
[0076] The intelligent sharding strategy, based on a deep reinforcement learning agent, takes log semantic density, source distribution, and node load as input data and outputs a dynamic partitioning adjustment scheme. In the application logic, the algorithm monitors the state space in real time through the agent, including density vectors and load metrics; the agent predicts actions, such as adjusting the number of shards or node allocation, using ring-consistent hashing to map data to virtual nodes for load balancing; the reward is calculated as the sum of negative latency and variance, and the agent network is updated; a fault prediction module is integrated to pre-migrate shards based on resource thresholds. This process optimizes resource allocation and improves distributed efficiency. Specific Implementation Example 4:
[0078] like Figures 1 to 2 As shown, based on Embodiment 1, this embodiment describes the specific practical application scenario of the method and explains the hardware / software environment configuration. At the same time, the effectiveness of the method is verified through implementation results.
[0079] This method is applicable to log analysis scenarios in large-scale distributed systems, such as security monitoring systems for cloud computing platforms. Unstructured logs originate from multiple microservice nodes, including user authentication services, data storage services, and network routing services. These logs contain anomalous events such as login failures, data access violations, and network latency, requiring semantic embedding and event serialization to transform them into structured sequences for detecting potential attack chains. In another scenario, it is suitable for the operation and maintenance management of IoT device clusters, such as sensor logs in smart factories, involving events like equipment failures and data transmission interruptions. The method processes scattered logs to generate event sequences, supporting real-time fault diagnosis and predictive maintenance. Furthermore, in financial transaction systems, this method can process transaction logs, audit logs, and system monitoring logs to identify fraud patterns such as abnormal transfer sequences.
[0080] Hardware / Software Environment Description: The hardware configuration includes one master server (CPU: Intel Xeon Gold 6240, Memory: 128GB, Storage: 2TB SSD) and at least four compute nodes (each node has a CPU: AMD EPYC7302, Memory: 64GB, Storage: 1TB SSD), with a 10Gbps Ethernet connection. The software environment is based on Ubuntu 20.04 operating system, Python 3.9 environment, with PyTorch 1.12.0, NetworkX 2.8, Scikit-learn 1.0.2, and Apache Spark 3.2.1 framework installed to ensure the reproducibility of the method. In the concrete implementation, log data is first obtained from the master server, distributed to the compute nodes via Spark for preprocessing and embedding, and then the results are synchronized between nodes to generate an event sequence.
[0081] The effectiveness was validated through comparative experiments. The experimental dataset consisted of 1 million unstructured log entries generated by a simulated distributed system, including normal operations and injected abnormal events. The baseline method was compared to a traditional log parser combined with an LSTM sequence model. This method was evaluated in terms of semantic accuracy, processing time, and privacy risk. The experimental data is shown in the table below:
[0082] Language accuracy (%) 92.5 78.3 Processing time (seconds / 100,000 records) 45 120 Privacy breach risk (differential privacy ε value) 1.2 5.8 Event sequence integrity (%) 95.1 82.4
[0083] Experimental data demonstrate that semantic accuracy, calculated using manually annotated semantic similarity, is 14.2 percentage points higher than the benchmark method, indicating the superiority of Transformer embedding. Processing time is reduced by 62.5% in this method, thanks to distributed parallelism and intelligent sharding. Privacy leakage risk is measured by differential privacy ε; the lower value of this method shows the effective protection of the federated learning mechanism. Event sequence integrity assessment shows a 12.7% improvement in the proportion of the sequence covering the original log, proving the robustness of serialization processing. These data validate the effectiveness and superiority of the method in real-world scenarios.
[0084] In practical applications, this method can be used in network monitoring systems of telecommunications network operators to process distributed unstructured logs from base stations, routers, and switches. These logs include events such as signal interference, connection drops, and abnormal traffic. By using Transformer semantic embedding to capture the semantic relationships in the logs and serializing them into event chains, it can be used for real-time network optimization and fault location, such as detecting DDoS attack sequences and improving network stability. When applied to user behavior analysis systems in e-commerce platforms, this method processes log entries from front-end applications, back-end servers, and payment gateways, involving events such as user browsing, shopping cart operations, and payment failures. It generates user behavior path sequences through event serialization, supporting the detection of abnormal behaviors such as fraudulent transactions or account theft, and optimizing recommendation algorithms to improve user experience. Specific Implementation Example 5:
[0086] like Figures 1 to 2 As shown, based on Embodiment 1, this embodiment is a variant embodiment that demonstrates the flexibility of the method by adjusting some modules and parameters to adapt to different log types and scales.
[0087] In this variant, for small-scale distributed systems such as edge computing environments, the enhanced Transformer model is simplified to 6 layers to reduce computational overhead. It retains the adaptive multi-head self-attention sub-layer and the temporal autoregressive convolution sub-layer, but omits the knowledge injection sub-layer to reduce external dependencies. In step Sp2, adaptive preprocessing uses only finite state automata and regular expressions, skipping the neural network log pattern discoverer, making it suitable for scenarios with simpler log patterns, such as single application server logs. In step Sp4, the dynamic weighted similarity hypergraph is replaced with a standard similarity graph, and spectral clustering combined with a community evolution algorithm simplifies the time window sliding mechanism. It only supports static clustering and does not process real-time log streams, making it suitable for offline analysis. In step Sp5, the graph traversal algorithm combined with reinforcement learning is changed to depth-first search combined with simple heuristics, and the variational attention mechanism is simplified to mean aggregation, reducing optimization complexity. In step Sp6, the intelligent sharding strategy is based on static hash allocation instead of a deep reinforcement learning agent, and the federated learning mechanism removes asynchronous differential privacy, using only synchronous parameter averaging, making it suitable for environments with lower privacy requirements.
[0088] Another variant targets scenarios with high privacy requirements, such as distributed log systems in healthcare. It enhances the differential privacy noise scale of the federated learning mechanism and introduces local data augmentation techniques. In step Sp3, the external knowledge graph embedding is limited to a privacy-preserving internal knowledge base to avoid external queries. In step Sp6, multiple rounds of federated iterations are added to ensure global model consistency, while the sharding strategy prioritizes data locality to minimize data transfer.
[0089] These variants maintain the integrity of the core steps and adapt to diverse needs, ranging from small edge systems to large, high-privacy clusters, through modular adjustments and parameter optimizations, demonstrating the flexibility and scalability of the approach. Specific Implementation Example Six:
[0091] like Figures 1 to 2 As shown, based on Embodiment 1, this embodiment is another variant embodiment that further demonstrates the flexibility of the method and focuses on real-time streaming log processing and multimodal log integration.
[0092] In this variant, a Kafka streaming framework is introduced for real-time streaming logs, such as dynamic logs from video surveillance systems. In step Sp1, log streams are subscribed to in real-time, with 5000 log entries processed in batches. In step Sp3, an enhanced Transformer model adds an online learning mechanism, fine-tuning model parameters after each batch of logs and supporting progressive embedding. In step Sp4, spectral clustering combined with a community evolution algorithm is extended to an incremental version, updating the hypergraph through a sliding window, recalculating only the similarity of newly added logs, reducing the overhead of full re-clustering. In step Sp5, a multi-layered event causal dependency graph supports dynamic updates, and a graph traversal algorithm combined with reinforcement learning employs an online reinforcement strategy, adjusting the reward model in real-time to adapt to evolving event patterns. In step Sp6, an intelligent sharding strategy integrates a real-time load feedback loop, adjusting shards every minute, and a federated learning mechanism adds edge node support, allowing mobile devices to participate in computation.
[0093] Furthermore, for multimodal logs, such as mixed logs containing text and images, step Sp2 is extended to parse image metadata and convert it into an additional token sequence that is then incorporated into a standardized token sequence. In step Sp3, the knowledge injection sublayer fuses a visual knowledge graph to enhance the embedded multimodal representation. In step Sp4, the multidimensional similarity metric adds image feature similarity calculation, supporting cross-modal clustering.
[0094] These variants extend the approach to dynamic and complex logging environments through streaming processing and multimodal integration, demonstrating their adaptability and innovative potential across different applications. Specific Implementation Example 7:
[0096] like Figures 1 to 2 As shown, based on Embodiments 1 to 6, this embodiment provides supplementary explanations of technical details that have not been fully disclosed or may not be clearly described in the foregoing embodiments, in order to ensure that the technical solution of the present invention is complete and clear, and is easy for those skilled in the art to implement.
[0097] The specific generation method of the standardized token sequence is as follows: After the adaptive preprocessing in step Sp2, the standardized token sequence uses a tokenizer to convert the log text into an integer ID sequence. The maximum length of the sequence is fixed at 512. If the length is insufficient, it is padded with padding tokens [PAD]. If the length exceeds 512, it is truncated and the first 512 tokens are retained. A classification token [CLS] is always added to the beginning of the sequence. Its final hidden state can be used as the aggregate representation of the entire log. The context hint token [LOG_TYPE] is dynamically inserted according to the log level or source identifier to guide the model to distinguish different types of logs.
[0098] Dimensions and usage of high-dimensional embedding vectors: After step Sp3, the high-dimensional embedding vectors output by the enhanced Transformer model are fixed at 768 dimensions. This vector is the output vector corresponding to the [CLS] token, or it is obtained by averaging the hidden states of all tokens. This vector is directly used as the input feature for similarity calculation in the subsequent step Sp4 without additional dimensionality reduction processing, so as to preserve the complete semantic information.
[0099] The specific storage and construction method of the dynamic weighted similarity hypergraph: In step Sp4, the dynamic weighted similarity hypergraph is stored using a sparse matrix, and only hyperedges with weights higher than the threshold are saved; each hyperedge connects a maximum of 5 log entries (i.e., each hyperedge contains a central entry and its 4 nearest neighbors) to control computational complexity; the joint calculation of multidimensional similarity indices adopts a weighted summation method, with the weight coefficients fixed as cosine similarity 0.6, Euclidean distance contribution 0.2, and temporal correlation coefficient 0.2.
[0100] The specific hierarchical structure of the multi-layer event causal dependency graph: In step Sp5, the multi-layer event causal dependency graph is clearly divided into two layers: the first layer is a strict time order layer, which connects a subset of log entries only according to the ascending order of timestamps to form a directed acyclic graph; the second layer is a semantic causal layer, which adds cross-time edges determined by predefined causal rule templates and high-dimensional embedding vector matching on the basis of the first layer; the nodes of the graph store the high-dimensional embedding vectors and variable fields of the corresponding log entries, and the edges store the weights determined jointly.
[0101] The source and update method of predefined causal rule templates: The predefined causal rule templates are initially obtained from manual configuration by domain experts or mining from historical logs and stored as a conditional probability table of "preceding event type → subsequent event type". During operation, they are dynamically updated through self-supervised learning, that is, the actual successor relationship within the event cluster is statistically analyzed and the template probability is smoothed by exponential moving average. The update cycle is once every 1 million logs processed.
[0102] The specific output form of the variational attention mechanism: In the parameter aggregation in step Sp5, the variational attention mechanism finally outputs a probability distribution vector with a dimension of 512. This vector includes the aggregated mean vector (384 dimensions) and the log-variance vector (128 dimensions). When used in downstream tasks, it is used to sample and generate specific parameter instances through reparameterization techniques to represent the uncertainty of variable fields.
[0103] Initialization and convergence judgment of the global model in the federated learning mechanism: In step Sp6, the initial parameters of the global model are obtained by the master server using a public dataset for pre-training. Each computing node downloads the initial parameters and performs local training. The convergence judgment adopts the condition that the change of the global loss function after three consecutive synchronization rounds is less than 0.001 as the stopping condition, or the maximum number of synchronization rounds is set to 50 rounds.
[0104] The specific location of noise addition for asynchronous differential privacy: In step Sp6, differential privacy noise is only added to the gradient update vector, not the original high-dimensional embedding vector; the noise is isotropic Gaussian noise, and the variance is adaptively adjusted after being clipped according to the gradient norm to ensure that the cumulative privacy budget in each round does not exceed the preset total budget.
[0105] The initial partitioning method of the intelligent sharding strategy: When the first execution of step Sp6 is performed, a simple hash based on the source identifier is used as the initial partition, and then the deep reinforcement learning agent is gradually optimized; the minimum partitioning granularity is 1000 log entries to ensure that each shard has enough data to support local training.
[0106] The specific criteria for the abnormal cluster detection submodule are as follows: In step Sp4, the abnormal cluster detection is based on the calculation of local abnormal factors. If the factor is greater than the threshold 3, the cluster is marked as noise and removed, and will not participate in the subsequent serialization process.
[0107] This application provides a system for semantic embedding and event serialization of distributed unstructured logs. The system includes a memory and a processor. The memory includes a program for a method of semantic embedding and event serialization of distributed unstructured logs. When the program for semantic embedding and event serialization of distributed unstructured logs is executed by the processor, it implements the steps of the method of semantic embedding and event serialization of distributed unstructured logs as described above.
[0108] This application provides a computer-readable storage medium storing program code. When the program code is executed by a processor, it implements the steps of a distributed unstructured log semantic embedding and event serialization method as described above.
[0109] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0110] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0111] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0112] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0113] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0114] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0115] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0116] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.< / num> < / ip>
Claims
1. A method for semantic embedding and event serialization of distributed unstructured logs, characterized in that, Includes the following steps: Acquire multi-source unstructured log data generated by a distributed system. This log data includes discrete log entries from heterogeneous nodes, and each log entry contains a text description, timestamp, source identifier, variable fields, and metadata tags. Adaptive preprocessing is performed on log entries. A neural network log pattern discoverer is used in combination with finite state automata and regular expressions to parse text descriptions, identify and extract variable fields, replace variable fields with semantically enhanced typed placeholders, and generate a standardized token sequence containing contextual hints. An enhanced Transformer model is constructed to perform semantic embedding on preprocessed log entries and generate high-dimensional embedding vectors. This enhanced Transformer model includes a multi-layer hybrid encoder. Each multi-layer hybrid encoder alternately stacks an adaptive multi-head self-attention sub-layer, a temporal autoregressive convolutional sub-layer, and a knowledge injection sub-layer. The data is input into the multi-layer hybrid encoder for processing through data interaction. A dynamic weighted similarity hypergraph is constructed based on high-dimensional embedding vectors. The hyperedges of this dynamic weighted similarity hypergraph are jointly calculated by the multi-dimensional similarity index of the high-dimensional embedding vectors. Log entries are clustered by spectral clustering combined with community evolution algorithm through instruction transmission to form event clusters of adaptive size, supporting incremental clustering processing of real-time log streams. Advanced event serialization processing is performed on each event cluster to construct a multi-layer event causal dependency graph within the cluster. A graph traversal algorithm combined with reinforcement learning is applied through communication connection to optimize the path and generate a structured event sequence. The parameter aggregation adopts a variational attention mechanism to fuse the variable fields within the cluster into a probability distribution vector. The above steps are executed in parallel in a distributed computing environment. Log data is distributed to computing nodes using an intelligent sharding strategy based on log semantic density, source heterogeneity, and real-time load. A federated learning mechanism combined with asynchronous differential privacy is used between nodes to synchronize high-dimensional embedding vectors, clustering results, and serialization parameters through data transmission paths, ensuring data privacy protection and global model consistency.
2. The semantic embedding and event serialization method for distributed unstructured logs according to claim 1, characterized in that, The input layer of the multi-layer hybrid encoder fuses adaptive triangular position encoding based on relative time difference, learnable source embedding matrix, and external knowledge graph embedding. The adaptive triangular position encoding, learnable source embedding matrix, and external knowledge graph embedding are added element-wise with the standardized token sequence through data interaction and then input into the multi-layer hybrid encoder. The adaptive multi-head self-attention sub-layer dynamically adjusts the number and dimension of attention heads according to the sequence length of log entries, the diversity of variable fields, and computing resources. Each attention head uses an independent query, key, and value projection matrix combined with a sparse attention mask to calculate semantic relevance and introduces a cross-attention mechanism to fuse the contextual information of adjacent log entries. The calculation results are passed to the temporal autoregressive convolutional sub-layer through data interaction. The temporal autoregressive convolutional sublayer uses an autoregressive dilated convolutional kernel with an adaptively increasing dilation rate based on the log time span. This kernel is alternately placed after the adaptive multi-head self-attention sublayer and captures long-distance temporal dependencies through instruction transmission. At the same time, a multilayer perceptron, layer normalization, and variable residual connections are incorporated into the feedforward network to optimize gradient flow and model convergence speed.
3. The semantic embedding and event serialization method for distributed unstructured logs according to claim 1, characterized in that, The knowledge injection sublayer links log-related entities to a pre-built domain knowledge base through an external knowledge graph embedding module, generates knowledge-enhanced vectors, and fuses them layer by layer with adaptive triangular position encoding and learnable source embedding matrix. The learnable source embedding matrix is pre-trained through contrastive learning to distinguish semantic deviations between nodes, and the fusion result is input into the next layer of multi-layer hybrid encoder through communication connection.
4. The semantic embedding and event serialization method for distributed unstructured logs according to claim 1, characterized in that, The multidimensional similarity metrics include cosine similarity, Euclidean distance, and temporal correlation coefficient. The dynamic weighted similarity hypergraph retains only high-confidence hyperedges through a threshold filtering mechanism and introduces a time window sliding mechanism to track the evolution of historical event clusters, supporting cluster merging, splitting, and extinction operations. At the same time, an abnormal cluster detection submodule is applied to remove noisy clusters based on the high-dimensional embedding vector density distribution, and the stability and accuracy of the clustering results are ensured through instruction transmission.
5. The semantic embedding and event serialization method for distributed unstructured logs according to claim 1, characterized in that, The nodes of the multi-layer event causal dependency graph correspond to a subset of log entries. The edges are jointly determined by timestamp sequences, high-dimensional embedding vector matching, and predefined causal rule templates. The rule inference engine is integrated to extract potential causal patterns, and the node features are propagated through a graph attention network to enhance the edge weights. The predefined causal rule templates are dynamically updated from self-supervised learning and adapt to the log patterns of different distributed systems through communication connection methods.
6. The semantic embedding and event serialization method for distributed unstructured logs according to claim 1, characterized in that, The graph traversal algorithm combines reinforcement learning with a state-action-reward model to optimize the sequence generation path. The state represents the current multi-layer event causal dependency graph node, the action corresponds to the edge selection, and the reward is based on sequence coherence and information coverage. Robust structured event sequences are generated through data transmission paths.
7. The semantic embedding and event serialization method for distributed unstructured logs according to claim 1, characterized in that, The intelligent sharding strategy is based on a deep reinforcement learning agent that monitors the semantic density of incoming logs, source distribution, and node computational load in real time, and performs dynamic partitioning adjustments. It uses ring consistent hashing combined with virtual nodes to achieve load balancing, and integrates a fault prediction module to pre-migrate high-risk shards. It also improves system fault tolerance through instruction transmission.
8. The semantic embedding and event serialization method for distributed unstructured logs according to claim 1, characterized in that, The federated learning mechanism trains sub-models locally on each node, only synchronizing gradient updates and parameter increments after noise injection. It also uses a secure multi-party computation protocol to distribute global thresholds, knowledge graph fragments, and predefined causal rule templates, ensuring the security and privacy of cross-node collaboration through data interaction.
9. A semantic embedding and event serialization system for distributed unstructured logs, characterized in that, The system includes a memory and a processor. The memory includes a program for a method of semantic embedding and event serialization of distributed unstructured logs. When the program for semantic embedding and event serialization of distributed unstructured logs is executed by the processor, it implements the steps of the method of semantic embedding and event serialization of distributed unstructured logs as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code, which, when executed by a processor, implements the steps of a distributed unstructured log semantic embedding and event serialization method as described in any one of claims 1 to 8.