A federated dual adaptive personalized log anomaly detection system and method thereof
Patent Information
- Application Number
- CN202610697230.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-20
- Publication Date
- 2026-09-25
AI Technical Summary
上述方案虽然具有一定效果,但大多聚焦于单一侧优化,即仅关注客户端个性化或仅关注服务器聚合策略,尚未形成“个性化程度自适应 +聚合范围自适应”相结合的双层联动机制
[0053]1、本发明能够直接根据客户端本地收益来调节个性化强度,使不同客户端获得差异化的参数共享边界,因此更适合处理日志场景中普遍存在的非独立同分布问题。特别是在跨域客户端混合条件下,本发明通过仅共享参数聚合或混合聚合,能够有效降低不相关分布引起的负迁移。
Smart Images

Figure CN122818142A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of log anomaly detection technology, specifically to a federated dual-adaptive personalized log anomaly detection system and method. Background Technology
[0002] In cloud computing, distributed systems, network security, and operations and maintenance analysis scenarios, servers, switching devices, storage systems, and applications continuously generate a large amount of logs. Log anomaly detection is a crucial means of identifying system failures, attacks, and operational anomalies. Existing centralized learning methods typically require aggregating logs from various nodes to a central location for unified training, which leads to problems such as data privacy leaks, difficulties in cross-organizational data sharing, and excessively high transmission costs. Federated learning, which allows the exchange of only model parameters or gradients without data leaving its domain, has thus become an important technological direction for intelligent log analysis.
[0003] However, log data inherently exhibits significant heterogeneity. On one hand, the distribution of log templates, anomaly rates, sequence lengths, and contextual patterns vary considerably across different systems, business processes, and time windows. On the other hand, even with a unified log parsing and vectorization method, the training sample size, anomaly sparsity, and statistical stability differ across clients. Traditional federated averaging methods aggregate all client parameters equally or by sample size, which can easily smooth out important local features during aggregation, leading to a phenomenon where some clients exhibit "global model usable but local performance poor."
[0004] To address this issue, existing personalized federated learning typically employs the following approaches: First, adding constraint terms to the loss function, such as FedProx methods which improve training stability by limiting the deviation of local updates from the global model; second, adding local adaptation layers or local aggregation modules outside the global model, such as FedALA methods which improve individual adaptation capabilities through adaptive local aggregation; and third, using sparse masks or lottery-related methods to selectively share different parameters to retain some client-specific knowledge. While these solutions have some effectiveness, most focus on single-sided optimization, i.e., only considering client personalization or only considering server aggregation strategies, and have not yet formed a two-layer linkage mechanism combining "adaptive personalization degree + adaptive aggregation scope".
[0005] The main drawbacks of existing technologies include: First, relying solely on global averaging or a single aggregation rule fails to distinguish the distribution differences between different clients, leading to negative transfer in heterogeneous scenarios; Second, most personalized methods employ fixed thresholds, fixed mask ratios, or manual empirical rules, lacking dynamic decision-making capabilities based on actual performance differences; Third, log anomaly detection models often have a large number of parameters or lack temporal inductive bias, which increases communication and training costs in federated environments; Fourth, existing solutions often focus on local algorithmic aspects, lacking complete process mechanisms and system device design, which is not conducive to engineering deployment. Summary of the Invention
[0006] This invention aims to provide a federated dual-adaptive personalized log anomaly detection system and method to solve the above problems.
[0007] First aspect
[0008] The technical solution of this invention is: a federated dual-adaptive personalized log anomaly detection system, comprising the following steps:
[0009] The data access and preprocessing module is used to convert raw log data into client-level sample data suitable for federated training, and to complete log parsing, sequence construction, label generation, client partitioning, and training / validation / test set partitioning.
[0010] The client-side local training module is used to train the anomaly detection model on the local log data of each client, enabling the model to learn the log patterns and anomaly features specific to that client.
[0011] The adaptive personalized decision-making module is used to dynamically determine which parameters should be shared globally and which parameters should be retained as local personalized parameters on the client, based on the differences between the client's local model and the global model.
[0012] The adaptive aggregation module is used to dynamically adjust the server-side aggregation weights based on the training quality, performance improvement, data scale, class distribution, and training stability of different clients, so that the global model can absorb more updates from high-quality clients.
[0013] The global model update module is used to perform weighted fusion of the model updates uploaded by the client based on the client weights generated by the adaptive aggregation module, to obtain a new round of global model, and then distribute it to each client.
[0014] The detection output module is used to apply the trained client-side personalized model to the local test log sequence, output the anomaly detection results, and calculate the final experimental evaluation index.
[0015] Second aspect
[0016] A federated dual-adaptive personalized log anomaly detection method, characterized in that it specifically includes:
[0017] S1. Log data preprocessing is used to convert raw unstructured log data into structured sequence samples that can be used to train the federated log anomaly detection model, and further construct a client dataset that meets the needs of federated learning experiments.
[0018] S2, Client-side Local Training and Metric Calculation, is used to train local anomaly detection models on private log data of each client and calculate the detection performance of the local model on the validation or test set, providing a basis for subsequent adaptive personalized decision-making and server-side dynamic adaptive aggregation; S1 is responsible for converting the original unstructured logs into structured log event sequences and constructing training, validation, and test sets for multiple clients according to the requirements of federated learning experiments; S2 uses the above preprocessing results as input to optimize the local anomaly detection model on the private training set of each client and calculates the detection performance metrics on the validation and test sets;
[0019] S3, TOPSIS-based client-side personalized pruning rate decision, is used to dynamically calculate the personalized pruning rate of each client in the current federation round based on the performance difference between the client's local model and the global model, thereby determining how many local personalized parameters the client should retain and how many global parameters should be shared. Among them, the log data preprocessing in S1 provides the data foundation for S3's TOPSIS-based client-side personalized pruning rate decision, and the client local training and metric calculation module in S2 provides direct performance input for S3's TOPSIS personalized pruning rate decision.
[0020] S4, parameter-level mask generation and update, is used to dynamically generate and update client-specific binary parameter masks based on the client-specific pruning rate obtained in S3 and the parameter change range of the client's local model. This allows for the division of globally shared parameters and locally personalized parameters at the parameter level. The S4 parameter-level mask generation and update module is used to further implement the client-specific pruning rate obtained in S3 at the level of specific model parameters, thereby dynamically dividing globally shared parameters and locally personalized parameters.
[0021] S5, server-adaptive aggregation, dynamically calculates client aggregation weights based on information such as model updates, performance improvements, data scale, category distribution, and training stability uploaded by each client. It then performs weighted fusion of globally shared parameters on the server side to generate a new global model more suitable for the overall federated log anomaly detection task. S5 is the aggregation and utilization stage of the information from the preceding steps: S1 provides the foundation for client data distribution, S2 provides local training and performance evaluation results, S3 provides decisions on the degree of personalization, and S4 provides parameter-level sharing range constraints. Based on this, S5 further calculates client-adaptive aggregation weights and performs weighted fusion of globally shared parameters to generate a new global model. Through this connection, FedDyna achieves synergy between client-side personalized parameter retention and server-side high-quality knowledge absorption.
[0022] S6, the aggregation weight calculation, is used to dynamically calculate the contribution weight of each client in server aggregation based on the data scale, category distribution, local model performance improvement, training stability, and parameter sharing status of each client. This allows the global model to absorb more high-quality, stable client updates that effectively contribute to anomaly detection. S6 calculates the dynamic weight of each client in server aggregation based on factors such as the number of client samples, anomaly ratio, local model performance improvement, and training stability. S1 provides data scale and category distribution information, S2 provides local training metrics and performance improvement information, S4 provides parameter sharing range or mask information, and S5 uses the aggregation weights output by S6 to complete the global model weighted aggregation.
[0023] S7, the log anomaly detection model, is a lightweight TCN model used for time-series feature extraction and anomaly classification of preprocessed log event sequences. It determines whether the input log sequence is normal or abnormal and provides a unified model carrier for FedDyna's local training, metric calculation, parameter pruning, mask update, and global aggregation. S7 not only handles anomaly detection but also serves as a unified model carrier for client-side personalized pruning, parameter mask generation, and server-side aggregation. The S2 client-side local training and metric calculation module trains and evaluates the S7 model on each client's private data and outputs model performance metrics and parameter update information. The S3 TOPSIS-based client-side personalized pruning rate decision module dynamically determines the proportion of personalized parameters retained by the client based on multiple performance differences between the local and global S7 models on the validation set. The S4 parameter-level mask generation and update module further operates on the specific model parameters of S7, generating a binary mask based on the parameter change magnitude to divide the S7 parameters into globally shared parameters and locally personalized parameters. The S6 aggregation weight calculation module calculates the weights based on the S7 model's performance. The model's performance improvement on each client, training stability, and client data distribution information are used to calculate dynamic aggregation weights. The S5 server adaptive aggregation module combines the aggregation weights output by S6 and the parameter mask generated by S4 to perform weighted fusion of the globally shared parameters in the S7 model of each client, generating a new round of global S7 model. Thus, S7 runs through the entire training process of FedDyna, serving as both the execution model for log anomaly detection and the carrier of parameter-level personalization and federated aggregation mechanisms.
[0024] S8, Result Output and Storage, is used to uniformly record, save, and output the FedDyna federated training process, client detection results, global model update results, personalized mask status, aggregate weight changes, and final experimental metrics. This provides data support for model evaluation, experimental reproduction, ablation analysis, and subsequent result presentation. S8 Result Output and Storage saves the model, metrics, masks, weights, communication overhead, and prediction results generated during FedDyna training and detection. S1 provides data preprocessing configuration and client statistics; S2 provides training logs and performance metrics; S3 provides TOPSIS scores and personalized pruning rates; S4 provides parameter masks and parameter splitting status; S6 provides aggregate weights and their constituent factors; S5 provides global model updates and communication information; and S7 provides final anomaly detection prediction results and model parameters. S8 uniformly saves these results for performance comparison, ablation analysis, and result presentation.
[0025] Preferably, step S1 specifically includes:
[0026] S11. Raw log reading and format unification: This step is used to read the raw log content from different log datasets such as HDFS, BGL, and Thunderbird, and to unify and organize the log fields.
[0027] S12. Perform template parsing on the raw log. This step is used to convert unstructured log text into structured log templates or event numbers.
[0028] S13. Encode log events by representing them as fixed-length vectors using word vectors or embedding models, and further map log templates into numerical representations that can be processed by the model.
[0029] S14. Log sequence construction: This step is used to organize single log events into a log event sequence, forming the basic input samples for the anomaly detection model. Different datasets can use different sequence construction methods. HDFS uses blocks as the basic unit to construct log sequences. BGL and Thunderbird usually do not have natural session identifiers similar to HDFS blocks, so a sliding time window can be used to construct log sequences.
[0030] S15, Label Generation and Alignment: This step is used to assign normal or abnormal labels to each log sequence. HDFS determines whether the block sequence is abnormal based on the block label. BGL generates window labels based on whether the window contains abnormal log lines. Thunderbird generates window labels based on log line-level abnormal annotations or key abnormal identifiers.
[0031] S16. Divide the data into multiple clients according to the actual scenario, and further divide them into training set, validation set and test set.
[0032] In S16, the actual scenarios include same-domain scenarios and cross-domain scenarios. For same-domain scenarios, multiple clients are selected from the same dataset. For cross-domain scenarios, clients are selected from different datasets to construct strong heterogeneity.
[0033] Preferably, step S2 specifically includes:
[0034] S21. Each client receives the global model initialization parameters broadcast by the server, performs several rounds of training on local data, and obtains the locally updated model parameters.
[0035] S22. After training locally, each client evaluates its local model and the previous round's global model using its local validation or test set, obtaining performance metrics and the performance gap between the local model and the previous round's global model. Performance metrics include: F1 score, accuracy, recall, precision, and PRAUC difference. The performance gap between the local model and the previous round's global model includes one or more of the following: ΔF1, accuracy gap, recall gap, and precision gap. Furthermore, each client also calculates the local sample size, anomaly ratio, and historical F1 fluctuations for subsequent personalized decision-making and aggregation weight calculation.
[0036] Preferably, step S3 specifically includes:
[0037] S31. In each federated round, after each client completes local training, the system collects the local evaluation metrics of each client and combines them with the corresponding evaluation metrics of the global model to construct a pruning gap matrix with the client as the evaluation object. The metrics of the pruning gap matrix include F1 gap, accuracy gap, precision gap and recall gap. Each gap metric is obtained by subtracting the corresponding metric of the global model from the corresponding metric of the client's local model. When the gap value is less than 0, it is set to 0 to indicate that the client does not have a performance advantage over the global model in that metric.
[0038] S32. The TOPSIS method is used to comprehensively evaluate each client and obtain the relative proximity score of each client in the pruning gap decision space. The relative proximity score reflects the degree of proximity of the client to the positive ideal gap vector, rather than the direct proximity between the local model and the global model. This is responsible for quantifying the local model advantages of each client into a pruning gap matrix. S32. Based on the pruning gap matrix in S31, TOPSIS is used to calculate the relative proximity score of each client.
[0039] S33. Map the relative proximity score to a preset pruning rate range to obtain the personalized pruning rate for each client; where the higher the relative proximity score, the smaller the pruning rate obtained by mapping.
[0040] Preferably, step S4 specifically includes:
[0041] S41. Sort the client parameters according to the magnitude of parameter changes;
[0042] S42. Select the corresponding number of parameters as local retained parameters according to the pruning rate p% or retention ratio (1-p%) determined in S3, and the remaining parameters as global shared parameters. Generate a parameter-level binary mask for each client model parameter; where the parameter with a mask value of 1 represents the local retained parameter, and the parameter with a mask value of 0 represents the global shared parameter, or use the equivalent reverse encoding method.
[0043] Preferably, step S5 specifically includes:
[0044] The aggregation range is adaptively selected based on the scenario. For weakly heterogeneous scenarios, where client logs are similarly distributed, a full-parameter aggregation mode is used, ignoring the mask and directly weighting all parameters according to client weights to maximize shared knowledge. For strongly heterogeneous scenarios, where clients differ greatly, a shared-parameter aggregation mode is used, normalizing and aggregating only parameters with a mask of 0 to reduce negative migration. For moderately heterogeneous scenarios, a hybrid mode is used, allowing local parameters with a mask of 1 to participate in aggregation with small weights, thus achieving a trade-off between global consistency and personalized preservation. Dividing clients within the same dataset corresponds to weakly heterogeneous scenarios; dividing clients with only two datasets corresponds to moderately heterogeneous scenarios; and dividing clients with three or more datasets corresponds to strongly heterogeneous scenarios.
[0045] Preferably, step S6 specifically includes:
[0046] S61. Before performing parameter aggregation, the server constructs an aggregation indicator matrix based on multiple aggregation evaluation indicators uploaded by each client. Among them, the aggregation evaluation indicators include: ΔF1, logarithmic transformation value of data volume, class balance, and stability score calculated based on historical F1 fluctuations.
[0047] S62. The server uses the entropy weight method to objectively assign weights to the indicators in S61, obtaining the weights of each indicator and the comprehensive score of each client. The comprehensive score can be further converted into normalized aggregate weights through the Softmax function for subsequent weighted aggregation of global parameters.
[0048] Preferably, step S7 specifically includes:
[0049] The lightweight TCN model consists of an input layer, a feature compression layer, a temporal convolutional backbone, a global pooling layer, a Dropout layer, and a fully connected classification layer. The input layer receives log samples with a shape equal to the sequence length multiplied by the embedding dimension. The feature compression layer reduces the high-dimensional embedding to a lower dimension. The temporal convolutional backbone includes multiple residual temporal convolutional blocks, each with a different dilation rate to expand the receptive field and extract multi-scale temporal features. Global average pooling and Dropout layers are used to obtain robust representations. The fully connected layer outputs the normal / abnormal classification result.
[0050] Preferably, step S8 specifically includes:
[0051] After reaching the preset number of federated rounds, the system outputs the final global model, personalized models for each client, corresponding parameter masks, round evaluation history, and analysis reports; thus, it can simultaneously support unified anomaly detection for the platform side and local optimal detection for the client side.
[0052] The beneficial effects of this invention are as follows:
[0053] 1. This invention can directly adjust the intensity of personalization based on the client's local benefits, enabling different clients to obtain differentiated parameter sharing boundaries. Therefore, it is more suitable for handling the non-independent identically distributed problem commonly found in log scenarios. Especially under the condition of mixed cross-domain clients, this invention can effectively reduce the negative migration caused by unrelated distributions through aggregation of shared parameters only or mixed aggregation.
[0054] 2. This invention employs a lightweight temporal convolutional model as the backbone for log detection. Compared to deep models with a larger number of parameters, it is easier to deploy in a federated environment and achieves a balance between communication efficiency, training stability, and detection performance. Therefore, this invention not only improves client-side personalization but also considers global model performance and engineering feasibility. Attached Figure Description
[0055] Figure 1 This is a diagram illustrating the overall architecture of federated dual-adaptive personalized log anomaly detection provided in an embodiment of the present invention.
[0056] Figure 2 A flowchart of the dual adaptive mechanism provided in an embodiment of the present invention;
[0057] Figure 3 This is a structural diagram of the lightweight temporal convolutional log detection module provided in an embodiment of the present invention;
[0058] Figure 4 This is a schematic diagram of the system deployment and functional devices provided in an embodiment of the present invention. Detailed Implementation
[0059] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. The embodiments of the present invention are not limited thereto.
[0060] Example 1
[0061] This invention provides a federated dual-adaptive personalized log anomaly detection system, including a data access and preprocessing module, a client-side local training module, an adaptive personalized decision-making module, an adaptive aggregation module, a global model update module, and a detection output module. During execution, the log data from each client is first uniformly preprocessed. Then, within each federated round, client-side local training, client-side metric calculation, personalized ratio decision-making, parameter-level mask generation, server-side weighted aggregation, and global update are executed sequentially until the final global model and personalized models for each client are obtained. Specifically, each client completes model training and performance metric calculation locally. Subsequently, a decision matrix is constructed based on multi-metric information from all clients. The TOPSIS method is used to determine the degree of personalization for each client and map it to the client parameter retention ratio. A parameter-level binary mask is then generated based on the parameter change magnitude. The server further performs weighted aggregation on the client parameters according to the aggregation weights.
[0062] like Figure 1 As shown, this invention employs a typical client-server federated learning architecture. The raw log data comes from datasets from different sources such as HDFS, BGL, and Thunderbird. After log parsing, event representation learning, sequence pruning / padding, and client-side partitioning, it is fed into each client's local model for training. The client output includes local model parameters, local evaluation metrics, and the performance gap with the previous round's global model. Upon receiving this information, the server generates parameter masks for each client through a personalized decision layer, calculates aggregation weights through an adaptive aggregation layer, updates the global model, and then broadcasts the updated global model to the next round.
[0063] Step 1: Log Data Preprocessing. The raw logs are parsed using templates to obtain event sequences; log events are represented as fixed-length vectors using word vectors or embedding models; sample sequences are standardized to a preset length, such as 50; subsequently, the data is divided into multiple clients based on the actual scenario, and further divided into training, validation, and test sets. For same-domain scenarios, multiple clients can be selected from the same dataset; for cross-domain scenarios, clients can be selected from different datasets to construct stronger heterogeneity.
[0064] Step Two: Local Training and Metric Calculation on Clients. Each client receives the global model initialization parameters broadcast by the server and performs several rounds of training on local data to obtain locally updated model parameters. After local training, each client evaluates its local model and the previous round's global model using the same local validation or test set, obtaining performance metrics such as F1 score, accuracy, recall, and precision. Furthermore, it calculates the performance gap between the local model and the previous round's global model, including one or more of ΔF1, accuracy gap, recall gap, and precision gap. In addition, each client also tracks the local sample size, outlier ratio, and historical F1 fluctuations for subsequent personalized decision-making and aggregation weight calculation.
[0065] Step 3: Client-specific pruning rate decision based on TOPSIS. See also Figure 2 In each federated round, after each client completes local training, the system collects the local evaluation metrics of each client and combines them with the corresponding evaluation metrics of the global model to construct a pruning gap matrix with the client as the evaluation object. The metrics in the pruning gap matrix include F1 gap, accuracy gap, precision gap, and recall gap, where each gap metric is obtained by subtracting the corresponding metric of the global model from the corresponding metric of the client's local model. When the gap value is less than 0, it is set to 0 to indicate that the client does not have a performance advantage over the global model on that metric. Subsequently, the TOPSIS method is used to comprehensively evaluate each client, obtaining the relative proximity score of each client in the pruning gap decision space. The relative proximity score reflects the degree of proximity of the client to the positive ideal gap vector, rather than the direct proximity between the local model and the global model. The relative proximity score is then mapped to a preset pruning rate range to obtain the personalized pruning rate corresponding to each client; where the higher the relative proximity score, the smaller the mapped pruning rate.
[0066] Step 4: Parameter-level Mask Generation and Update. After obtaining the personalized pruning rate or parameter retention ratio for each client, a parameter-level binary mask is generated for the model parameters of each client. First, the client parameters are sorted according to the parameter change magnitude or other parameter importance indicators. Then, according to the pruning rate or retention ratio determined in Step 3, a corresponding number of parameters are selected as locally retained parameters, and the remaining parameters are used as globally shared parameters, thus forming a parameter-level binary mask. Parameters with a mask value of 1 represent locally retained parameters, and parameters with a mask value of 0 represent globally shared parameters, or an equivalent reverse encoding method can be used.
[0067] Step 5: Server-Adaptive Aggregation. Unlike existing methods that only use a single aggregation mode, this invention adaptively selects the aggregation scope based on the scenario. For weakly heterogeneous scenarios, where client logs are similarly distributed, a full-parameter aggregation mode can be used, ignoring the mask and directly weighting all parameters according to client weights to maximize shared knowledge. For strongly heterogeneous scenarios, where client differences are significant, a shared-parameter aggregation mode can be used, normalizing and aggregating only parameters with a mask of 0 to reduce negative migration. Furthermore, a hybrid mode can be used, allowing local parameters with a mask of 1 to participate in aggregation with a smaller weight, thus achieving a trade-off between global consistency and personalized preservation.
[0068] Step Six: Aggregate Weight Calculation. Before parameter aggregation, the server constructs an aggregation index matrix based on multiple aggregation evaluation indicators uploaded by each client. The preferred aggregation evaluation indicators include: ΔF1, the logarithmic transformation value of the data volume, class balance, and a stability score calculated based on historical F1 fluctuations. The server uses the entropy weight method to objectively assign weights to the above indicators, obtaining the weights of each indicator and the comprehensive score of each client. The comprehensive score can then be further converted into normalized aggregation weights using the Softmax function for subsequent weighted aggregation of global parameters.
[0069] Step 7: Log Anomaly Detection Model. See also... Figure 3 The lightweight TCN model comprises an input layer, an optional feature compression layer, a temporal convolutional backbone, a global pooling layer, a Dropout layer, and a fully connected classification layer. The input layer receives log samples with a shape equal to the sequence length multiplied by the embedding dimension. The feature compression layer reduces the high-dimensional embedding to a lower dimension. The backbone is preferably composed of multiple residual temporal convolutional blocks, each with a different dilation rate to expand the receptive field and extract multi-scale temporal features. Robust representations are then obtained through global average pooling and Dropout, and finally, a fully connected layer outputs the normal / abnormal classification result. Compared to heavier deep models, this structure has fewer parameters and is more suitable for frequent communication and parallel training with multiple clients.
[0070] Step 8: Result Output and Storage. After reaching the preset number of federated rounds, the system outputs the final global model, the personalized models for each client, the corresponding parameter masks, the round evaluation history, and the analysis report. This enables simultaneous support for unified anomaly detection on the platform side and local optimal detection on the client side.
[0071] like Figure 4As shown, from a device perspective, this invention can also be implemented as a federated log anomaly detection system. This system includes: a data access device for completing log collection, template parsing, and vectorization preprocessing; a client training device for performing lightweight log model training on local data; a personalization control device for generating masks based on client performance differences; an aggregation control device for selecting the aggregation range based on scenarios and weight rules; a server coordination device for updating the global model and broadcasting parameters; and a detection output device for outputting anomaly alarms, performance reports, and model files. Each device can be deployed on the same server or distributed across multiple hosts or cloud-edge nodes.
[0072] It should be noted that the order of the above steps is not absolutely fixed. Without departing from the core idea of this invention, the order of mask update and weight calculation, the selection method of aggregation range, and the local structure of the log model can all be adjusted according to actual engineering conditions.
[0073] Example 2
[0074] In addition to TOPSIS, personalized decision-making can also employ analytic hierarchy process (AHP), fuzzy comprehensive evaluation method, Bayesian decision method, or policy network based on reinforcement learning to determine the mask ratio.
[0075] In addition to using the three modes of all, masked_only, and hybrid, the aggregation control part can also use hierarchical aggregation, clustering-based intra-group aggregation, gated weighted aggregation, or meta-learning aggregation to complete global parameter updates.
[0076] In addition to using the LightLog lightweight temporal convolutional model, the log anomaly detection model can also be replaced with a recurrent neural network, Transformer, graph neural network or its lightweight variants, as long as it still meets the requirement of local training in a federated environment and is combined with the dual adaptive decision-making mechanism of this invention.
[0077] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A federated dual-adaptive personalized log anomaly detection system, characterized in that, Includes the following steps: The data access and preprocessing module is used to convert raw log data into client-level sample data suitable for federated training, and to complete log parsing, sequence construction, label generation, client partitioning, and training / validation / test set partitioning. The client-side local training module is used to train the anomaly detection model on the local log data of each client, enabling the model to learn the log patterns and anomaly features specific to that client. The adaptive personalized decision-making module is used to dynamically determine which parameters should be shared globally and which parameters should be retained as local personalized parameters on the client, based on the differences between the client's local model and the global model. The adaptive aggregation module is used to dynamically adjust the server-side aggregation weights based on the training quality, performance improvement, data scale, class distribution, and training stability of different clients, so that the global model can absorb more updates from high-quality clients. The global model update module is used to perform weighted fusion of the model updates uploaded by the client based on the client weights generated by the adaptive aggregation module, to obtain a new round of global model, and then distribute it to each client. The detection output module is used to apply the trained client-side personalized model to the local test log sequence, output the anomaly detection results, and calculate the final experimental evaluation index.
2. A federated dual-adaptive personalized log anomaly detection method, characterized in that, Specifically, it includes: S1. Log data preprocessing is used to convert raw unstructured log data into structured sequence samples that can be used to train the federated log anomaly detection model, and further construct a client dataset that meets the needs of federated learning experiments. S2, Client-side Local Training and Metric Calculation, is used to train local anomaly detection models on private log data of each client and calculate the detection performance of the local model on the validation or test set, providing a basis for subsequent adaptive personalized decision-making and server-side dynamic adaptive aggregation; S1 is responsible for converting the original unstructured logs into structured log event sequences and constructing training, validation, and test sets for multiple clients according to the requirements of federated learning experiments; S2 uses the above preprocessing results as input to optimize the local anomaly detection model on the private training set of each client and calculates the detection performance metrics on the validation and test sets; S3, TOPSIS-based client-side personalized pruning rate decision, is used to dynamically calculate the personalized pruning rate of each client in the current federation round based on the performance difference between the client's local model and the global model, thereby determining how many local personalized parameters the client should retain and how many global parameters should be shared. Among them, the log data preprocessing in S1 provides the data foundation for S3 TOPSIS-based client-side personalized pruning rate decision, and the client local training and metric calculation module in S2 provides direct performance input for S3 TOPSIS personalized pruning rate decision. S4, parameter-level mask generation and update, is used to dynamically generate and update client-specific binary parameter masks based on the client-specific pruning rate obtained in S3 and the parameter change range of the client's local model. This allows for the division of globally shared parameters and locally personalized parameters at the parameter level. The S4 parameter-level mask generation and update module is used to further implement the client-specific pruning rate obtained in S3 at the level of specific model parameters, thereby dynamically dividing globally shared parameters and locally personalized parameters. S5, server-adaptive aggregation, dynamically calculates client aggregation weights based on information such as model updates, performance improvements, data scale, category distribution, and training stability uploaded by each client. It then performs weighted fusion of globally shared parameters on the server side to generate a new global model more suitable for the overall federated log anomaly detection task. S5 is the aggregation and utilization stage of the information from the preceding steps: S1 provides the foundation for client data distribution, S2 provides local training and performance evaluation results, S3 provides decisions on the degree of personalization, and S4 provides parameter-level sharing range constraints. Based on this, S5 further calculates client-adaptive aggregation weights and performs weighted fusion of globally shared parameters to generate a new global model. Through this connection, FedDyna achieves synergy between client-side personalized parameter retention and server-side high-quality knowledge absorption. S6, the aggregation weight calculation, is used to dynamically calculate the contribution weight of each client in server aggregation based on the data scale, category distribution, local model performance improvement, training stability, and parameter sharing status of each client. This allows the global model to absorb more high-quality, stable client updates that effectively contribute to anomaly detection. S6 calculates the dynamic weight of each client in server aggregation based on factors such as the number of client samples, anomaly ratio, local model performance improvement, and training stability. S1 provides data scale and category distribution information, S2 provides local training metrics and performance improvement information, S4 provides parameter sharing range or mask information, and S5 uses the aggregation weights output by S6 to complete the global model weighted aggregation. S7, the log anomaly detection model, is a lightweight TCN model used for time-series feature extraction and anomaly classification of preprocessed log event sequences. It determines whether the input log sequence is normal or abnormal, and provides a unified model carrier for FedDyna's local training, metric calculation, parameter pruning, mask update, and global aggregation. S7 not only handles anomaly detection but also serves as a unified model carrier for client-side personalized pruning, parameter mask generation, and server-side aggregation. The S2 client-side local training and metric calculation module trains and evaluates the S7 model on each client's private data and outputs model performance metrics and parameter update information. The S3 TOPSIS-based client-side personalized pruning rate decision module dynamically determines the proportion of personalized parameters retained by the client based on multiple performance differences between the local and global S7 models on the validation set. The S4 parameter-level mask generation and update module further operates on the specific model parameters of S7, generating a binary mask based on the parameter change magnitude, dividing the S7 parameters into globally shared parameters and locally personalized parameters. The S6 aggregation weight calculation module calculates the weights based on the S7 model's performance. The model's performance improvement on each client, training stability, and client data distribution information are used to calculate dynamic aggregation weights. The S5 server adaptive aggregation module combines the aggregation weights output by S6 and the parameter mask generated by S4 to perform weighted fusion of the globally shared parameters in the S7 model of each client, generating a new round of global S7 model. Thus, S7 runs through the entire training process of FedDyna, serving as both the execution model for log anomaly detection and the carrier of parameter-level personalization and federated aggregation mechanisms. S8, Result Output and Storage, is used to uniformly record, save, and output the FedDyna federated training process, client detection results, global model update results, personalized mask status, aggregate weight changes, and final experimental metrics. This provides data support for model evaluation, experimental reproduction, ablation analysis, and subsequent result display. S8 Result Output and Storage saves the model, metrics, masks, weights, communication overhead, and prediction results generated during FedDyna training and detection. S1 provides data preprocessing configuration and client statistics; S2 provides training logs and performance metrics; S3 provides TOPSIS scores and personalized pruning rates; S4 provides parameter masks and parameter splitting status; S6 provides aggregate weights and their constituent factors; S5 provides global model updates and communication information; and S7 provides final anomaly detection prediction results and model parameters. S8 uniformly saves these results for performance comparison, ablation analysis, and result display.
3. The federated dual-adaptive personalized log anomaly detection method according to claim 2, characterized in that, Step S1 specifically includes: S11. Raw log reading and format unification: This step is used to read the raw log content from different log datasets such as HDFS, BGL, and Thunderbird, and to unify and organize the log fields. S12. Perform template parsing on the raw log. This step is used to convert unstructured log text into structured log templates or event numbers. S13. Encode log events by representing them as fixed-length vectors using word vectors or embedding models, and further map log templates into numerical representations that can be processed by the model. S14. Log sequence construction: This step is used to organize single log events into a log event sequence, forming the basic input samples for the anomaly detection model. Different datasets can use different sequence construction methods. HDFS uses blocks as the basic unit to construct log sequences. BGL and Thunderbird usually do not have natural session identifiers similar to HDFS blocks, so a sliding time window can be used to construct log sequences. S15, Label Generation and Alignment: This step is used to assign normal or abnormal labels to each log sequence. HDFS determines whether the block sequence is abnormal based on the block label. BGL generates window labels based on whether the window contains abnormal log lines. Thunderbird generates window labels based on log line-level abnormal annotations or key abnormal identifiers. S16. Divide the data into multiple clients according to the actual scenario, and further divide them into training set, validation set and test set. In S16, the actual scenarios include same-domain scenarios and cross-domain scenarios. For same-domain scenarios, multiple clients are selected from the same dataset. For cross-domain scenarios, clients are selected from different datasets to construct strong heterogeneity.
4. The federated dual-adaptive personalized log anomaly detection method according to claim 2, characterized in that, Step S2 specifically includes: S21. Each client receives the global model initialization parameters broadcast by the server, performs several rounds of training on local data, and obtains the locally updated model parameters. S22. After training locally, each client evaluates its local model and the previous round's global model using its local validation or test set, obtaining performance metrics and the performance gap between the local model and the previous round's global model. Performance metrics include: F1 score, accuracy, recall, precision, and PRAUC difference. The performance gap between the local model and the previous round's global model includes one or more of the following: ΔF1, accuracy gap, recall gap, and precision gap. Furthermore, each client also calculates the local sample size, anomaly ratio, and historical F1 fluctuations for subsequent personalized decision-making and aggregation weight calculation.
5. The federated dual-adaptive personalized log anomaly detection method according to claim 2, characterized in that, Step S3 specifically includes: S31. In each federated round, after each client completes local training, the system collects the local evaluation metrics of each client and combines them with the corresponding evaluation metrics of the global model to construct a pruning gap matrix with the client as the evaluation object. The metrics of the pruning gap matrix include F1 gap, accuracy gap, precision gap and recall gap. Each gap metric is obtained by subtracting the corresponding metric of the global model from the corresponding metric of the client's local model. When the gap value is less than 0, it is set to 0 to indicate that the client does not have a performance advantage over the global model in that metric. S32. The TOPSIS method is used to comprehensively evaluate each client and obtain the relative proximity score of each client in the pruning gap decision space. The relative proximity score reflects the degree of proximity of the client to the positive ideal gap vector, rather than the direct proximity between the local model and the global model. This is responsible for quantifying the local model advantages of each client into a pruning gap matrix. S32. Based on the pruning gap matrix in S31, TOPSIS is used to calculate the relative proximity score of each client. S33. Map the relative proximity score to a preset pruning rate range to obtain the personalized pruning rate for each client; where the higher the relative proximity score, the smaller the pruning rate obtained by mapping.
6. The federated dual-adaptive personalized log anomaly detection method according to claim 2, characterized in that, Step S4 specifically includes: S41. Sort the client parameters according to the magnitude of parameter changes; S42. Select the corresponding number of parameters as local retained parameters according to the pruning rate p% or retention ratio (1-p%) determined in S3, and the remaining parameters as global shared parameters. Generate a parameter-level binary mask for each client model parameter; where the parameter with a mask value of 1 represents the local retained parameter, and the parameter with a mask value of 0 represents the global shared parameter, or use the equivalent reverse encoding method.
7. The federated dual-adaptive personalized log anomaly detection method according to claim 2, characterized in that, Step S5 specifically includes: The aggregation range is adaptively selected based on the scenario. For weakly heterogeneous scenarios, where client logs are similarly distributed, a full-parameter aggregation mode is used, ignoring the mask and directly weighting all parameters according to client weights to maximize shared knowledge. For strongly heterogeneous scenarios, where clients differ greatly, a shared-parameter aggregation mode is used, normalizing and aggregating only parameters with a mask of 0 to reduce negative migration. For moderately heterogeneous scenarios, a hybrid mode is used, allowing local parameters with a mask of 1 to participate in aggregation with small weights, thus achieving a trade-off between global consistency and personalized preservation. Dividing clients within the same dataset corresponds to weakly heterogeneous scenarios; dividing clients with only two datasets corresponds to moderately heterogeneous scenarios; and dividing clients with three or more datasets corresponds to strongly heterogeneous scenarios.
8. The federated dual-adaptive personalized log anomaly detection method according to claim 2, characterized in that, Step S6 specifically includes: S61. Before performing parameter aggregation, the server constructs an aggregation indicator matrix based on multiple aggregation evaluation indicators uploaded by each client. Among them, the aggregation evaluation indicators include: ΔF1, logarithmic transformation value of data volume, class balance, and stability score calculated based on historical F1 fluctuations. S62. The server uses the entropy weight method to objectively assign weights to the indicators in S61, obtaining the weights of each indicator and the comprehensive score of each client. The comprehensive score can be further converted into normalized aggregate weights through the Softmax function for subsequent weighted aggregation of global parameters.
9. The federated dual-adaptive personalized log anomaly detection method according to claim 2, characterized in that, Step S7 specifically includes: The lightweight TCN model consists of an input layer, a feature compression layer, a temporal convolutional backbone, a global pooling layer, a Dropout layer, and a fully connected classification layer. The input layer receives log samples with a shape equal to the sequence length multiplied by the embedding dimension. The feature compression layer reduces the high-dimensional embedding to a lower dimension. The temporal convolutional backbone includes multiple residual temporal convolutional blocks, each with a different dilation rate to expand the receptive field and extract multi-scale temporal features. Global average pooling and Dropout layers are used to obtain robust representations. The fully connected layer outputs the normal / abnormal classification result.
10. The federated dual-adaptive personalized log anomaly detection method according to claim 2, characterized in that, Step S8 specifically includes: After reaching the preset number of federated rounds, the system outputs the final global model, personalized models for each client, corresponding parameter masks, round evaluation history, and analysis reports; thus, it can simultaneously support unified anomaly detection for the platform side and local optimal detection for the client side.