Cross-domain federated learning system and cross-domain DDoS attack detection method and device

CN121309058BActive Publication Date: 2026-09-29NORTHEASTERN UNIV CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511396451.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2026-09-29
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

现有联邦学习相关方法,无论是基于客户端模型聚合的方案,还是引入服务端蒸馏的方案,均难以充分提取并融合多域中稀缺的攻击类知识,无法高效提升模型对尾部攻击类别的识别性能

Benefits of technology

[0021]借由上述技术方案,本申请提供的一种跨域联邦学习系统及跨域DDoS攻击检测方法及装置,通过构建包含中央协调节点与多域节点的跨域联邦学习系统,结合双输出本地模型与公共数据集,在保护数据隐私的基础上,通过多轮联邦训练实现跨域知识协同与少数类攻击样本建模;各域节点基于本地私有流量数据训练双输出模型并上传参数,中央协调节点利用公共数据集推理生成逻辑值,既充分利用了各域本地数据,又避免了原始数据跨域传输,能够有效缓解现有方法未适配跨域场景的问题;中央协调节点通过样本量加权平均聚合域节点参数,可确保数据贡献的合理性,同时基于预测置信度融合多域逻辑值,可强化攻击识别能力强的域节点知识贡献,避免低质量知识掩盖稀缺攻击类信息,可以解决现有方法难以充分聚合多域攻击类知识的痛点;通过在公共数据集中筛选潜在攻击样本,结合融合逻辑值对初步跨域联合模型执行知识蒸馏,聚焦少数类攻击样本强化模型学习,可针对性缓解域内类别失衡与尾部攻击样本稀缺导致的识别性能不足问题;最后通过多轮联邦训练迭代,持续优化模型泛化能力与鲁棒性,最终获得的全局共享模型能让各域节点高效执行DDoS攻击检测,既能够兼顾数据隐私保护,又可以显著增强模型对少数类攻击样本的感知与识别能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121309058B_ABST
    Figure CN121309058B_ABST
Patent Text Reader

Abstract

The application discloses a cross-domain federated learning system and a cross-domain DDoS attack detection method and device, relates to the technical field of network security, and comprises a central coordination node and multiple domain nodes. The domain nodes maintain local models of a double-output structure, the central coordination node holds a public data set for distillation training, and the system adopts a multi-round federated training strategy. In each round of training, the domain nodes train local models based on local private traffic data sets, upload the updated parameters to the central coordination node, the central coordination node uses the public data set to infer the domain models to generate logical values, aggregates the parameters by weighting to obtain a preliminary cross-domain joint model, and simultaneously determines fusion logical values; the central coordination node screens a potential attack sample set, combines the fusion logical values to perform distillation training on the preliminary model, updates an optimized global model, and distributes the global model to the domain nodes for initialization of parameters of the next round of local models to perform multi-round collaborative optimization, so that a global shared model is obtained, and each domain node performs DDoS attack detection according to the global shared model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of cybersecurity technology, and in particular to a cross-domain federated learning system and a cross-domain DDoS attack detection method and apparatus. Background Technology

[0002] With the continuous expansion of network scale and the widespread deployment of Internet of Things (IoT) devices, network security faces severe challenges. Among these, Distributed Denial of Service (DDoS) attacks, due to their high frequency, covert methods, and strong destructiveness, have become a key threat to network security. To improve the intelligence level of DDoS attack detection, traditional methods often employ centralized machine learning architectures, requiring the uploading of large amounts of network traffic data to a central server for unified modeling. However, in real-world scenarios, DDoS attacks often originate across multiple network domains, with network traffic data distributed across nodes in each domain. Due to limitations imposed by data sovereignty, privacy protection, and network isolation requirements, it is difficult to aggregate the raw data to a central node, thus limiting the effectiveness of centralized methods and leading to a severe data silo problem in cross-domain DDoS detection. Federated Learning (FL) offers a feasible solution to this problem. It trains models on local nodes in each domain, and a central coordinating node aggregates model parameters or knowledge, achieving joint modeling without sharing raw data. However, due to significant differences in network environments, attack types, and collection strategies across different domains, local data exhibits non-independent identically distributed (Non-IID) characteristics, and attack traffic is naturally scarce, forming an extreme long-tail distribution. This makes it difficult for the global model constructed by federated learning to learn effective discrimination patterns of tail attack samples. Data heterogeneity and class imbalance have become key bottlenecks restricting the generalization ability and attack identification ability of federated DDoS detection models.

[0003] To alleviate these problems, existing technologies have been researched. For example, the FLDDoS method achieves joint attack identification through multi-client model aggregation and combines K-Means clustering and SMOTEENN sampling mechanisms to address the global class imbalance problem; the FLAD method employs an adaptive client scheduling strategy to improve the training efficiency of heterogeneous clients, mitigating the impact of data imbalance to some extent. Furthermore, methods such as FedDF, Fedic, and CDKT-FL introduce server-side distillation mechanisms, leveraging proxy datasets to achieve cross-client knowledge transfer, attempting to alleviate non-IID problems. However, FLDDoS and FLAD mainly address the problem of data heterogeneity between clients, failing to effectively solve the severe class imbalance and scarcity of tail attack samples within each client or domain. Furthermore, their performance evaluations are mostly based on single-domain or homogeneous environments, without fully considering the data distribution characteristics and training requirements in cross-domain scenarios. Although methods such as FedDF, Fedic, and CDKT-FL introduce distillation mechanisms, they are mainly applied to general tasks such as image recognition and text classification, without being specifically designed for the sparsity of DDoS attack samples and the differences in features between domains in network traffic scenarios, making it difficult to adapt to the special needs of cross-domain DDoS attack detection.

[0004] In practical tasks of cross-domain DDoS attack detection, attack samples constitute an extremely low percentage of the total attack and exhibit strong dispersion in their distribution. Furthermore, the data collection strategies differ across domains, further exacerbating the challenges of robustness and generalization during model training. Existing federated learning methods, whether based on client-side model aggregation or server-side distillation, struggle to fully extract and integrate scarce attack-type knowledge from multiple domains, failing to efficiently improve the model's performance in identifying tail-type attacks. Currently, the industry lacks a DDoS detection method suitable for cross-domain federated learning scenarios that balances data privacy protection with attack sample modeling requirements. How to effectively aggregate multi-domain node knowledge and enhance the model's ability to perceive minority attack samples without sharing original data has become a core technical problem urgently needing to be solved in the field of cross-domain federated DDoS attack detection. Summary of the Invention

[0005] In view of this, this application provides a cross-domain federated learning system and a cross-domain DDoS attack detection method and device, which can be applied to cross-domain federated learning scenarios, can take into account both data privacy protection and attack sample modeling needs, fully extract and integrate scarce attack-type knowledge in multiple domains, and effectively improve the model's recognition performance of tail attack categories.

[0006] According to a first aspect of this application, a cross-domain federated learning system is provided, comprising a central coordinating node and multiple domain nodes, each domain node maintaining a local model with a dual-output structure, the central coordinating node holding a public dataset for distillation training, and the cross-domain federated learning system employing a multi-round federated training strategy, wherein each round of federated training includes:

[0007] Each domain node trains its local model based on its local private traffic dataset, updates the local model parameters, and uploads it to the central coordination node. The central coordination node uses the public dataset to perform inference on the local models uploaded by each domain node, generating multiple logical values ​​for knowledge distillation.

[0008] The central coordination node aggregates the local model parameters of each domain node through a weighted average strategy to obtain a preliminary cross-domain joint model. At the same time, it calculates the prediction confidence of each logical value and uses the prediction confidence as a weight to fuse the multiple logical values ​​to generate a fused logical value.

[0009] The central coordination node filters a potential attack sample set from the public dataset, performs knowledge distillation training on the preliminary cross-domain joint model based on the potential attack sample set and the fusion logic value, and updates the optimized global model.

[0010] The central coordinating node distributes the optimized global model to each domain node. Each domain node uses the global model as the initialization parameter for its local model in the next round of federated training and executes each round of federated training in a loop. Through continuous collaborative optimization of multiple rounds of federated training, a global shared model is finally obtained. Each domain node performs DDoS attack detection processing based on the global shared model.

[0011] According to a second aspect of this application, a method for detecting cross-domain DDoS attacks is provided, the method being applied to the cross-domain federated learning system described in the first aspect, the method comprising:

[0012] Each domain node collects local network traffic in real time through network probes and extracts the feature vector of the local network traffic;

[0013] The feature vector is input into the locally configured global shared model to obtain the binary classification result output by the first output branch of the global shared model;

[0014] The DDoS attack detection result of the local network traffic is determined based on the binary classification result.

[0015] According to a third aspect of this application, a cross-domain DDoS attack detection device is provided, the device being applied to the cross-domain federated learning system described in the first aspect, the device comprising:

[0016] The extraction module is used to collect local network traffic in real time by each domain node through network probes, and extract the feature vector of the local network traffic.

[0017] The input module is used to input the feature vector into a locally configured global shared model and obtain the binary classification result output by the first output branch in the global shared model;

[0018] The determination module is used to determine the DDoS attack detection result of the local network traffic based on the binary classification result.

[0019] According to a fourth aspect of this application, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the above-described cross-domain DDoS attack detection method.

[0020] According to a fifth aspect of this application, an electronic device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described cross-domain DDoS attack detection method.

[0021] By employing the aforementioned technical solutions, this application provides a cross-domain federated learning system and a cross-domain DDoS attack detection method and apparatus. Through constructing a cross-domain federated learning system comprising a central coordinating node and multiple domain nodes, and combining a dual-output local model with a public dataset, it achieves cross-domain knowledge collaboration and minority class attack sample modeling through multiple rounds of federated training while protecting data privacy. Each domain node trains a dual-output model based on its local private traffic data and uploads the parameters. The central coordinating node uses the public dataset to infer and generate logical values, fully utilizing local data from each domain while avoiding cross-domain transmission of raw data, effectively alleviating the problem of existing methods not being adapted to cross-domain scenarios. The central coordinating node aggregates domain node parameters through a sample-weighted average, ensuring the rationality of data contributions, and simultaneously based on pre-... By fusing multi-domain logistic values ​​to measure confidence, the knowledge contribution of domain nodes with strong attack identification capabilities can be enhanced, avoiding the obscuring of scarce attack-type information by low-quality knowledge. This addresses the pain point of existing methods that struggle to fully aggregate multi-domain attack-type knowledge. By screening potential attack samples in public datasets and performing knowledge distillation on the initial cross-domain joint model using fused logistic values, the model learning is strengthened by focusing on minority attack samples. This can specifically alleviate the problem of insufficient identification performance caused by intra-domain class imbalance and the scarcity of tail attack samples. Finally, through multiple rounds of federated training iterations, the model's generalization ability and robustness are continuously optimized. The resulting globally shared model enables each domain node to efficiently perform DDoS attack detection, which can both protect data privacy and significantly enhance the model's ability to perceive and identify minority attack samples.

[0022] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description

[0023] Figure 1 This paper illustrates a schematic diagram of the system architecture of a cross-domain federated learning system provided in an embodiment of this application.

[0024] Figure 2 The diagram illustrates a flowchart of a cross-domain DDoS attack detection method provided in an embodiment of this application.

[0025] Figure 3 This application provides an example of an F1_Attack comparison curve.

[0026] Figure 4 This application provides a Recall_Attack comparison curve diagram according to an embodiment of the present application;

[0027] Figure 5This illustration shows a structural schematic diagram of a cross-domain DDoS attack detection device provided in an embodiment of this application;

[0028] In the picture:

[0029] 10 - Domain nodes, 20 - Central coordination nodes. Detailed Implementation

[0030] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.

[0031] With the continuous expansion of network scale and the widespread deployment of Internet of Things (IoT) devices, network security faces severe challenges. Among these, Distributed Denial of Service (DDoS) attacks, due to their high frequency, covert methods, and strong destructiveness, have become a key threat to network security. To improve the intelligence level of DDoS attack detection, traditional methods often employ centralized machine learning architectures, requiring the uploading of large amounts of network traffic data to a central server for unified modeling. However, in real-world scenarios, DDoS attacks often originate across multiple network domains, with network traffic data distributed across nodes in each domain. Due to limitations imposed by data sovereignty, privacy protection, and network isolation requirements, it is difficult to aggregate the raw data to a central node, thus limiting the effectiveness of centralized methods and leading to a severe data silo problem in cross-domain DDoS detection. Federated Learning (FL) offers a feasible solution to this problem. It trains models on local nodes in each domain, and a central coordinating node aggregates model parameters or knowledge, achieving joint modeling without sharing raw data. However, due to significant differences in network environments, attack types, and collection strategies across different domains, local data exhibits non-independent identically distributed (Non-IID) characteristics, and attack traffic is naturally scarce, forming an extreme long-tail distribution. This makes it difficult for the global model constructed by federated learning to learn effective discrimination patterns of tail attack samples. Data heterogeneity and class imbalance have become key bottlenecks restricting the generalization ability and attack identification ability of federated DDoS detection models.

[0032] To alleviate these problems, existing technologies have been researched. For example, the FLDDoS method achieves joint attack identification through multi-client model aggregation and combines K-Means clustering and SMOTEENN sampling mechanisms to address the global class imbalance problem; the FLAD method employs an adaptive client scheduling strategy to improve the training efficiency of heterogeneous clients, mitigating the impact of data imbalance to some extent. Furthermore, methods such as FedDF, Fedic, and CDKT-FL introduce server-side distillation mechanisms, leveraging proxy datasets to achieve cross-client knowledge transfer, attempting to alleviate non-IID problems. However, FLDDoS and FLAD mainly address the problem of data heterogeneity between clients, failing to effectively solve the severe class imbalance and scarcity of tail attack samples within each client or domain. Furthermore, their performance evaluations are mostly based on single-domain or homogeneous environments, without fully considering the data distribution characteristics and training requirements in cross-domain scenarios. Although methods such as FedDF, Fedic, and CDKT-FL introduce distillation mechanisms, they are mainly applied to general tasks such as image recognition and text classification, without being specifically designed for the sparsity of DDoS attack samples and the differences in features between domains in network traffic scenarios, making it difficult to adapt to the special needs of cross-domain DDoS attack detection.

[0033] In practical tasks of cross-domain DDoS attack detection, attack samples constitute an extremely low percentage of the total attack and exhibit strong dispersion in their distribution. Furthermore, the data collection strategies differ across domains, further exacerbating the challenges of robustness and generalization during model training. Existing federated learning methods, whether based on client-side model aggregation or server-side distillation, struggle to fully extract and integrate scarce attack-type knowledge from multiple domains, failing to efficiently improve the model's performance in identifying tail-type attacks. Currently, the industry lacks a DDoS detection method suitable for cross-domain federated learning scenarios that balances data privacy protection with attack sample modeling requirements. How to effectively aggregate multi-domain node knowledge and enhance the model's ability to perceive minority attack samples without sharing original data has become a core technical problem urgently needing to be solved in the field of cross-domain federated DDoS attack detection.

[0034] To address the aforementioned problems, embodiments of the present invention provide a cross-domain federated learning system, such as... Figure 1As shown, the cross-domain federated learning system comprises multiple domain nodes 10 (e.g., Domain1, Domain2, Domain3) and a central coordinating node 20. Each domain node 10 maintains a local model with a dual-output structure. The central coordinating node 20 holds a public dataset for distillation training. The cross-domain federated learning system employs a multi-round federated training strategy, and each round of federated training includes: each domain node 10 trains its local model based on its local private traffic dataset, updates its local model parameters, and uploads them to the central coordinating node 20; the central coordinating node 20 uses the public dataset to perform inference on the local models uploaded by each domain node 10, generating multiple logits for knowledge distillation; the central coordinating node 20 aggregates the local model parameters of each domain node 10 using a weighted averaging strategy to obtain a preliminary cross-domain joint model (FedAvg). The model simultaneously calculates the prediction confidence of each logical value and merges multiple logical values ​​with the prediction confidence as weights to generate a fused logical value. The central coordinating node 20 filters a potential attack sample set from the public dataset, performs knowledge distillation training on the preliminary cross-domain joint model based on the potential attack sample set and the fused logical value, and updates the optimized global model. The central coordinating node 20 distributes the optimized global model to each domain node 10. Each domain node 10 uses the global model as the initialization parameter of the local model for the next round of federated training and repeats the aforementioned federated training process for each round. Through continuous collaborative optimization of multiple rounds of federated training, a global shared model is finally obtained. Each domain node 10 performs DDoS attack detection processing based on the global shared model.

[0035] In this cross-domain federated learning system, each domain node holds a local private traffic dataset. Each domain node maintains a local model with a dual-output structure, responsible for local model training, parameter uploading, and DDoS attack detection based on the final globally shared model. The central coordination node is the core control node of the cross-domain federated learning system, holding a public dataset for distillation training. It is responsible for receiving model parameters uploaded from domain nodes, performing parameter aggregation, logistic value fusion, potential attack sample screening, knowledge distillation training, and distributing the optimized global model. The dual-output local model is the model maintained by the domain nodes, with two outputs: one is the DDoS attack binary classification result processed by the activation function (0 = normal traffic, 1 = attack traffic). The other layer consists of unprocessed logits, used for knowledge distillation by the central coordinating node. The public dataset is a small-scale auxiliary dataset held by the central coordinating node, available in labeled (containing flow features and 0 / 1 labels, suitable for supervised distillation) and unlabeled (containing only flow features, suitable for unsupervised distillation) formats, used to support knowledge interaction and distillation training for multi-domain models. The multi-round federated training strategy refers to the system employing a T-round iterative optimization mode. After each round of training, the central coordinating node distributes the optimized global model to each domain node as initialization parameters for the next round of local models, continuously improving the global model performance through iterative iteration. Logits are the outputs used for knowledge distillation in the dual-output structure of the local model. The model carries the original decision information on sample categories, serving as the teacher signal source for knowledge distillation. The weighted averaging strategy, used by the central coordinating node to aggregate domain node model parameters, uses the proportion of samples from each domain node's local private traffic dataset to the total number of samples from all domain nodes as weights. The aggregated parameters are obtained by weighted summation of the parameters from each domain node. The initial cross-domain joint model, obtained by the central coordinating node aggregating the model parameters from each domain node using the weighted averaging strategy, forms the basis for subsequent student models trained by knowledge distillation. The prediction confidence score, obtained by mapping the logical values ​​output by each domain node using the Sigmoid function, represents a 0-1 range value, reflecting the confidence of the corresponding domain node model in determining whether a sample in the public dataset constitutes attack traffic. The closer the value is to 1, the higher the confidence level. The fusion logic value is the logic value output by the central coordinating node to the multi-domain nodes. It is a unified output obtained by weighting and summing the values ​​according to the prediction confidence ratio of each domain. It serves as the basis for the teacher signal in knowledge distillation, strengthens the feature expression of attack-type samples, and suppresses knowledge interference from low-quality domain nodes. The potential attack sample set is a subset of samples selected from the public dataset. Knowledge distillation training is the process of optimizing the model by minimizing the difference between the output probability distributions of the two, using the soft label obtained after the fusion logic value is transformed by the activation function as the teacher signal and the preliminary cross-domain joint model as the student model. The global model is the optimized model obtained by the central coordinating node after each round of knowledge distillation training. It is distributed to each domain node as the initialization parameters for the next round of local training.The globally shared model is an optimized model with stable performance and strong generalization ability, obtained after multiple rounds of federated training iterations. Each domain node performs local DDoS attack detection based on this model, achieving cross-domain collaborative detection.

[0036] This cross-domain federated learning system can fundamentally solve the core pain points of existing federated learning in cross-domain DDoS attack detection, while achieving the dual goals of protecting data privacy and improving attack identification performance. The specific effects are reflected in two aspects:

[0037] First, it resolves the conflict between cross-domain data silos and privacy protection. By training models locally and uploading only model parameters, the system avoids the cross-domain transmission of raw network traffic data. This satisfies the data sovereignty, privacy protection, and network isolation requirements of each domain, while enabling multi-domain knowledge exchange through a public dataset of a central coordination node. This overcomes the limitations of traditional centralized methods that require the aggregation of raw data and are difficult to adapt to cross-domain scenarios, ensuring cross-domain collaborative modeling is completed without compromising privacy.

[0038] Secondly, this system specifically addresses the issues of data heterogeneity and class imbalance. Existing methods, such as FLDDoS, can only solve the problem of data heterogeneity between clients, but cannot handle class imbalance within a domain. Methods like FedDF and Fedic, which introduce distillation mechanisms, are mostly used for general tasks such as images and text, and are not suitable for the sparse nature of DDoS attack samples and the large differences in features between domains in network traffic. This system, however, uses confidence-weighted fusion logic values ​​to allow domain nodes with strong attack identification capabilities to contribute more knowledge, preventing low-quality information from obscuring scarce attack class features. Furthermore, it uses dynamic thresholding to filter potential attack samples and oversampling enhancement to focus on tail-end attack samples to alleviate their scarcity.

[0039] In specific application scenarios, the local model's dual-output structure includes a first output branch for binary classification of DDoS attacks and a second output branch for outputting logical values ​​for knowledge distillation. The first output branch focuses on the practicality of attack identification, specifically outputting the binary classification results of DDoS attacks (0 representing normal traffic and 1 representing attack traffic), directly supporting the training of the local model's basic classification capabilities. The second output branch focuses on cross-domain knowledge transfer, specifically outputting logical values ​​for knowledge distillation (the model's raw decision information before activation function processing), providing a data foundation for cross-domain knowledge fusion at the central coordination node.

[0040] Correspondingly, each domain node 10 trains a local model based on its local private traffic dataset, updates the local model parameters, and uploads them to the central coordination node 20. When the central coordination node 20 uses the public dataset to perform inference on the local models uploaded by each domain node 10 and generates multiple logical values ​​for knowledge distillation, the specific steps may include the following: Each domain node 10 uses its local private traffic dataset as training data to iteratively train its local model. During the iterative training, it repeatedly calculates the loss function of the local model based on the output results of the first output branch until it is determined that the loss function is less than a preset threshold, thus confirming that the local model has completed the update of its local model parameters. Each domain node 10 uploads the updated local model and local model parameters to the central coordination node 20 through an encrypted transmission channel. The central coordination node 20 does not need to obtain the original traffic data of each domain; it only inputs its own public dataset into the local models uploaded by each domain node 10 to perform the inference process and obtains the logical values ​​for knowledge distillation output by the second output branch.

[0041] The preset threshold is a numerical standard for the loss function used to determine whether the local model training has met the standards. It is set by technical personnel based on the actual attack detection accuracy requirements. When the loss function is less than the threshold during iterative training, it indicates that the local model has a reliable basic classification ability, and local training can be stopped and the model parameters updated. The encrypted transmission channel is a secure transmission link used by each domain node to upload models and parameters to the central coordination node. The transmitted model data is encrypted using encryption algorithms (such as SSL / TLS) to prevent parameters from being illegally stolen or tampered with during transmission, thus ensuring the security of model data.

[0042] By leveraging the synergistic effects of local model dual-output structure, iterative training to the loss threshold, encrypted transmission channels, and central public data inference, it is possible to ensure that local models at each domain node can fully learn the attack and normal traffic characteristics within private traffic. Reliable basic classification accuracy is achieved through loss calculation and iterative optimization of the first output branch, laying a high-quality model foundation for cross-domain collaboration. By transmitting only model parameters and using encrypted transmission, data privacy and model security are guaranteed from both the perspective of transmission content and channel security, mitigating the risk of cross-domain data leakage. Simultaneously, by utilizing the logical values ​​derived from public datasets on models across domains, a standardized, fine-grained knowledge carrier can be provided for cross-domain knowledge distillation, effectively addressing the issue of differences in knowledge representation between domains. This provides crucial support for the global model to integrate multi-domain attack identification experience through distillation and improve cross-domain DDoS detection performance.

[0043] In specific application scenarios, when the central coordination node 20 aggregates the local model parameters of each domain node 10 through a weighted average strategy to obtain a preliminary cross-domain joint model, the process includes: counting the number of samples in the local private traffic dataset corresponding to each domain node; calculating the weight value of each domain node based on the number of samples and the total number of samples corresponding to multiple domain nodes, where the weight value is the ratio of the number of samples to the total number of samples; multiplying the local model parameters of each domain node by the corresponding weight value to obtain the weighted model parameters; summing all the weighted model parameters element-wise to obtain the aggregated model parameters; and constructing a preliminary cross-domain joint model based on the aggregated model parameters.

[0044] By leveraging the synergistic effects of techniques such as statistically analyzing the number of domain node samples, calculating sample proportion weights, weighted parameter multiplication, and element-level summation aggregation, this approach ensures that model parameters of domain nodes with larger sample sizes receive higher contribution weights, preventing excessive dilution of parameters from small-sample domain nodes or excessive dominance from parameters from large-sample domain nodes. This guarantees a balanced aggregation process across multiple domain data characteristics. Furthermore, element-level summation precisely integrates the weighted model parameters from each domain, ensuring the rationality and effectiveness of the aggregation parameters. This allows the constructed preliminary cross-domain joint model to initially integrate traffic classification experience from multiple domains, laying a high-quality foundation for subsequent knowledge distillation-based enhancement of attack sample identification capabilities. Simultaneously, it adapts to the characteristics of large data distribution differences in cross-domain scenarios, avoiding model performance deviations caused by traditional equal-weight aggregation.

[0045] In specific application scenarios, when the central coordination node 20 calculates the prediction confidence of each logical value and merges multiple logical values ​​using the prediction confidence as weight to generate a fused logical value, the process includes: for a single network traffic sample in the public dataset, extracting the logical values ​​output by the local models uploaded by all domain nodes for the network traffic sample; inputting the logical values ​​of each domain node for the network traffic sample into the Sigmoid function for transformation to obtain the prediction confidence of each domain node for the network traffic sample; calculating the proportion of the prediction confidence of each domain node to the sum of the prediction confidence of all domain nodes for the network traffic sample, and determining the proportion as the fusion weight of the logical value corresponding to the domain node; multiplying the logical value of each domain node for the network traffic sample by the corresponding fusion weight to obtain the weighted result of each domain node with respect to the logical value, and then summing the weighted results of all domain nodes to generate the fused logical value of the network traffic sample; iterating through all network traffic samples in the public dataset and repeating the above steps to generate the fused logical value of each network traffic sample in sequence. Among them, network traffic samples are the basic data units in public datasets and private datasets of domain nodes. They contain feature vectors of network traffic (such as IP fragmentation frequency, TCP connection count, etc.) and are the basic objects for model reasoning and knowledge fusion.

[0046] By combining techniques such as extracting single-sample multi-domain logistic values, Sigmoid conversion of confidence scores, weighting based on confidence score proportions, and generating fused logistic values ​​through weighted summation, the Sigmoid function can be used to quantitatively convert logistic values ​​into predicted confidence scores, accurately representing the reliability of each domain model's judgment on samples. Furthermore, using confidence score proportions as fusion weights allows domain nodes with strong attack detection capabilities and high confidence to lead knowledge fusion, effectively suppressing the masking of scarce attack-related knowledge by domain nodes with large data volumes but low attack detection accuracy. Simultaneously, by integrating effective knowledge from multiple domains through weighted summation, the generated fused logistic values ​​unify the cross-domain knowledge expression format, avoiding fusion bias caused by differences in inter-domain knowledge. This provides high-quality teacher signals for subsequent attack sample selection based on fused logistic values ​​and knowledge distillation, ultimately helping the global model learn cross-domain attack-related features more accurately and improving cross-domain DDoS attack detection performance.

[0047] In specific application scenarios, when the central coordination node 20 filters the potential attack sample set in the public dataset, it includes: inputting the fusion logic value corresponding to each network traffic sample in the public dataset into the Sigmoid function for transformation to obtain the first attack confidence of each network traffic sample (the closer the first attack confidence is to 1, the higher the probability that the sample belongs to DDoS attack traffic); calculating the dynamic threshold for this round of distillation training based on the core parameters of the dynamic threshold, including the initial threshold, descent step size and minimum threshold; identifying network traffic samples in the public dataset whose first attack confidence is greater than the dynamic threshold as potential attack samples for this round of distillation training; and collecting all network traffic samples marked as potential attack samples to form the potential attack sample set for this round of federated training.

[0048] The first attack confidence score is a 0-1 range value obtained by transforming the fusion logic value corresponding to the network traffic sample using the Sigmoid function. It is used to quantify the probability that the sample belongs to DDoS attack traffic. The higher the value, the greater the probability that the sample is attack traffic, and it is the core judgment indicator for sample selection. The dynamic threshold is a sample selection standard that is dynamically adjusted with each round of federated training. It is used to determine whether a network traffic sample is a potential attack sample. Its value is calculated from the core parameters and can adapt to the sample selection needs of different training stages. The core parameters of the dynamic threshold are a set of key parameters used to calculate the dynamic threshold, including the initial threshold (the selection standard for the first round of training to ensure that high-confidence samples are selected), the descent step size (controlling the decrease of the threshold in each round to gradually relax the selection standard), and the minimum threshold (the lower limit of the threshold to avoid the inclusion of normal samples due to excessively lenient selection standards). Potential attack samples are network traffic samples in the public dataset whose first attack confidence score is greater than the dynamic threshold of the current round. They have a high probability of being attack traffic and are the core data objects for subsequent knowledge distillation training.

[0049] Accordingly, when calculating the dynamic threshold for this round of distillation training based on the core parameters of the dynamic threshold, the current federated training round can be used as a variable. The initial threshold is obtained by calculating "current dynamic threshold = initial threshold - descent step size × current training round". Finally, the initial threshold is compared with the minimum threshold. If the initial threshold is higher than the minimum threshold, the initial threshold is used as the dynamic threshold for this round of distillation training. If the initial threshold is lower than the minimum threshold, the minimum threshold is used as the dynamic threshold for this round. This achieves dynamic adjustment of the screening criteria from strict to lenient as the training progresses, balancing the accuracy of sample screening in the early stage of training with the generalization of sample coverage in the later stage.

[0050] By combining technical features such as fusing logical values ​​to convert them into first attack confidence, calculating dynamic thresholds based on core parameters, screening potential samples based on confidence, and collecting samples to form a sample set, this approach not only uses the Sigmoid function to transform abstract fused logical values ​​into intuitive quantitative indicators of attack probability (i.e., first attack confidence), providing a clear basis for sample screening; but also, through the design of core parameters for dynamic thresholds, allows the screening criteria to be flexibly adjusted as the training process progresses, avoiding the problems of overly strict screening leading to missed samples or overly lenient screening leading to noise due to fixed thresholds; and simultaneously, through confidence screening and sample collection, it can accurately focus on potential attack samples in public datasets, forming targeted training subsets, enabling subsequent knowledge distillation to concentrate resources on learning scarce attack-type features, effectively alleviating the problem of low proportion and scattered distribution of attack samples in cross-domain scenarios.

[0051] In specific application scenarios, knowledge distillation training is performed on the initial cross-domain joint model based on the potential attack sample set and the fusion logic value to update and obtain an optimized global model. Specifically, this can include the following steps: oversampling the potential attack sample set by means of repeated sampling or feature perturbation to increase the number of attack samples and obtain an enhanced attack sample set; using the enhanced attack sample set as input and the fusion logic value as the teacher signal source, the distillation loss of the initial cross-domain joint model is calculated, and then the parameters of the initial cross-domain joint model are continuously adjusted through optimization algorithms such as gradient descent to continuously minimize the distillation loss until the loss converges or reaches the preset training rounds. At this time, the model with the adjusted parameters is the optimized global model, which has a stronger ability to identify cross-domain DDoS attacks.

[0052] Correspondingly, the distillation loss is flexibly adapted to the label status of the public dataset: in unlabeled public data scenarios, only KL divergence is used as the sole loss term, achieving cross-domain knowledge transfer of attack-type features in an unsupervised manner, allowing the student model to learn the common features of attack samples without relying on manual annotation; in labeled public data scenarios, category-aware Focal Loss is introduced on top of KL divergence to construct a joint loss function. Focal Loss effectively alleviates the class imbalance problem of few attack samples and many normal samples by dynamically adjusting class weights (giving higher weights to attack samples) and sample difficulty weights (paying more attention to difficult-to-distinguish attack samples), preventing the model from being biased towards learning normal traffic features due to the high proportion of normal samples; while KL divergence continuously ensures the effective transfer of cross-domain attack knowledge. The two work together to enable the distillation loss to accurately align multi-domain attack knowledge and focus on strengthening the discrimination ability of minority class attack samples, ultimately driving the student model to transform into an optimized global model that can efficiently identify cross-domain DDoS attacks through parameter updates.

[0053] Among these methods, resampling is a common oversampling technique that expands the sample size by directly copying samples from the potential attack sample set. Feature perturbation is another common oversampling technique that makes small, reasonable adjustments to the feature vectors of samples in the potential attack sample set (such as adding low-amplitude Gaussian noise or fine-tuning the IP sharding frequency and other feature dimension values), increasing sample diversity while expanding the sample size and improving the model's robustness to attack sample variants. Oversampling is a sample augmentation technique that addresses the problem of scarce potential attack sample sets by expanding the number of attack samples through resampling or feature perturbation, providing sufficient attack-class training data for knowledge distillation and alleviating the impact of class imbalance. The enhanced attack sample set is a sample set obtained by sampling the potential attack sample set. It has a sufficient number of samples and retains the core features of the attack samples. It is the core input data for knowledge distillation training. The teacher signal (i.e., the soft label obtained by transforming the fused logic value by the Sigmoid function) is the standard signal used in knowledge distillation to guide the learning of the student model (i.e., the preliminary cross-domain joint model). The cross-domain attack knowledge contained in the fused logic value can help the student model improve its attack recognition ability. The distillation loss is an indicator that measures the difference between the output probability distribution of the preliminary cross-domain joint model and the teacher signal. The smaller the value, the higher the alignment between the two. It is the core optimization target for updating the model parameters.

[0054] By combining techniques such as oversampling to expand attack samples, fusing logical values ​​as teacher signal sources, and minimizing distillation loss to update parameters, the problem of scarce potential attack sample sets and difficulty in the model fully learning attack features can be effectively solved by oversampling through repeated sampling or feature perturbation. At the same time, feature perturbation can also improve the robustness of the model to attack sample variants. Using soft labels generated by fusing logical values ​​as teacher signals, high-quality attack identification knowledge from multiple domains is passed to the initial cross-domain joint model, which can avoid insufficient generalization caused by the model relying solely on single-domain experience. By minimizing distillation loss to update parameters, it is ensured that the global model can accurately align with cross-domain attack knowledge. While strengthening the ability to identify minority class attack samples, it also takes into account the adaptability of the model in multi-domain scenarios. The resulting optimized global model can effectively alleviate the class imbalance and data heterogeneity problems in cross-domain DDoS detection, and improve attack identification accuracy and generalization performance.

[0055] In summary, the cross-domain federated learning system provided by this invention constructs a cross-domain federated learning system comprising a central coordinating node and multiple domain nodes. Combining a dual-output local model with a public dataset, it achieves cross-domain knowledge collaboration and minority class attack sample modeling through multiple rounds of federated training while protecting data privacy. Each domain node trains a dual-output model based on its local private traffic data and uploads its parameters. The central coordinating node uses the public dataset to infer and generate logical values, fully utilizing local data from each domain while avoiding cross-domain transmission of raw data, effectively alleviating the problem of existing methods not being adapted to cross-domain scenarios. The central coordinating node aggregates domain node parameters through a sample-weighted average, ensuring the rationality of data contributions, and simultaneously fuses multi-domain logical values ​​based on prediction confidence. This approach enhances the knowledge contribution of domain nodes with strong attack identification capabilities, preventing low-quality knowledge from masking scarce attack information and addressing the pain point of existing methods struggling to fully aggregate multi-domain attack knowledge. By screening potential attack samples in public datasets and performing knowledge distillation on the initial cross-domain joint model using fusion logic values, it focuses on minority attack samples to strengthen model learning, specifically alleviating the problem of insufficient identification performance caused by intra-domain class imbalance and scarcity of tail attack samples. Finally, through multiple rounds of federated training iterations, it continuously optimizes the model's generalization ability and robustness. The resulting globally shared model enables each domain node to efficiently perform DDoS attack detection, balancing data privacy protection with significantly enhanced model perception and identification capabilities for minority attack samples.

[0056] Furthermore, to fully illustrate the implementation of this embodiment, this embodiment also provides a cross-domain DDoS attack detection method, which is applied to the aforementioned cross-domain federated learning system, such as... Figure 2 As shown, the method includes:

[0057] Step 210: Each domain node collects local network traffic in real time through network probes and extracts the feature vector of local network traffic.

[0058] Among them, the network probe is a dedicated data collection tool deployed in the local network of the domain node. It has the function of capturing and filtering network traffic in real time and can continuously acquire raw traffic data flowing through the local network segment. It is the core hardware / software component for traffic collection. Local network traffic refers to all network data streams generated or flowing through the domain node's jurisdiction, including legitimate traffic such as normal user access and business data transmission, as well as potential DDoS attack traffic. The feature vector is a standardized numerical vector extracted from the raw local network traffic that reflects the key attributes of the traffic. The extraction dimensions usually cover transmission protocol features, connection behavior features, packet features, etc., and it is the core input data for the global shared model to perform attack discrimination.

[0059] In this embodiment of the disclosure, each domain node deploys a dedicated data collection tool, a network probe, only within its own jurisdiction to capture all traffic data flowing through the local network in real time (including normal business traffic and potential attack traffic), avoiding data sovereignty conflicts and privacy leaks caused by cross-domain collection. At the same time, key indicators that characterize the essential attributes of the traffic (such as TCP connection establishment frequency, IP fragmentation quantity, UDP packet ratio, etc.) are extracted from the collected raw network traffic, and these indicators are standardized into fixed-dimensional numerical vectors (i.e., feature vectors), transforming unstructured raw traffic data into structured data that can be identified and calculated by the subsequent globally shared model, providing an input basis for accurately performing DDoS attack detection.

[0060] By combining local data collection at domain nodes with feature vector extraction, the real-time collection capabilities of network probes can ensure that each domain node can synchronously capture dynamic traffic changes in its local network, avoiding missed detections of attack behaviors due to traffic data delays. At the same time, the local collection mode can avoid security risks and compliance issues related to cross-domain data transmission. By extracting feature vectors, the messy information in the original network traffic is transformed into structured data focusing on key dimensions for attack identification. Redundant information interference can be eliminated, enabling the subsequent globally shared model to quickly locate attack-related feature signals. This lays the data foundation for the model to accurately output binary classification results and logical values, ultimately ensuring the real-time performance and accuracy of cross-domain DDoS attack detection.

[0061] Step 220: Input the feature vector into the locally configured global shared model and obtain the binary classification result output by the first output branch in the global shared model.

[0062] During the inference phase of cross-domain DDoS attack detection, each domain node inputs the previously extracted local network traffic feature vectors into the locally deployed global shared model. This model is the final model after multiple rounds of cross-domain federated training and knowledge distillation optimization, and it has the ability to identify cross-domain attacks. The global shared model directly outputs binary classification results through the first output branch (value 0 represents normal traffic, and value 1 represents attack traffic), providing an intuitive preliminary judgment for attack detection.

[0063] The core logic of outputting binary classification results through the first output branch can rely on the cross-domain attack identification knowledge integrated by the globally shared model (especially the ability to distinguish minority class and tail attack samples) to ensure that the binary classification results of the first output branch have high accuracy and cross-domain generalization, and avoid the identification bias caused by data heterogeneity in single-domain models.

[0064] Step 230: Determine the DDoS attack detection results of local network traffic based on the binary classification results.

[0065] Among them, the DDoS attack detection result is the final determination of whether local network traffic belongs to a DDoS attack. It is usually divided into two categories: "normal traffic" and "DDoS attack traffic", which are used to guide subsequent network security operations (such as blocking attacks and allowing normal traffic).

[0066] In this embodiment of the disclosure, the discrete classification result (0 or 1) output by the first output branch of the globally shared model can be used as the core basis to directly determine the attack attribute of local network traffic: if the binary classification result is 1, the traffic is determined to be DDoS attack traffic; if the binary classification result is 0, it is determined to be normal traffic. This process relies on the attack identification knowledge learned by the globally shared model through multiple rounds of cross-domain federated training, transforming the model's deep discrimination of feature vectors into intuitive detection conclusions, providing a clear basis for domain nodes to quickly respond to network security status.

[0067] To verify the effectiveness of this invention in DDoS attack identification tasks, a comparative experiment was designed, focusing on comparing the performance of this invention (FAIR-Distil) with the traditional federated averaging method (FedAvg) in attack class identification. Considering the importance of "recall rate" to the defense system in DDoS identification tasks, and the overall identification performance, the experiment selected the following two core indicators for evaluation:

[0068] (1) F1_Attack: The F1 score of the attack class, which measures the overall recognition effect of the attack sample (taking into account both precision and recall).

[0069] (2) Recall_Attack: The recall rate of attack classes measures the model's ability to capture attack samples and is a key indicator for DDoS detection tasks.

[0070] Table 1 below shows the final performance metrics of the method of this invention and FedAvg under the same number of rounds (e.g., 80 training rounds):

[0071] Table 1 Performance Comparison

[0072] FedAvg 0.958 0.928 FAIR-Distil 0.991 0.988

[0073] As shown in Table 1, within the first 80 rounds of training, this invention comprehensively outperforms the traditional FedAvg method in attack type identification performance, with a 6% improvement in Recall_Attack and a 3.3% improvement in F1_Attack, demonstrating faster model learning speed and stronger attack behavior identification capabilities. This performance lead is particularly significant in the early training stages, providing direct benefits for reducing communication costs and shortening the training cycle in practical deployments.

[0074] Figure 3 and Figure 4 The presentation shows the evolution curves of F1 score and recall for attack samples in the federated learning DDoS attack identification task, with respect to training rounds. The comparison reveals that while FedAvg gradually improves its identification performance with each round, it suffers from significant drawbacks: a gradual performance ramp-up and low training efficiency. In contrast, FAIR-Distill demonstrates a significant advantage through its multi-dimensional design: in the early training phase (first 20 rounds), the model's dual-output structure and the confidence-weighted aggregation strategy of the central coordinating node quickly take effect, resulting in a recall rate significantly higher than FedAvg. As the number of rounds increases, dynamic threshold scheduling and attack sample oversampling work together to accurately optimize the identification of minority class attack samples, leading to a continuous surge in performance that stabilizes at a high level, converging after 40 rounds. More importantly, FAIR-Distill achieves high attack identification performance within the first 30-40 rounds, while FedAvg requires training up to 40 rounds to reach a high level, and even after 40 rounds, FAIR-Distill's attack identification performance remains significantly superior to FedAvg.

[0075] In summary, the technical solution in this application ensures the real-time nature of attack detection by having each domain node collect local traffic and extract feature vectors in real time through network probes, thus avoiding missed attacks due to data latency. Furthermore, it can quickly obtain preliminary binary classification results of normal / attack traffic by leveraging the first output branch of the globally shared model. Finally, it can comprehensively determine the detection results based on the binary classification results, and rely on the attack-type features learned across domains by the globally shared model to ensure detection accuracy, reduce false positives for normal traffic and false negatives for attack traffic. Ultimately, this allows each domain node to obtain DDoS attack detection results that are real-time, accurate, and cross-domain generalizable without sharing the original data, effectively addressing the challenges of scarce attack samples and large differences in features between domains in cross-domain scenarios.

[0076] Furthermore, as Figure 2 The specific implementation of the method shown in this embodiment provides a cross-domain DDoS attack detection device, such as... Figure 5 As shown, the device includes: an extraction module 51, an input module 52, and a determination module 53.

[0077] The extraction module 51 can be used to collect local network traffic in real time by each domain node through network probes and extract the feature vector of local network traffic.

[0078] The input module 52 can be used to input feature vectors into the locally configured global shared model and obtain the binary classification result output by the first output branch in the global shared model;

[0079] The determination module 53 can be used to determine the DDoS attack detection results of local network traffic based on the binary classification results.

[0080] Based on the above, Figure 2 Accordingly, this embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the above-described method. Figure 2 The method for detecting cross-domain DDoS attacks is shown.

[0081] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause an electronic device (such as a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0082] Based on the above, Figure 2 To achieve the above objectives, this application also provides an electronic device, specifically a personal computer, tablet computer, server, or other network device, etc. This device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figure 2 The method for detecting cross-domain DDoS attacks is shown.

[0083] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.

[0084] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0085] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.

[0086] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware.

[0087] This solution constructs a systematic cross-domain DDoS attack detection process. Each domain node collects local traffic and extracts feature vectors in real time through network probes, ensuring real-time attack detection and avoiding missed attacks due to data latency. Furthermore, it can quickly obtain preliminary binary classification results of normal / attack traffic by leveraging the first output branch of the globally shared model. Finally, the detection result can be comprehensively judged based on the binary classification results. It can rely on the attack-type features learned across domains by the globally shared model to ensure detection accuracy, reduce false positives for normal traffic and false negatives for attack traffic. Ultimately, it enables each domain node to obtain DDoS attack detection results that combine real-time performance, accuracy, and cross-domain generalization without sharing the original data, effectively addressing the challenges of scarce attack samples and large differences in features between domains in cross-domain scenarios.

[0088] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or can be modified to be located in one or more apparatuses different from this embodiment. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.

[0089] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of this application.

Claims

1. A cross-domain federated learning system, characterized in that, The cross-domain federated learning system comprises a central coordinating node and multiple domain nodes. Each domain node maintains a local model with a dual-output structure. The dual-output structure of the local model includes a first output branch for binary classification of DDoS attacks and a second output branch for outputting logical values ​​for knowledge distillation. The central coordinating node holds a public dataset for distillation training. The cross-domain federated learning system employs a multi-round federated training strategy, and each round of federated training includes: Each domain node trains its local model based on its local private traffic dataset, updates its local model parameters, and uploads it to the central coordination node. The central coordination node uses the public dataset to perform inference on the local models uploaded by each domain node, generating multiple logical values ​​for knowledge distillation. This includes: each domain node using the local private traffic dataset as training data to iteratively train its local model, repeatedly calculating the loss function of the local model based on the output of the first output branch during the iterative training until the loss function is determined to be less than a preset threshold, indicating that the local model has completed the update of its local model parameters; each domain node uploading the updated local model and its local model parameters to the central coordination node via an encrypted transmission channel; and the central coordination node inputting the public dataset into the local models uploaded by each domain node to perform the inference process, obtaining the logical values ​​for knowledge distillation output by the second output branch. The central coordination node aggregates the local model parameters of each domain node through a weighted average strategy to obtain a preliminary cross-domain joint model. At the same time, it calculates the prediction confidence of each logical value and uses the prediction confidence as a weight to fuse the multiple logical values ​​to generate a fused logical value. The central coordination node filters a potential attack sample set from the public dataset, and performs knowledge distillation training on the preliminary cross-domain joint model based on the potential attack sample set and the fusion logic value to update and obtain an optimized global model. The filtering of the potential attack sample set by the central coordination node in the public dataset includes: inputting the fusion logic value corresponding to each network traffic sample in the public dataset into a Sigmoid function for transformation to obtain a first attack confidence for each network traffic sample; calculating a dynamic threshold for this round of distillation training based on core parameters of a dynamic threshold, including an initial threshold, a descent step size, and a minimum threshold; identifying network traffic samples in the public dataset whose first attack confidence is greater than the dynamic threshold as potential attack samples for this round of distillation training; and collecting all network traffic samples marked as potential attack samples to form a potential attack sample set for this round of federated training. The central coordinating node distributes the optimized global model to each domain node. Each domain node uses the global model as the initialization parameter for its local model in the next round of federated training and executes each round of federated training in a loop. Through continuous collaborative optimization of multiple rounds of federated training, a global shared model is finally obtained. Each domain node performs DDoS attack detection processing based on the global shared model.

2. The cross-domain federated learning system according to claim 1, characterized in that, When the central coordinating node aggregates the local model parameters of each domain node using a weighted averaging strategy to obtain a preliminary cross-domain joint model, it includes: Count the number of samples in the local private traffic dataset corresponding to each domain node; Based on the number of samples and the total number of samples corresponding to the plurality of domain nodes, the weight value of each domain node is calculated, and the weight value is the ratio of the number of samples to the total number of samples; The local model parameters of each domain node are multiplied by the corresponding weight values ​​to obtain the weighted model parameters; Element-wise summation is performed on all weighted model parameters to obtain aggregated model parameters, and a preliminary cross-domain joint model is constructed based on the aggregated model parameters.

3. The cross-domain federated learning system according to claim 1, characterized in that, When the central coordination node calculates the prediction confidence of each of the logical values ​​and merges the multiple logical values ​​using the prediction confidence as weights to generate a merged logical value, the process includes: For a single network traffic sample in the public dataset, extract the logical values ​​output by the local models of all domain nodes for the network traffic sample; The logical value of each domain node for the network traffic sample is input into the Sigmoid function for transformation to obtain the prediction confidence of each domain node for the network traffic sample. Calculate the proportion of the prediction confidence of each domain node to the sum of the prediction confidence of all domain nodes for the network traffic sample, and determine the proportion as the fusion weight of the logical value corresponding to the domain node. The logical value of each domain node for the network traffic sample is multiplied by the corresponding fusion weight to obtain the weighted result of the logical value of each domain node. Then, the weighted results of all domain nodes are summed to generate the fusion logical value of the network traffic sample. Iterate through all network traffic samples in the public dataset, repeat the above steps, and generate the fused logical value for each network traffic sample in sequence.

4. The cross-domain federated learning system according to claim 1, characterized in that, Based on the potential attack sample set and the fusion logic value, knowledge distillation training is performed on the initial cross-domain joint model to update the optimized global model, including: By employing resampling or feature perturbation, the potential attack sample set is oversampled to increase the number of attack samples, resulting in an enhanced attack sample set. Using the enhanced attack sample set as input and the fusion logic value as the teacher signal source, the distillation loss of the preliminary cross-domain joint model is calculated. The model parameters are updated by minimizing the distillation loss to obtain the optimized global model.

5. A method for detecting cross-domain DDoS attacks, characterized in that, The method is applied to the cross-domain federated learning system according to any one of claims 1 to 4, and the method includes: Each domain node collects local network traffic in real time through network probes and extracts the feature vector of the local network traffic; The feature vector is input into the locally configured global shared model to obtain the binary classification result output by the first output branch of the global shared model; The DDoS attack detection result of the local network traffic is determined based on the binary classification result.

6. A cross-domain DDoS attack detection device, characterized in that, The apparatus is applied to the cross-domain federated learning system according to any one of claims 1 to 4, the apparatus comprising: The extraction module is used to collect local network traffic in real time by each domain node through network probes, and extract the feature vector of the local network traffic. The input module is used to input the feature vector into a locally configured global shared model and obtain the binary classification result output by the first output branch in the global shared model; The determination module is used to determine the DDoS attack detection result of the local network traffic based on the binary classification result.

7. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of claim 5.

8. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of claim 5.

Citation Information

Patent Citations

  • Network attack model cooperative training method based on federated learning

    CN120321030A

  • Internet of vehicles anomaly detection method based on double knowledge distillation and federated learning

    CN120342781A