Transform and knowledge distillation-based privacy protection federated learning method and system
By introducing Transformer as a local feature extractor and Paillier encryption protocol into federated learning, combined with knowledge distillation technology, the problems of low computational efficiency and generalization under Non-IID data are solved, achieving a balance between privacy protection and detection performance. This improves the computational efficiency and detection accuracy of edge devices and is suitable for industrial IoT and sensitive fields.
Patent Information
- Application Number
- CN202511329995.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-17
AI Technical Summary
Existing federated learning frameworks have significant shortcomings in data privacy protection, model performance optimization, and applicability to resource-constrained devices. They suffer from low computational efficiency, unresolved generalization problems under non-IID data, and a lack of full integration of Transformer's attention mechanism and knowledge distillation technology, making it difficult to balance privacy protection and detection performance.
The Transformer is used as the local feature extractor, combined with the Paillier encryption protocol. The central server uses knowledge distillation technology to transfer and aggregate model weights in an encrypted state, improve feature quality by leveraging the attention mechanism of the Transformer, and optimize the generalization ability of small models through knowledge distillation technology, thus achieving a balance between data security and detection efficiency.
It significantly improves detection performance under Non-IID data, achieving an intrusion detection accuracy of 96.5%, reduces computational overhead, is suitable for resource-constrained edge devices, reduces the risk of privacy leaks, and is applicable to industrial IoT and sensitive fields.
Smart Images

Figure CN121098591A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of artificial intelligence and network security, and particularly relates to a privacy protection federated learning method and system based on a Transformer and knowledge distillation. BACKGROUND
[0002] With the rise of IIoT (Industrial Internet of Things) and distributed cloud platforms, the massive data generated by edge devices such as sensors and cloud nodes has become a core asset, but at the same time, it faces serious privacy leakage risks. Traditional centralized machine learning requires uploading all data to a central server for training, which not only increases security risks during data transmission, but also may violate international privacy regulations, especially in sensitive industries such as chemical industry and power industry, data leakage may cause significant economic losses or safety accidents.
[0003] To address these challenges, federated learning (Federated Learning) has emerged as a distributed learning paradigm. It allows clients to train models locally and only share model updates (such as weights) rather than raw data, thereby protecting data privacy to some extent. This technology aims to solve mobile device privacy problems and gradually extends to the industrial field. However, in practical applications, federated learning still faces model performance degradation caused by Non-IID (Non-IID) data distribution, excessive computational overhead of resource-constrained edge devices, and the need for high precision in intrusion detection tasks. At the same time, advanced mechanisms such as attention mechanisms (such as Transformer models, which perform well in natural language processing and computer vision, and can capture long-range dependencies) and knowledge distillation techniques (used for model compression and generalization improvement) have not been fully integrated into the federated learning framework, making it difficult to balance privacy protection and detection efficiency.
[0004] Existing federated learning frameworks combine homomorphic encryption protocols, such as introducing Paillier (a probabilistic public key encryption algorithm that supports homomorphic operations, commonly used in privacy protection scenarios) or similar encryption mechanisms in federated learning systems (such as FedAvg variants combined with homomorphic encryption). The core of these schemes is that clients train models locally (such as using CNN or RNN to extract features), then transmit model updates to the central server for aggregation through encryption protocols, and finally distribute the global model back to the clients. This technology has been applied in some industrial internet of things intrusion detection systems, aiming to solve data privacy problems while supporting distributed training.
[0005] However, existing federated learning frameworks have significant shortcomings in data privacy protection, model performance optimization, and resource-constrained device applicability. Specifically, existing technologies or scenarios have the following key problems:
[0006] Low computational efficiency: Homomorphic encryption supports operations in an encrypted state, but introduces high computational overhead, especially on resource-constrained edge devices, resulting in slow training iterations and failing to meet real-time intrusion detection requirements.
[0007] Generalization problem under Non-IID data is not effectively solved: In the scenario of uneven data distribution, the model performance decreases significantly, and the intrusion detection accuracy is difficult to stabilize above 95%. In sensitive fields such as chemical industry and electric power, the weak generalization ability leads to an increase in false positives or false negatives.
[0008] Insufficient mechanism fusion: Existing solutions rely on traditional feature extractors (such as CNN) and do not fully utilize the attention mechanism of Transformer to improve feature quality. Meanwhile, there is a lack of knowledge distillation technology to compress the model and improve the generalization of small models, resulting in insufficient computing power of edge devices and poor robustness of the overall system, which cannot balance privacy protection and detection efficiency. SUMMARY
[0009] To address the problems mentioned in the background art, the present application proposes a privacy protection federated learning method and system based on Transformer and knowledge distillation. By using Transformer as a local feature extractor, combining Paillier encryption protocol, and using knowledge distillation technology after aggregation by the central server, the detection performance under Non-IID data is optimized, and the balance between data security and detection efficiency is achieved.
[0010] Technical solution: To solve the above technical problems, the technical solution adopted by the present application is as follows:
[0011] A privacy protection federated learning method based on Transformer and knowledge distillation, which includes using the attention mechanism of the Transformer model as the core component of local feature extraction to capture long-distance dependency relationships in data and improve feature representation quality; introducing Paillier encryption protocol to realize homomorphic encryption transmission of model weights; using knowledge distillation technology on the central server side to extract soft knowledge from the aggregated global model as a teacher model, and feeding back to the client to compress the model and optimize, including the following steps:
[0012] S1, data acquisition and preprocessing: collecting data from local sensors and preprocessing it into sequence format;
[0013] S2, feature extraction: deploying a Transformer model as a local feature extractor on each client device for feature extraction;
[0014] S3, local model training: training an intrusion detection model based on extracted features and optimizing the loss function;
[0015] S31, the output of the input self-attention mechanism is connected with the original input, and the residual connection is performed;
[0016] S32, the input of S31 is connected with the residual, and the normalization of the residual result is performed;
[0017] S33, the normalized result of S32 is added to the feedforward network;
[0018] S34, the output of the second layer linear transformation in S33 and the normalized result in S32 are connected with the residual and layer normalized;
[0019] S35, output high-dimensional feature representation;
[0020] S4, encrypted transmission: after the client completes the local training, the model weight is encrypted by applying the Paillier encryption protocol, and then the weight is transmitted to the central server through the secure channel;
[0021] S5, central server aggregation: after the central server receives the encrypted weights of multiple clients, it performs aggregation operation to generate global model;
[0022] S6, knowledge distillation backwash: the central server decrypts the aggregated global model as a teacher model, then applies knowledge distillation technology to refine knowledge, and encrypts the refined knowledge to backwash each client for updating the local model;
[0023] The distillation process includes generating soft labels: the teacher model calculates the softmax output for the shared pseudo data set, and the temperature parameter t softens the probability distribution, specifically:
[0024] ,
[0025] wherein, represents the softened probability of the i-th class of soft label; t represents the temperature parameter; represents the score of the j-th class in logits vector, j represents the class index in summation;
[0026] Then, the student model minimizes the difference with the teacher soft label by KL divergence loss, specifically:
[0027] ,
[0028] wherein, LOSS represents the total loss function; CE represents cross entropy, KL represents divergence, and α represents balance factor, represents the real label; represents the predicted probability distribution of the student model; represents the soft label probability distribution of the teacher model;
[0029] S7, Iterative optimization: close-loop iteration is performed on the above steps until the global model converges.
[0030] As preferred, in S2, the specific implementation process is:
[0031] The encoder layer of the Transformer model uses the multi-head self-attention mechanism to calculate the dependency weight between the input sequences. Specifically, the dot product attention calculation is performed using the query, key and value matrices, specifically:
[0032] ,
[0033] Where softmax() represents the normalization function; Q represents the query matrix, K represents the key matrix, represents the key dimension, and T represents the transpose operation of the matrix.
[0034] As preferred, in S4, the specific implementation process is:
[0035] The Paillier encryption protocol is applied to encrypt the model weights, and the specific formula for calculating the ciphertext c is:
[0036] ,
[0037] Where r represents a random number; represents the plaintext message; n represents the modulus; g represents the base; represents the m-th power of the base g, which is used to encode the plaintext message m; represents the n-th power of the random number r, which is used to introduce randomness.
[0038] As preferred, in S5, the specific implementation process is:
[0039] Using the homomorphic property of Paillier, the encrypted weights are executed for weighted average, and the specific calculation formula is:
[0040] ,
[0041] Where, represents the client weight coefficient; represents the global model weight; represents the local model weight of the i-th client; represents the encryption function;
[0042] Then initialize the global model, iteratively accumulate the encrypted contribution, and finally output the global model in encrypted form.
[0043] As preferred, the specific content of initializing the global model, iteratively accumulating the encrypted contribution, and finally outputting the global model in encrypted form is:
[0044] Step 1, initialize global model: global model weights initialized as a zero vector;
[0045] Step 2, iteratively accumulate encrypted contributions: for each client, the server calculates the encrypted contribution of the client, then adds noise to the encrypted contribution, and finally accumulates to the global model;
[0046] Step 3, output the global model in encrypted form.
[0047] A privacy protection federated learning system based on Transformer and knowledge distillation, which implements the privacy protection federated learning method based on Transformer and knowledge distillation described in any of the above, including a client and a central server;
[0048] The client includes a local sensor, a feature extractor, a local trainer, and an encrypted transmitter.
[0049] Raw data is collected from the local sensor and preprocessed to convert it into a sequence format; the feature extractor is used to extract features from the preprocessed data; the local trainer is used to train the intrusion detection model based on the extracted high-dimensional features, and the loss function is optimized; finally, the encrypted transmitter is used to apply the Paillier encryption protocol to the trained model weights for homomorphic encryption, and transmit them to the central server through a secure channel;
[0050] The central server includes an aggregator, a knowledge distiller, and a counter-feeding distributor.
[0051] The encrypted weights are received from multiple clients, the aggregator is used to perform weighted average aggregation on the encrypted weights using the Paillier homomorphic property, and the global model in encrypted form is generated; the knowledge distiller decrypts the aggregated global model as a teacher model, applies knowledge distillation technology to the shared pseudo data set or anonymous samples, generates soft labels and intermediate representations, and refines knowledge to improve model generalization and compression; the counter-feeding distributor applies the refined knowledge to the Paillier protocol for encryption, and then distributes it back to each client through a secure channel, which is used to update the local model and optimize the performance under Non-IID data.
[0052] Advantages: Compared with the prior art, the present application has the following advantages:
[0053] (1) The application constructs a privacy protection type federated learning framework, which uses Transformer as a local feature extractor (uses its attention mechanism to improve feature quality), combines Paillier encryption protocol (supports homomorphic operation and ensures aggregation of model weights in an encrypted state without leaking original information), and uses knowledge distillation technology after central server aggregation (uses a global model as a teacher model to improve the generalization ability of the small model), thereby optimizing the detection performance under Non-IID data, achieving the balance between data security and detection efficiency. This concept not only solves the bottleneck of the prior art, but is also applicable to industrial Internet of Things, distributed cloud platforms and other scenarios, and can be extended to medical, financial and other privacy-sensitive fields, ultimately significantly reducing the risk of data leakage and improving the accuracy of intrusion detection (> 95%).
[0054] (2) The application uses Transformer as a local feature extractor, combines Paillier encryption protocol to transmit model weights, and uses knowledge distillation technology to feed back each client model after central server aggregation, achieving a balance between privacy protection and detection efficiency. The method and system of the application can effectively solve the key problems existing in the prior art, including performance degradation of federated learning under Non-IID data, large computational overhead of edge devices, low computational efficiency caused by encryption protocols, and insufficient fusion of attention mechanisms with federated learning and knowledge distillation, thereby completing the closed loop of the invention logic: from identification of background problems to proposal of technical solutions, and then to verification of actual effects. Specifically:
[0055] Solve the privacy leakage risk and improve data security: In the prior art, traditional centralized machine learning requires data to be uploaded to the central server, increasing the transmission security risk and possibly violating privacy regulations (such as GDPR); even if an encryption protocol (such as homomorphic encryption) is introduced, it often leads to low computational efficiency and does not fully protect the privacy of model updates. The application supports homomorphic operation through Paillier encryption protocol and performs model aggregation in an encrypted state to avoid decryption operations and original weight leakage, thereby significantly reducing the risk of data leakage. Experimental verification shows that the privacy leakage risk is reduced to 1.0 through differential privacy measurement, ensuring the security of data in industrial control systems (such as chemical and power industries). This effect achieves high-level privacy protection without sacrificing computational efficiency, making it suitable for sensitive industries.
[0056] Optimize model performance and generalization ability under Non-IID data: The invention innovatively integrates the attention mechanism of Transformer for local feature extraction, which can capture long-range dependencies in data and improve feature representation quality; combined with knowledge distillation technology, the global model is used as a teacher model to extract knowledge (soft labels or intermediate representations) and feed back to the client, improving the generalization ability of small models. Experimental results: On the Non-IID dataset (such as KDD Cup 99 variants), the intrusion detection accuracy reaches 96.5%, which is 8% higher than traditional federated learning. This effect solves the performance bottleneck under Non-IID data, improves the robustness and detection efficiency of the system (precision > 95%), especially on resource-constrained edge devices.
[0057] Reduce computational overhead and improve edge device applicability: The invention optimizes small models through efficient feature extraction and knowledge distillation of Transformer, reducing the computational requirements of local training; the homomorphic property of Paillier protocol ensures efficient aggregation without additional decryption overhead. This effect solves the problem of insufficient computing power on edge devices, improves computational efficiency, maintains high-precision intrusion detection, and is suitable for distributed environments such as industrial Internet of Things and distributed cloud platforms.
[0058] Expand application scenarios and improve business value: The invention framework is applicable to intrusion detection in industrial Internet of Things and distributed cloud platforms, and can be extended to other privacy-sensitive fields such as medical and financial. Through multi-mechanism fusion (Transformer, federated learning, knowledge distillation), the overall robustness and applicability of the system are improved. This effect completes the invention logic loop: from solving the core problem of data privacy and efficiency balance to achieving precision improvement and risk reduction, promoting safe applications in sensitive industries (such as chemical industry and power industry). BRIEF DESCRIPTION OF DRAWINGS
[0059] Figure 1 is the network topology diagram of the invention;
[0060] Figure 2 is the data processing flowchart of the invention;
[0061] Figure 3 is the timing diagram of the invention;
[0062] Figure 4 is the model performance comparison curve of the invention;
[0063] Figure 5 is the final accuracy comparison graph under different data heterogeneity of the invention;
[0064] Figure 6 is the privacy-utility trade-off analysis graph of the invention;
[0065] Figure 7 is a client local computing overhead comparison chart of the present application;
[0066] Figure 8 is a Transformer local feature extraction flowchart of the present application;
[0067] Figure 9 is a central server knowledge distillation counter-feeding flowchart of the present application;
[0068] Figure 10 is a privacy-precision trade-off curve chart of the present application. DETAILED DESCRIPTION
[0069] The present application will be further illustrated below in combination with specific embodiments, which are implemented on the premise of the technical solutions of the present application, and it should be understood that these embodiments are only used to illustrate the present application and not to limit the scope of the present application.
[0070] Embodiment 1
[0071] Based on the comprehensive needs of privacy protection and performance optimization of federated learning in distributed environments such as industrial Internet of Things and distributed cloud platforms, the present application aims to build an efficient and secure framework through deep integration of multiple mechanisms. Specifically, this concept takes the attention mechanism of the Transformer model as the core component of local feature extraction to capture long-range dependencies of data and improve the quality of feature representation; at the same time, it introduces the Paillier encryption protocol to realize the homomorphic encryption transmission of model weights, ensuring that the aggregation process does not leak original information; further, at the central server end, knowledge distillation technology is used to extract soft knowledge (including soft labels and intermediate representations) from the aggregated global model as a teacher model, and to feed back to the client to compress the model and optimize generalization. The core of this concept lies in the synergy between mechanisms: Transformer provides high-quality local features, encryption protocols guarantee transmission security, and knowledge distillation solves Non-IID data distribution and edge computing limitations, thereby achieving a dynamic balance between privacy protection and intrusion detection efficiency (precision > 95%). This concept is not simply additive, but forms a closed-loop system through an iterative optimization cycle (such as repeated local training-encryption-aggregation-distillation), which is suitable for intrusion detection tasks in sensitive scenarios such as chemical industry and power industry.
[0072] The following are explanations of terms in this embodiment:
[0073] LLM: Large Language Models, large language models, such as ChatGPT and other generative AI models.
[0074] MoE: Mixture of Experts, a mixed expert model, an architecture that dynamically routes to multiple expert sub-models.
[0075] PU Learning: Positive-Unlabeled Learning, a semi-supervised learning method that only requires positive samples and unlabeled data.
[0076] Jailbreak Attack: A form of bypassing AI model security restrictions, including prompt injection and malicious instruction bypass.
[0077] Prompt Injection: A type of attack that bypasses model restrictions by inputting prompts.
[0078] Malicious Instruction Bypass: A type of attack that bypasses security mechanisms by inputting instructions.
[0079] Transformer: A deep learning model based on self-attention mechanism for sequence processing.
[0080] Gating Network: A component in MoE that dynamically routes input to expert sub-models.
[0081] Experts: Independent sub-networks in MoE for processing specific input shards.
[0082] Reliable Negative Extraction: An algorithm that filters high-confidence negative samples from unlabeled data.
[0083] AdvBench: Adversarial Benchmark, a public dataset for testing adversarial attacks.
[0084] API: Application Programming Interface, used for system integration.
[0085] CNN: Convolutional Neural Network, used for text feature extraction.
[0086] RNN: Recurrent Neural Network, used for sequence data processing.
[0087] LSTM: Long Short-Term Memory, a variant of RNN that captures long dependencies.
[0088] GRU: Gated Recurrent Unit, a simplified variant of RNN.
[0089] Bi-RNN: Bidirectional RNN, a variant of RNN that processes data in both forward and backward directions.
[0090] GNN: Graph Neural Network, a type of neural network used for graph-structured data.
[0091] GAT: Graph Attention Network, a variant of GNN that focuses on attention mechanisms.
[0092] GAN: Generative Adversarial Network, a method for data generation and discrimination.
[0093] RL: Reinforcement Learning, an optimization method for decision-making through reward signals.
[0094] BPE: Byte Pair Encoding, a subword tokenization method.
[0095] POS Tagging: Part-of-Speech Tagging, a method for enhancing text semantic processing.
[0096] ELMo: Embeddings from Language Models, a dynamic embedding method.
[0097] IIoT: Industrial Internet of Things.
[0098] Non-IID: Non-Independent and Identically Distributed.
[0099] GDPR: General Data Protection Regulation, a data privacy law.
[0100] Federated Learning: a distributed learning paradigm that allows clients to train locally and share model updates.
[0101] Paillier: Paillier encryption protocol, a public-key encryption scheme that supports homomorphic operations for privacy protection.
[0102] FedAvg: Federated Averaging, a model aggregation algorithm in federated learning.
[0103] KDD Cup 99: KDD Cup 99 dataset, a benchmark dataset for intrusion detection.
[0104] CKKS: Cheon-Kim-Kim-Song, a fully homomorphic encryption scheme that supports multiplication and addition operations.
[0105] BFV: Brakerski-Fan-Vercauteren, a lattice-based homomorphic encryption scheme that supports integer operations.
[0106] PHE: Partial Homomorphic Encryption, a partial homomorphic encryption that supports specific operations such as multiplication.
[0107] DP: Differential Privacy, a mechanism for quantifying privacy protection.
[0108] MPC: Multi-Party Computation, a joint computation technique for privacy protection.
[0109] KD: Knowledge Distillation, a model compression and knowledge transfer technique.
[0110] ViT: Vision Transformer, a Transformer variant for image processing.
[0111] BERT: Bidirectional Encoder Representations from Transformers, a pre-trained Transformer model.
[0112] TTF: Tab Transformer, a Transformer variant for tabular data.
[0113] AE: Autoencoder, a feature compression and data reconstruction technique.
[0114] ECA-Net: Efficient Channel Attention Network, an attention-enhanced CNN variant.
[0115] ResNet: Residual Network, a CNN architecture.
[0116] Shamir: Shamir secret sharing scheme, a protocol for splitting a secret among multiple parties.
[0117] ElGamal: ElGamal encryption scheme, a public-key encryption that supports multiplicative homomorphism.
[0118] Skefl: Skefl protocol, a lightweight homomorphic encryption framework for federated learning.
[0119] Laplace: Laplace mechanism, a mechanism for adding noise in differential privacy.
[0120] FedSplit: FedSplit framework, a federated learning variant with dynamic aggregation in stages.
[0121] FedSSD: Federated Selective Self-Distillation, a local self-distillation method.
[0122] Data-free KD: Data-free Knowledge Distillation, distillation using synthetic data.
[0123] FedGen: Federated Generator, a generator framework for data-free distillation.
[0124] FD: Federated Distillation, a client-to-client knowledge transfer method.
[0125] AdaptiveKD: Adaptive Knowledge Distillation, combining mechanisms such as clustering.
[0126] FedAsync: Federated Asynchronous, an asynchronous federated learning variant.
[0127] HIPAA: Health Insurance Portability and Accountability Act, a US medical privacy regulation.
[0128] CAN: Controller Area Network, used for vehicle communication.
[0129] The framework of the present application consists of multiple client devices, a central server, and a communication module, wherein the client devices integrate a Transformer-based local model for data processing; the central server is responsible for weight aggregation and knowledge distillation calculation; the communication module embeds the Paillier encryption protocol to support secure transmission under homomorphic operation.
[0130] The privacy protection federated learning method based on Transformer and knowledge distillation provided by the embodiment has the following specific implementation steps:
[0131] S1, data acquisition and preprocessing: collecting data from local sensors and preprocessing them into sequence format;
[0132] First, the local data set (such as traffic logs or video stream data in industrial Internet of Things) is preprocessed, including tokenization and positional encoding, to generate input embedding vectors.
[0133] S2, local feature extraction: deploy a Transformer model as a local feature extractor on each client device (such as an edge sensor or a cloud node); as shown in Figure 8 .
[0134] Then, the encoder layer (multi-layer stacked, usually 6-12 layers) of the Transformer calculates the dependency weight between the input sequences through the multi-head self-attention mechanism, for example, using the query (Query), key (Key) and value (Value) matrix to perform dot product attention calculation:
[0135]
[0136] wherein, is a normalization function used to convert the input vector into a probability distribution, so that the sum of all elements is 1 and each element is between 0 and 1; Q is the query matrix, K is the key matrix, is the key dimension, and T represents the transpose operation, which is used to interchange the rows and columns of the key matrix K (i.e. K T ) so as to perform matrix multiplication (dot product operation) with the query matrix Q.
[0137] This mechanism captures long-range dependencies, such as temporal patterns in intrusion data.
[0138] S3, Local model training: Train intrusion detection model based on extracted features, optimize loss function;
[0139] Add feed-forward network and layer normalization, output high-dimensional feature representation, for downstream intrusion detection model training (such as classification head).
[0140] In the encoder layer of the Transformer model, after the Multi-Head Self-Attention calculation is completed, the output needs to be further processed to enhance the model's non-linear representation ability and stability. Therefore, a Feed-Forward Network (FFN) and Layer Normalization (LN) need to be added. These components are standard parts of the Transformer architecture, which are used to convert the intermediate representation of the self-attention output into a higher-quality high-dimensional feature representation, facilitating downstream tasks (such as intrusion detection model training).
[0141] The core goal of this process is:
[0142] 1. Introduce non-linear transformation through feed-forward network to improve feature representation ability.
[0143] 2. Through layer normalization and residual connection (Add & Norm), alleviate the problem of gradient vanishing / explosion, ensure training stability.
[0144] 3. Output high-dimensional feature representation, usually with a dimension of the model's hidden layer size (e.g. 512 or 768), which captures the long-distance dependencies and semantic information of the input sequence.
[0145] The specific operation steps are:
[0146] S31, input the output of the self-attention mechanism and the original input, and perform residual connection on it;
[0147] The self-attention mechanism outputs an intermediate tensor, denoted as AttentionOutput, with a shape of (batch_size, sequence_length, hidden_dim), where hidden_dim is the hidden dimension (e.g. 512); batch_size represents the number of data samples; sequence_length represents the length of a single sample.
[0148] S32, perform residual connection on the input of S31, and normalize the residual result;
[0149] Residual connection: The self-attention output is added to the original input to preserve the information flow. The specific formula is as follows:
[0150] ;
[0151] in, This represents the residual connection result; The intermediate tensor represents the output of the self-attention mechanism; Input is the embedding vector before entering self-attention, i.e., the original input.
[0152] Layer normalization is applied to normalize the residual connection results and stabilize their distribution. The layer normalization formula is as follows:
[0153]
[0154]
[0155]
[0156] in, Indicates to Perform normalization. This represents the i-th input vector. It is the input vector (calculated along the last dimension), i.e., the result of the residual connection; d is the mean, and d is the hidden dimension. It is variance. It is a small constant (usually 1e-5) to prevent division by zero. and These are learnable parameters (initialized to 1 and 0).
[0157] The output is denoted as Norm1.
[0158] Objective: To prevent gradient problems and provide standardized inputs for feedforward networks.
[0159] S33. Add a feed-forward network to the normalized result of S32.
[0160] Input: Norm1 from the output of S32.
[0161] A feedforward network is a two-layer fully connected network (MLP) that uses an activation function (usually ReLU or GELU) to introduce nonlinearity.
[0162] The first layer of linear transformation expands the input to a higher dimension (typically 4 times the hidden dimension, such as 2048) to increase expressive power.
[0163] Activation: Apply the ReLU activation function.
[0164] Second linear transformation: Project the dimensions back to the hidden dimensions.
[0165] The output is denoted as FFNOutput.
[0166] Objective: Enhance the nonlinear representation of features, capture more complex patterns.
[0167] The formula for adding a Feed-Forward Network to the normalized result of S32 is:
[0168]
[0169] where, FFNOutput represents the output of the Feed-Forward Network; Norm1 is the input (from the first layer normalization result); W1 is the first layer weight matrix (shape: hidden_dim x intermediate_dim, e.g. 512 x 2048); b1 is the first layer bias term (shape: intermediate_dim); ReLU is the ReLU activation function. W2 is the second layer weight matrix (shape: intermediate_dim x hidden_dim); b2 is the second layer bias term (shape: hidden_dim).
[0170] In this embodiment, the ReLU activation function can be replaced by GELU:
[0171] ,
[0172] where, Norm1 is the input (from the first layer normalization result).
[0173] S34, residual connection and layer normalization are performed on the output of the second layer linear transformation in S33 and the normalized result in S32;
[0174] Add residual connection: ;
[0175] where, Residual2 represents the residual connection result; FFNOutput represents the output of the second layer linear transformation in S33, and Norm1 represents the output of the normalized result in S32.
[0176] Apply layer normalization: Normalize the residual connection result, and the output is denoted as Norm2.
[0177] Objective: To further stabilize the training and output the final high-dimensional feature representation.
[0178] The formula for layer normalization is:
[0179]
[0180]
[0181]
[0182] in, Indicates to Perform normalization. This represents the i-th input vector. It is the input vector (calculated along the last dimension), i.e., the residual connection result Residual2; d is the mean, and d is the hidden dimension. It is variance. It is a small constant (usually 1e-5) to prevent division by zero. and These are learnable parameters (initialized to 1 and 0).
[0183] S35. Output high-dimensional feature representation:
[0184] Norm2 is the high-dimensional feature representation of this encoder layer, which can be directly used in downstream intrusion detection models (such as by adding a classification head).
[0185] If it is a multi-layer Transformer, then Norm2 is used as the input of the next layer.
[0186] In the context of intrusion detection, these features can be fed into a simple linear classifier that optimizes the cross-entropy loss function.
[0187] The entire process is performed layer by layer and implemented in the Transformer model on the local client. The computational overhead mainly comes from matrix multiplication, which can be executed efficiently using frameworks such as PyTorch (e.g., using GPU acceleration on edge devices).
[0188] In this embodiment, the complete Transformer sublayer formula (combining residuals and normalization) is:
[0189]
[0190] In practice: all operations are batch processed, supporting parallel computing.
[0191] parameter , , 、 、 、 optimized by backpropagation.
[0192] For intrusion detection, high_dim_features can be flattened and connected to a softmax classifier with cross-entropy loss , the specific formula is:
[0193]
[0194] where, represents the one-hot encoding of the true label (1 for the i-th class, and 0 for others). represents the softmax probability of the i-th class predicted by the model. C represents the number of classes.
[0195] Example parameters: hidden_dim=512, intermediate_dim=2048, number of layers=6, to ensure running time <1 second / iteration on edge devices.
[0196] Local training uses gradient descent optimization (such as Adam optimizer) to update model weights, but does not upload raw data. Specific parameters for this step include the number of attention heads (e.g. 8 heads) and hidden dimensions (e.g. 512) to ensure efficient operation on resource-constrained devices.
[0197] Specifically, after the high-dimensional feature representation is output, it is input to a simple classification head (e.g. linear layer) with cross-entropy loss function as the optimization objective, which is suitable for classification problems of intrusion detection (such as binary or multi-class intrusion classification). The optimization of the loss function is achieved by gradient descent (such as Adam optimizer), which only updates the local model weights without uploading raw data.
[0198] These are consistent with the Transformer feature extraction in this application (output high-dimensional features after adding feedforward network and layer normalization), parameter settings (such as 8 heads of attention heads and 512 of hidden dimensions) and efficient operation requirements. The entire process is executed locally on the client side, supporting resource-constrained devices (such as edge sensors), ensuring efficiency through small batch training and parameter optimization. The specific optimization process is:
[0199] 1. Obtain features and labels: obtain high-dimensional feature representation in S3, flatten or pool features, and input real labels in local dataset;
[0200] Get high-dimensional feature representation from S3 (Norm2, shape (batch_size, sequence_length, hidden_dim), e.g., hidden_dim=512).
[0201] Flatten or pool features (e.g., by average pooling or taking [CLS] token representation) to get fixed-dimension vector for downstream classification.
[0202] Input true labels (y) of local dataset, e.g., class labels (one-hot encoding, e.g., normal / attack) in intrusion detection.
[0203] 2. Build classification head: add a linear classification layer to map high-dimensional features to class number;
[0204] Add a linear classification layer: map high-dimensional features to class number C (e.g., C=2 for binary classification).
[0205] Output logits (z): ;
[0206] where logits(z) represents the original output of the classification head (a linear layer); represents a logits vector; represents a weight matrix (hidden_dim × C); represents a bias term. represents the refined, high-dimensional numerical representation output by the Transformer encoder.
[0207] 3. Calculate prediction probability;
[0208] Apply softmax function to logits to get prediction probability distribution, specifically:
[0209]
[0210] where, represents the softmax probability of the i-th class predicted by the model; represents the original output score of the i-th class (class) in the logits vector z (i=1,2,3,…,C).
[0211] 4. Calculate loss function: use cross-entropy loss to quantify the difference between prediction and true label.
[0212] Cross-entropy loss function (optimization objective) is specifically:
[0213]
[0214] where, represents one-hot encoding of true labels (1 for the i-th class, 0 otherwise). represents softmax probability of the i-th class predicted by the model. C represents the number of classes. This formula measures distribution difference, suitable for classification tasks, and promotes model learning of intrusion patterns.
[0215] 5. Optimization update: use Adam optimizer to calculate gradient, backpropagation to update model weights (including Transformer parameters and classification head).
[0216] Adam optimizer update (example simplified formula for gradient descent):
[0217] Adam combines momentum and adaptive learning rate, with update rule:
[0218]
[0219] where, represents model weights (including attention head, hidden layer). represents learning rate (e.g. 1e-3). 、 represents first and second order momentum estimates, respectively. represents a small constant (1e-8). represents the value of the model weight at the t-th iteration (or time step).
[0220] Adam is specified in this embodiment to efficiently handle resource-constrained devices.
[0221] Iterate through the local dataset until the local epoch ends or the loss converges (e.g. fix 5-10 epochs to adapt to resource-constrained devices).
[0222] 6. Monitoring and termination:
[0223] Evaluate the loss after each mini-batch (small batch) to ensure efficient operation on edge devices (small batch_size, such as 32, to avoid memory overflow).
[0224] Do not upload raw data, only transmit after encrypting weights at S4.
[0225] Calculation process:
[0226] Initialization: load local dataset (Non-IID data such as traffic logs), set parameters (number of attention heads = 8, hidden dimension = 512, learning rate = 1e-3, batch_size = 32).
[0227] Forward propagation: input batch data, extract features by Transformer → classification head computes logits → softmax gets p → compute Loss.
[0228] Backward propagation: compute gradient based on Loss (∂Loss / ∂θ), Adam adjusts weights.
[0229] Iteration loop: for each epoch (e.g., 5 rounds): iterate through local dataset batches.
[0230] Compute average Loss (monitor decrease).
[0231] S4, encrypted transmission: after the client completes local training, apply Paillier encryption protocol to encrypt model weights (tensor form), then transmit weights to the central server through a secure channel;
[0232] Paillier is a public-key homomorphic encryption scheme that generates public key (n, g) and private key (λ, μ), where n represents the modulus (modulus), which is the product of two large prime numbers p and q, i.e., n = p × q, as part of the public key, it defines the modulus range (mod n²) for encryption operations; g represents the base or generator (base or generator), which is part of the public key, a randomly selected integer, belongs to the multiplicative group of n² under the condition of being coprime with n², usually takes the value of n+1 or other integers that ensure the order is a multiple of n, used to encode plaintext information and support homomorphic properties; λ represents the Carmichael function value (Carmichael's totient function), which is lcm(p-1, q-1) (lcm represents the least common multiple of p-1 and q-1), as part of the private key, used in the decryption process and to ensure the security of the scheme; μ represents the modular inverse, which is the inverse of n under the modulus, i.e.,
[0233]
[0234]
[0235] as part of the private key, used to recover plaintext from ciphertext; n represents the modulus; L represents the modulus operation, i.e., the modulus n² operation on the encrypted number. represents the base g raised to the power of.
[0236] The encryption process is: for weight m, the specific formula for computing ciphertext c is:
[0237] ;
[0238] where r is a random number; denotes plaintext message, in the context of this application specifically refers to model weight, i.e. the numerical data to be encrypted; n denotes modulus; g denotes base or generator; denotes m-th power of base g, used to encode plaintext information m (model weight). denotes n-th power of random number r, used to introduce randomness, enhance the security of encryption (probabilistic encryption).
[0239] The protocol supports additive homomorphism and scalar multiplication homomorphism, which facilitates subsequent aggregation. The additive homomorphism is:
[0240] ,
[0241] where a denotes a plaintext value, in the demonstration of Paillier encryption additive homomorphism, it is the numerical value or message (e.g. part of the model weight) to be encrypted, used to demonstrate that the product of two encrypted values corresponds to the operation of plaintext addition; b denotes another plaintext value, similar to a, as the plaintext of the second encrypted input in the additive homomorphism, used to calculate .
[0242] Paillier encryption supports scalar multiplication homomorphism, i.e. for plaintext scalar k (constant integer) and encrypted ciphertext Enc(m), multiplication in the encrypted state can be achieved through power operation without decryption. This property plays an important role in federated learning aggregation, for example, when calculating weighted contribution , it is directly performed in the encrypted domain.
[0243] The calculation formula is:
[0244]
[0245] where k denotes scalar constant, m denotes plaintext message. Enc(m) denotes Paillier encryption ciphertext of m. n denotes modulus (as defined by public key).
[0246] Calculation process (using square-multiply algorithm to efficiently implement power operation, avoiding direct calculation of large exponent):
[0247] Step 1, Initialization: Take the ciphertext c = Enc(m), the result res = 1 (corresponding to Enc(0), but actually the unit element).
[0248] Step 2, Binary decomposition: Convert the scalar k to binary representation.
[0249] Step 3, Iterative calculation:
[0250] From the highest bit of k, traverse the binary bits.
[0251] Each time: .
[0252] If the current bit is 1, then (multiply Enc(m)).
[0253] Step 4, Output: ; k represents the scalar constant; m represents the plaintext message.
[0254] Efficiency considerations: In actual implementation, use a large integer library to speed up; for floating-point k1, first quantize (such as k1*10 d Convert to integer, then dequantize) ; The process is applied when the central server aggregates, ensuring privacy.
[0255] The encrypted weights are transmitted to the central server through a secure channel. In specific implementation, the key length is set to 1024 bits or higher to balance security and computational overhead; before transmission, the weights can be quantized and compressed (such as 8-bit quantization) to reduce data volume.
[0256] S5, Central server aggregation: After the central server receives the encrypted weights of multiple clients, it performs aggregation operations to generate an encrypted global model;
[0257] Using the homomorphic property of Paillier, perform weighted average on encrypted weights, the calculation formula of global weights is:
[0258] ,
[0259] Where, represents the client weight coefficient (based on data volume or performance), realized through homomorphic multiplication and addition. represents the global model weight (global model weights), which is the aggregation result generated by the central server through weighted averaging of multiple client encrypted weights, used as a teacher model for subsequent knowledge distillation; represents the local model weight of the i-th client (local model weights of the i-th client), which is the model parameter encrypted and uploaded to the central server by the client after completing local training. denotes the encryption function, specifically the encryption operation of the Paillier encryption protocol, used for homomorphic encryption of model weights to ensure privacy protection during transmission and aggregation without revealing original weight information.
[0260] This step generates an encrypted global model, avoiding the risk of centralized decryption. The specific algorithm is similar to FedAvg but integrates homomorphic operations: first, initialize the global model, then iteratively accumulate encrypted contributions, and finally output the global model in encrypted form.
[0261] This step mainly describes the use of the homomorphic property of the Paillier encryption protocol to perform weighted average aggregation on the encrypted model weights uploaded by multiple clients, generating an encrypted global model. The core of this process is a variant of the FedAvg (Federated Averaging) algorithm, but integrates homomorphic operations to ensure that the entire aggregation is performed in the encrypted domain, avoiding decryption operations that leak original weight information. The specific steps are:
[0262] Step 1, initialize the global model;
[0263] Global model weights initialized to a zero vector (in the encrypted domain, represents Enc(0), but in practice Enc(0)=1 if g=n+1).
[0264] Formula for all weight elements; denotes the encryption function.
[0265] Process: The server generates a zero encrypted tensor with the same structure as the local model as the starting point. This avoids introducing leakage risks from plaintext initialization.
[0266] Step 2, iteratively accumulate encrypted contributions:
[0267] For each client i (i=1 to N), the server calculates the client's encrypted contribution then adds it to the global model.
[0268] Using Paillier homomorphism, the scalar multiplication calculation formula is:
[0269] ,
[0270] where denotes the encryption function; n denotes the modulus; denotes the client weight coefficient; denotes the local model weight of the i-th client; assume is an integer; if is a floating point, which can be quantized to an integer, e.g. multiplied by a precision factor 10 k .
[0271] Additive accumulation:
[0272]
[0273] where, denotes an encryption function; n denotes a modulus; denotes global model weights; denotes a client weight coefficient; denotes the local model weights of the i-th client.
[0274] If it is a weighted average, it needs to be additionally divided by the total weight , but since the encryption domain does not support division, it can be processed after decryption or adjusted as a pre-normalized value.
[0275] In this embodiment, noise can be introduced to enhance differential privacy, specifically: before accumulation, add Laplace noise to each . The specific formula is:
[0276] ,
[0277]
[0278] ,
[0279] where, denotes an encryption function; denotes noisy or perturbed model weights. Tilde is commonly used to represent an approximation or perturbed value, in the context of differential privacy, it is the result of adding noise η to the original to prevent the original data from being inversely deduced through the weights; denotes the local model weights of the i-th client. denotes a noise term, sampled from a Laplace distribution, used to achieve differential privacy. b denotes a scale parameter, which controls the dispersion of noise (the larger b, the larger the noise, the stronger the privacy protection, but the model accuracy may be slightly reduced). denotes a sensitivity, denotes a constant, = 1.0.
[0280] Noise addition utilizes homomorphism: .
[0281] Process: Loop through N clients, accumulate one by one. The weight is a high-dimensional tensor, so apply homomorphic operations independently on each element (parallelizable).
[0282] Step 3, output the global model in encrypted form:
[0283] After iteration is complete, is the final encrypted global model, directly used in subsequent steps (such as decryption before knowledge distillation, or further operations in encrypted state).
[0284] If normalization is needed, then: (where, represents the global model weight; represents the client weight coefficient; represents the local model weight of the i-th client), the division can be performed after decryption (only server private), or is designed as , but scalar multiplication needs to handle fractions (through fixed-point quantization).
[0285] During aggregation, noise can be introduced (such as differential privacy mechanism, ε=1.0) to further enhance privacy.
[0286] S6, Knowledge Distillation Backfeeding: The central server decrypts the aggregated global model (using the private key) as the teacher model, then applies the knowledge distillation technique to refine knowledge, and encrypts the refined knowledge to backfeed each client for updating the local model; as shown in Figure 9 .
[0287] The distillation process includes generating soft labels: the teacher model calculates the softmax output on the shared pseudo dataset (or the anonymous samples uploaded by the clients), and the temperature parameter t (for example, t=5) softens the probability distribution, specifically:
[0288] ,
[0289] where, represents the softened probability for class i of the soft label, which is the softmax output of the teacher model softened by temperature, used to refine knowledge and backfeed the student model, providing more rich distribution information than hard labels; t represents the temperature parameter (temperature parameter), a positive real number (for example, t=5), used to control the softening degree of the softmax distribution; when t>1, the distribution is smoother (softened), facilitating knowledge transfer; when t=1, it degenerates to standard softmax; denotes the logit score for class j in the logits vector, j denotes the summation index over classes, which goes through all classes.
[0290] Subsequently, the student model (client local model) minimizes the difference with the teacher soft labels through the KL divergence loss, specifically:
[0291] ,
[0292] where LOSS denotes the total loss function, which is the optimization goal of the student model in the knowledge distillation process, and is used to update the client local model by combining the cross-entropy loss and the KL divergence loss to minimize the difference between the student model and the teacher model soft labels, while considering the influence of the real labels; CE denotes cross-entropy, KL denotes divergence, and a denotes a balance factor, denotes the real label. denotes the predicted probability distribution of the student model (student model's predicted probabilities), i.e., the softened probability output by the softmax function of the client local model (student model), which is used for comparison with the real label and the teacher soft label; denotes the soft label probability distribution of the teacher model (teacher model's soft labels), i.e., the softened softmax output of the global model (teacher model) through the temperature parameter t, which provides rich distribution knowledge for the student model optimization. At the same time, the distillable intermediate representation (such as the attention layer output) is used to transfer feature knowledge.
[0293] The refined knowledge (soft labels or intermediate vectors) is encrypted and fed back to each client for updating the local model. This step specifically optimizes the Non-IID distribution and improves the student's generalization through teacher guidance.
[0294] Based on the existing Paillier protocol framework in this application (supporting additive homomorphism and scalar multiplication), and adapting to the characteristics of knowledge distillation output (soft labels are probability vectors, and intermediate vectors are high-dimensional tensors), the specific process of encrypting the refined knowledge is as follows:
[0295] Step 1, quantization processing: soft labels or intermediate vectors are usually floating-point numbers (in the range of 0-1), which need to be quantized to integers to be compatible with Paillier (integer encryption). For example, fixed-point quantization is used: multiply by the precision factor (k=6, to ensure precision). The formula is:
[0296]
[0297] where, denotes the decrypted plaintext message, which is the original m recovered from the ciphertext c (if no perturbation). denotes the i-th component or instance of the query vector in the attention mechanism. k denotes a scalar constant for weight scaling operation.
[0298] Step 2, element-wise encryption: Soft labels / intermediate vectors are tensors, apply Paillier encryption independently to each element, accelerate with parallel computing.
[0299] For each element of the distilled knowledge (quantized soft label value or intermediate vector value):
[0300]
[0301] where m denotes the plaintext message, specifically the model weight, i.e., the numerical data to be encrypted; g denotes the base (public key part, usually g = n + 1). r denotes a random number (from uniform sampling, providing probabilistic security). n denotes the modulus (public key, n = p × q, p and q are large prime numbers). denotes the m-th power of the base g, used to encode the plaintext message m; denotes the n-th power of the random number r, used to introduce randomness.
[0302] Encryption loop: traverse tensor elements, generate random r, calculate c.
[0303] Step 3, package transmission: form encrypted tensor after encryption, distribute through secure channel (TLS encryption).
[0304] Transmission: package the encrypted tensor into serialized format (such as JSON or Protocol Buffers), send through secure channel.
[0305] Overhead consideration: encryption time is proportional to vector dimension, can be optimized by batch processing.
[0306] Step 4, client decryption and integration: the client uses the private key to decrypt, and integrates into the local loss function after dequantization.
[0307] Client integration process: after receiving the encrypted knowledge, the client uses the private key to decrypt, specifically:
[0308]
[0309]
[0310]
[0311]
[0312] Where m represents the plaintext message, specifically the model weights, i.e., the numerical data to be encrypted. λ represents the Carmichael function value; This represents a function used to calculate the least common multiple of two or more integers; p and q represent prime numbers; It represents the modular inverse, which is The inverse modulo n; n represents the modulus. Representing the base g Power of 1. This is part of the private key and is used to recover the plaintext from the ciphertext. n represents the modulus. This represents the ciphertext c. Power of 1.
[0313] Inverse quantization: .
[0314] in, This represents the original value of "knowledge" extracted through knowledge distillation. It is a scaling factor. This represents the decrypted plaintext message. k represents a scalar constant.
[0315] Update the local model: Integrate the decrypted knowledge into the loss function and fine-tune the local Transformer model using gradient descent.
[0316] Process: Receive - Decrypt element by element - Dequantize - Calculate loss - Optimizer update.
[0317] This step is used to update the local model. Specifically, it optimizes the Non-IID distribution and improves student generalization through teacher guidance.
[0318] S7. Iterative optimization: Perform closed-loop iterations of the above steps until the global model converges.
[0319] The above steps form a closed-loop iteration until the global model converges (e.g., based on a validation set accuracy threshold > 95% or a fixed number of rounds).
[0320] In this embodiment, the convergence criterion is:
[0321] Step 1, using the validation set accuracy as the main indicator: after each round of aggregation, the central server evaluates the global model using the shared validation set (anonymous samples). The calculation formula is:
[0322]
[0323]
[0324] where, is the client weight coefficient (based on data volume), is the accuracy of the i-th client on the local validation set; N represents the total number of clients.
[0325] Threshold judgment: if the global accuracy > 95% and there is no significant improvement (such as improvement <0.5%) for 3 consecutive rounds, then converge.
[0326] Step 2, the calculation process of the early stopping mechanism is:
[0327] Initialization: set the patience parameter (patience = 5, consecutive no improvement round threshold) and the best accuracy (best_acc = 0).
[0328] After each iteration: calculate the current global accuracy (acc_current).
[0329] Judgment: if acc_current > best_acc, then update best_acc = acc_current and reset the counter; otherwise, the counter is incremented by 1.
[0330] If the counter > patience, or the number of iterations > max_rounds (such as 100), stop the iteration.
[0331] Step 3, loss convergence supplementary formula: monitor the global loss ;
[0332] where, is the client local loss (cross-entropy); is the client weight coefficient. N represents the total number of clients.
[0333] If (ε = 0.001, for 3 consecutive rounds), then converge.
[0334] where, represents the global loss value calculated at the end of the t-th iteration. represents the global loss. t represents the time step or round of training.
[0335] represents the global loss value calculated at the end of the t-1th iteration (previous round of the tth round). represents the global loss. t-1 represents the previous training round.
[0336] Asynchronous mode extension: In the FedAsync variant, the convergence criterion can be based on the sliding window average accuracy (e.g., the last 10 rounds average > 95%).
[0337] After each iteration, the client fine-tunes the local Transformer model with the knowledge distillation, and the central server monitors the aggregation stability. The specific implementation includes synchronous / asynchronous aggregation mode (e.g., the FedAsync variant) and convergence criterion (e.g., early stopping mechanism).
[0338] The embodiment also provides a privacy protection federated learning system based on Transformer and knowledge distillation, which implements the privacy protection federated learning method based on Transformer and knowledge distillation.
[0339] 1. Application scenarios;
[0340] The application is applied to privacy-sensitive distributed environment products, for example:
[0341] 1) Industrial Internet of Things (IIoT) platform: such as intelligent factory monitoring system, in the chemical or power industry, edge devices (such as sensors) collect flow log or video stream data for real-time intrusion detection. Product example: a software platform named "SecureIIoT Guardian", deployed on the edge server of the factory, without uploading sensitive production data to the cloud.
[0342] 2) Distributed cloud platform: such as multi-cloud management tool, realizing intrusion monitoring between power grid or cloud nodes. Product example: a cloud security service product "FedSecure Cloud", suitable for enterprise-level distributed systems, helping administrators monitor intrusion events across regional nodes while complying with privacy regulations such as GDPR.
[0343] 3) Extended scenarios: can be extended to medical image analysis platforms (processing patient data) or financial transaction monitoring systems, and is applicable to any privacy protection scenario involving Non-IID data distribution and edge computing. This solution solves the data centralization risk of traditional products and improves the applicability on resource-constrained devices (such as low-power sensors).
[0344] Through these scenarios, the product helps users implement a "zero trust" security model in sensitive industries, i.e., local data processing, while ensuring detection accuracy > 95%.
[0345] 2. Functional characteristics;
[0346] The functional features focus on privacy protection, performance optimization, and ease of use, including:
[0347] 1) Privacy protection function: using Paillier encryption protocol to ensure that data is not leaked during model weight transmission. Product features: support homomorphic encryption aggregation, generate global model without decryption; integrate differential privacy mechanism (epsilon = 1.0) to quantify privacy risk. Users can view privacy compliance reports in the product, such as the "Data Leakage Risk Assessment" module showing the current value.
[0348] 2) Intrusion detection efficiency: based on the attention mechanism of Transformer to extract features, combined with knowledge distillation to optimize Non-IID data generalization. Product features: detection accuracy up to 96.5% (such as on the KDD Cup 99 dataset), support real-time alerts (such as anomaly traffic detection); knowledge distillation compresses model size (reduces 50% parameters), suitable for low-power edge device scenarios. Functional modules include "Intrusion Alert Center" to display detection results and confidence.
[0349] 3) Federated learning optimization: iterative closed-loop training, supporting synchronous / asynchronous mode. Product features: automatically handle Non-IID distribution bias, improve system robustness; experiments show that iteration time is shortened by 20%, suitable for large-scale clients (such as hundreds of edge nodes).
[0350] 4) Expansion and integration: the product supports parameter adjustment (such as encryption key length 1024 bits, distillation temperature t = 5), and can integrate third-party APIs (such as cloud service interfaces). Other features: multi-user role support (administrator / operator), log audit and performance dashboard.
[0351] These features distinguish the product from traditional intrusion detection software (such as centralized ML-based tools), emphasizing distributed privacy balance, achieving the core values of "efficient, secure, and easy to expand".
[0352] The privacy-precision trade-off curve of the invention is shown in Figure 10 .
[0353] 3, operation mode;
[0354] The product operation mode adopts modular design, users (such as system administrators or security engineers) configure, train and monitor through graphical interface (GUI). The operation process is divided into three stages of deployment, training iteration and result viewing, emphasizing automation to reduce user threshold. The following explains in combination with the interaction process.
[0355] 1) Deployment stage:
[0356] User logs into the product Web dashboard (main interface: left navigation bar includes "Client Management", "Server Configuration", "Training Settings").
[0357] Operation: Click the "Add Client" button and enter edge device information (e.g., IP address, data type: traffic log or video stream). The product automatically deploys a Transformer-based local model (e.g., choose "BERT variant" or "ViT visual variant").
[0358] Interaction: A configuration wizard UI pops up to set parameters such as attention head count (default 8 heads), hidden dimension (default 512). After confirmation, the system generates a Paillier public / private key pair and embeds a communication module.
[0359] Example UI: A form interface with a drop-down menu to select scenarios (chemical industry / power) and a preview framework structure diagram (showing embedded layers - Encoder Layers - output).
[0360] 2) Training iteration phase:
[0361] Operation: Click the "Start Federated Training" button and the system enters a closed-loop iteration. The client locally extracts features (Transformer processes data, capturing temporal dependencies), and encrypted weights are uploaded to the server.
[0362] Interaction process: Real-time progress bar UI displays each step (e.g., "Local extraction in progress: Client 1 completed 50%"). After server aggregation, knowledge distillation is performed to benefit the client. Users can monitor the dashboard: curve graph shows accuracy changes (target > 95%), and adjust parameters.
[0363] Example interaction diagram: Assume a flowchart UI where users click on nodes to view details, such as clicking on "Knowledge Distillation" to pop up a dialog box showing the loss function; and allow manual intervention (e.g., pause iteration).
[0364] Automation: The system supports early stopping mechanism when the accuracy threshold is reached; users can choose synchronous mode (all clients upload at the same time) or asynchronous (FedAsync variant).
[0365] 3) Results viewing and maintenance phase:
[0366] Operation: After training converges, view the "Detection Report" module, which displays intrusion alerts (e.g., "Abnormal traffic detected, accuracy 96.5%") and performance metrics (iteration time reduced by 20%).
[0367] Interaction: Export logs or reports; if problems are detected, users can restart iterations or adjust encryption key length.
[0368] Example UI: The dashboard home page displays a list of real-time alerts (in table form: time, client, detection type, confidence). Clicking on an alert jumps to a detailed view, including feature visualizations.
[0369] The hardware and software environment of the present invention is set up in a distributed architecture, suitable for resource-constrained edge devices and centralized servers, supporting high concurrency and secure communication. The hardware environment includes multiple client devices (edge nodes) and a central server; the software environment is based on programming languages such as Python, integrating deep learning frameworks (such as PyTorch for Transformer implementation) and encryption libraries (such as python-paillier for Paillier protocol).
[0370] 1. Hardware Environment:
[0371] 1) Client Devices: Edge devices such as IIoT sensor nodes or cloud edge nodes (such as Raspberry Pi or industrial-grade embedded devices). Each client is equipped with CPU / GPU (minimum configuration: 4-core CPU, 2GB RAM) for local model training. Supports multiple clients (N≥2), suitable for distributed scenarios.
[0372] 2) Central Server: High-performance server (such as Intel Xeon processor, ≥16GB RAM, NVIDIA GPU optional), responsible for model aggregation and knowledge distillation. The server needs to support high-bandwidth network interfaces (≥1Gbps) to handle encrypted data transmission.
[0373] 3) Communication Module: Integrated network interface between client and server, supporting TCP / IP protocol. Hardware can use wireless / wired networks (such as 5G or Ethernet) to ensure low-latency transmission.
[0374] 2. Software Environment:
[0375] 1) Operating System: Client uses lightweight Linux (such as Ubuntu ARM version); server uses Ubuntu Server.
[0376] 2) Core Libraries and Frameworks:
[0377] Deep Learning: PyTorch (used for Transformer model implementation, supporting attention mechanism).
[0378] Encryption Protocol: Paillier library (supports homomorphic addition operations, key length ≥1024 bits).
[0379] Federated Learning Framework: Custom implementation based on variants of the FedAvg algorithm.
[0380] Knowledge distillation: integrate Softmax temperature parameter (t=5 by default).
[0381] Data processing: NumPy / Pandas for pre-processing, suitable for Non-IID datasets (e.g., KDD Cup 99).
[0382] 3) Deployment: containerized deployment, easy to scale to cloud platforms.
[0383] The hardware environment of the present application emphasizes edge computing, reducing the risk of data center transmission; the software integrates Transformer (attention mechanism captures long-distance dependencies) and Paillier (homomorphic encryption avoids decryption), achieving privacy-performance balance, suitable for sensitive industries such as chemical industry / power industry.
[0384] As shown in Figure 1 , it is the network topology diagram of the present application; the topology diagram shows a star structure: the client connects to the server through an encrypted channel, avoiding point-to-point leakage.
[0385] The system of the present application includes a front end (client) and a back end (central server), using an iterative optimization mechanism. The front end is responsible for local processing, and the back end is responsible for global coordination. The entire logic is based on the federated learning paradigm, integrating Transformer extraction and knowledge distillation.
[0386] 1. Front-end implementation logic (client):
[0387] 1) Module composition: feature extractor (Transformer-based), local trainer, encrypted transmitter.
[0388] 2) Implementation steps:
[0389] Data input: collect data from local sensors (e.g., flow logs), pre-process into sequence format.
[0390] Feature extraction: use Transformer model (multi-layer Encoder, including Self-Attention and Feed-Forward layers) to capture long-distance dependencies, improve the representation quality of Non-IID data.
[0391] Local training: train intrusion detection model (e.g., classifier) based on extracted features, optimize loss function.
[0392] Encrypted transmission: encrypt model weights using Paillier protocol (public key distributed by server), transmit to server.
[0393] The front end uses Transformer as a feature extractor to address the generalization deficiency of traditional CNNs on Non-IID data; edge devices only require small models, resulting in low computational resource expenditure.
[0394] Background implementation logic (central server):
[0395] 1) Module composition: aggregator, knowledge distiller, and counter-feeding distributor.
[0396] 2) Implementation steps:
[0397] Receive encrypted weights: collected from multiple clients.
[0398] Aggregation: use Paillier homomorphic properties for weighted averaging to generate a global model without decryption.
[0399] Knowledge distillation: the global model serves as a teacher model to generate soft labels (temperature t=5) and refine knowledge.
[0400] Counter-feeding: distribute distilled knowledge back to clients to update local models.
[0401] Background knowledge distillation for fusion and optimization of small model generalization (8% precision improvement); homomorphic encryption ensures privacy and security during aggregation, suitable for high-precision demand (>95%) scenarios.
[0402] As shown in Figure 2 , the data processing flow.
[0403] The data processing flow adopts a closed-loop iteration: from local data collection to global optimization, repeated until convergence (e.g., precision >95% or iterations >50 rounds). The process emphasizes privacy protection: data never leaves the client, only encrypted weights are transmitted.
[0404] 1. Detailed process:
[0405] 1) Local data collection and extraction: clients collect Non-IID data and use Transformer to extract features.
[0406] 2) Training and encryption: after local training, encrypted weights are transmitted.
[0407] 3) Aggregation and distillation: the server aggregates to generate a global model and distills knowledge.
[0408] 4) Counter-feeding and iteration: knowledge is returned to clients to update models.
[0409] 5) Evaluation: in the intrusion detection task, measure precision and privacy.
[0410] As shown in Figure 3As shown, the timing chart demonstrates multi-client parallelism, emphasizing the asynchronous nature of encrypted transmission. There is no data decryption in the timing of the invention, and the privacy risk is minimized; distillation counteracts to improve overall performance.
[0411] Through the above description, the invention realizes efficient and privacy-safe federated learning in technical implementation. Experimental verification on the KDD Cup 99 dataset achieves an accuracy of 96.5%, suitable for IIoT intrusion detection and other scenarios.
[0412] The experimental verification process of the present application aims to evaluate the performance of the system in terms of privacy protection, model performance, and computational efficiency. The experiment uses a simulated distributed environment to simulate the intrusion detection task in the industrial Internet of Things (IIoT) scenario. The following is the detailed experimental process:
[0413] 1. Data set preparation:
[0414] The KDD Cup 99 dataset variant (a classic intrusion detection benchmark dataset containing 41-dimensional features and multiple attack types such as DoS (Denial of Service attack), Probe (probe attack), etc.) is used. To simulate Non-IID (Non-IID) data distribution, the dataset is divided into multiple clients (N=10 clients), and each client's data subset is unevenly distributed in attack type and sample size (for example, client 1 is biased towards DoS attacks, and client 2 is biased towards normal traffic).
[0415] The NSL-KDD dataset is additionally introduced as a supplementary verification to reduce redundant samples and improve generalization testing.
[0416] Data preprocessing: normalize feature values, convert to sequence format (sequence length = 128), and add position encoding. The training / validation / test set ratio is 7:2:1, with a total sample size of about 500,000.
[0417] Prepare a shared pseudo-dataset for knowledge distillation: use a GAN generator to synthesize anonymous samples (about 10,000), avoiding the leakage of real data.
[0418] 2. Experimental environment setup:
[0419] Hardware: Client simulates edge devices (Raspberry Pi 4, 4GB RAM, CPU-only); central server uses NVIDIA RTX 3080 GPU, 16GB RAM.
[0420] Software: PyTorch 1.12 framework implementing Transformer model (hidden dimension = 512, number of attention heads = 8, number of layers = 6); Paillier encryption library (python-paillier, key length = 1024 bits); Optimizer Adam (learning rate = 1e-3); Knowledge distillation temperature t = 5, balancing factor a = 0.5.
[0421] Parameter settings: number of federated learning rounds = 50; local epoch = 5; differential privacy budget e = 1.0 (add Laplace noise); client participation rate = 100% (synchronous mode) or 50% (asynchronous FedAsync variant).
[0422] Baseline comparison: (1) traditional FedAvg (no Transformer, no distillation); (2) FedAvg + Paillier (no distillation); (3) FedAvg + Transformer (no distillation).
[0423] 3. Evaluation indicators:
[0424] Performance: intrusion detection accuracy (Accuracy), recall (Recall), F1 score (F1-Score).
[0425] Privacy: differential privacy budget e (the smaller the privacy is stronger), use privacy auditing tools to quantify the risk of leakage.
[0426] Efficiency: local computing overhead (FLOPs / iteration time, unit: seconds); communication overhead (transferred data volume, unit: MB); overall iteration time.
[0427] Non-IID heterogeneity: use Dirichlet distribution parameter a (balancing factor) to control (a = 0.1 for high heterogeneity, a = 1.0 for low heterogeneity).
[0428] 4. Experimental process:
[0429] Initialization: randomly initialize model weights, generate Paillier key pair.
[0430] Iterative training: execute steps S1-S7. Evaluate the performance of the global model on the shared validation set every 10 rounds.
[0431] Convergence judgment: use early stopping mechanism (patience = 5), if the accuracy improvement is <0.5% for 5 consecutive rounds or the accuracy >95%, stop iteration.
[0432] Repeated experiments: run multiple independent runs, take the average ± standard deviation.
[0433] Extended Test: Repeat under different heterogeneity (a = 0.1, 0.5, 1.0) and privacy budget (e = 0.5, 1.0, 2.0).
[0434] The experimental results are based on the KDD Cup 99 variant dataset, the number of clients N = 10, and the number of rounds = 50. The following is a summary of key data (mean ± standard deviation).
[0435] Results and analysis:
[0436] (1) Model performance comparison;
[0437] The FedTD framework achieved the best performance on the test set, with accuracy, recall rate, and F1 score significantly higher than all baseline models. The specific numerical comparison is shown in Table 1 below:
[0438] Table 1 Model performance comparison
[0439]
[0440] The results show that simple encryption (FedAvg + Paillier) will introduce a slight performance loss. The Transformer structure (FedAvg + Transformer) can significantly improve the model's ability, proving its advantage in capturing long-distance dependencies. The FedTD of the present invention combines the feature extraction capability of Transformer, the privacy protection capability of Paillier, and the generalization enhancement capability of knowledge distillation, achieving the best performance, with an accuracy improvement of more than 8% compared to the basic FedAvg, fully proving its innovation and effectiveness.
[0441] (2) Performance under different data heterogeneity;
[0442] By adjusting the Dirichlet parameter a, different degrees of data heterogeneity are simulated. As the data heterogeneity increases (a value decreases), the performance of all models decreases, but the FedTD framework always maintains the highest accuracy and the strongest robustness.
[0443] High heterogeneity (a = 0.1): FedTD = 94.8% ± 1.0% vs. FedAvg = 82.5% ± 1.5% (improvement of 12.3%).
[0444] Medium heterogeneity (a = 0.5): FedTD = 95.7% ± 0.9% vs. FedAvg = 86.0% ± 1.3% (improvement of 9.7%).
[0445] Low heterogeneity (a = 1.0): FedTD = 96.5% ± 0.8% vs. FedAvg = 88.2% ± 1.2% (improvement of 8.3%).
[0446] Experiments demonstrate the key role of the knowledge distillation mechanism in alleviating the Non-IID data distribution problem. The soft label knowledge provided by the teacher model effectively benefits each client, guiding the local model to learn more generalizable features, so that it can still maintain excellent performance in a highly heterogeneous environment.
[0447] (3) Privacy-utility trade-off analysis;
[0448] There is a natural trade-off between privacy protection strength (measured by differential privacy budget ε) and model utility (accuracy). The performance changes of FedTD under different ε settings are evaluated.
[0449] ε=0.5 (strong privacy): Accuracy=94.2%±1.1%, with extremely low privacy leakage risk (<0.01%).
[0450] ε=1.0 (balance): Accuracy=96.5%±0.8%, with controllable privacy leakage risk (<0.05%).
[0451] ε=2.0 (weak privacy): Accuracy=97.1%±0.7%, with higher privacy leakage risk (<0.1%).
[0452] The results show that by flexibly adjusting the ε value, the framework can provide customizable solutions according to different requirements for privacy and security in actual application scenarios. In most applications requiring privacy protection (ε=1.0), the framework can maintain high detection accuracy while ensuring high security.
[0453] (4) Computational and communication efficiency analysis;
[0454] The FedTD framework has slightly higher FLOPs (floating-point operations per second) per iteration (1.2 GFLOPs) than the CNN-based FedAvg baseline (0.9 GFLOPs) due to the integration of Transformer. However, thanks to the optimization of knowledge distillation for model convergence, the iteration time required by FedTD (0.85 seconds ± 0.1 seconds) is actually shortened by about 20% compared to FedAvg (1.1 seconds ± 0.15 seconds). In terms of communication, Paillier encryption and quantization compression increase the data volume transmitted per round to 15MB, which is an increase compared to the baseline (12MB), but in exchange for strong privacy protection.
[0455] The results prove the efficiency of the FedTD framework. Instead of simply stacking complex modules, it optimizes through component collaboration (such as distillation to accelerate convergence), achieving net gains in computational efficiency while improving performance and protecting privacy, making it very suitable for resource-constrained edge computing environments.
[0456] In summary, the FedTD framework proposed by the present application is significantly superior to the existing federated learning baseline method in terms of privacy protection, model performance and computing efficiency. It not only solves the problem of performance decline caused by Non-IID data, but also provides provable privacy protection through homomorphic encryption and differential privacy, and optimizes the overall efficiency through techniques such as knowledge distillation. The deep integration of the three components of Transformer, Paillier and knowledge distillation produces a significant synergistic effect, verifying the huge application potential of the present application in privacy-sensitive scenarios such as industrial Internet of Things intrusion detection.
[0457] The core content of the present application is a privacy protection federated learning system and method based on Transformer and knowledge distillation, aiming to build a distributed learning framework that balances data security and detection efficiency. The framework is suitable for intrusion detection systems in industrial Internet of Things (IIoT) and distributed cloud platforms. By processing data locally on the client side, encrypting model updates for transmission, and aggregating and knowledge feedback on the central server, it realizes data privacy protection (avoiding original data leakage, complying with GDPR and other regulations), model generalization optimization under Non-IID (non-independent and identically distributed) data, and high-precision intrusion detection (precision > 95%). The overall process includes local feature extraction, encrypted transmission, server aggregation, knowledge distillation feedback and iterative optimization, forming a closed-loop system to solve the bottlenecks of traditional federated learning in privacy, efficiency and computing overhead.
[0458] The present application provides specific improvements to address the shortcomings of existing federated learning frameworks combined with homomorphic encryption:
[0459] 1) Solve the problem of low computing efficiency: The homomorphic encryption of existing solutions leads to high overhead. The present application reduces the local / central computing burden through parallel attention calculation of Transformer (“O(“n 2 ”)” complexity but efficient GPU implementation) and optimization of Paillier (such as batch encryption); knowledge distillation further compresses the model size (reduces parameters by more than 50%), making edge device training faster, and experiments show that the iteration time is shortened by 20% compared with traditional solutions.
[0460] 2) Solve the generalization problem under Non-IID data: The performance of existing technologies decreases significantly. The present application uses Transformer to capture cross-client data dependencies and improve feature robustness; knowledge distillation balances distribution bias through soft knowledge transfer, with an accuracy of 96.5% on the KDD Cup 99 Non-IID dataset, an increase of 8%, solving the bottleneck.
[0461] 3) Solve the problem of insufficient mechanism fusion: the existing scheme lacks deep integration, and the application organically fuses attention mechanism, federated learning and distillation: Transformer provides high-quality input, encryption guarantees safe transmission, and distillation optimizes output to achieve end-to-end balance; In sensitive scenarios such as power grids, the detection accuracy is >95%, the privacy risk is reduced to epsilon=1.0, and it is extended to medical and other fields.
[0462] Through these detailed technical implementations, the application not only solves the root cause of the existing problems, but also ensures the operability and high performance of the system in actual deployment.
[0463] The above only describes the preferred embodiments of the application, and it should be pointed out that for ordinary skilled persons in the art, without departing from the principles of the application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the application.
Claims
1. A privacy-preserving federated learning method based on Transformer and knowledge distillation, characterized in that: The attention mechanism of the Transformer model is included as a core component of local feature extraction to capture long-range dependencies in data and improve the quality of feature representation. The Paillier encryption protocol is introduced to implement homomorphic encryption transmission of model weights. The knowledge distillation technique is used on the central server side to extract soft knowledge from the aggregated global model as a teacher model, which is then fed back to the client side to compress the model and optimize it, including the following steps: S1, data acquisition and preprocessing: data is collected from local sensors and preprocessed into a sequence format; S2, feature extraction: deploy the Transformer model as a local feature extractor on each client device for feature extraction; S3, local model training: train the intrusion detection model based on the extracted features and optimize the loss function; S31, input the output of the self-attention mechanism and the original input, and perform residual connection; S32, perform residual connection on the input of S31 and normalize the residual result; S33, add a feedforward network to the normalized result of S32; S34, perform residual connection and layer normalization on the output of the second linear transformation in S33 and the normalized result in S32; S35, output high-dimensional feature representation; S4, encrypted transmission: after the client completes local training, the model weights are encrypted using the Paillier encryption protocol, and then transmitted to the central server through a secure channel; S5, central server aggregation: after the central server receives the encrypted weights from multiple clients, it performs aggregation operations to generate a global model; S6, knowledge distillation backfeeding: the central server decrypts the aggregated global model as a teacher model, then applies the knowledge distillation technique to extract knowledge, and encrypts the extracted knowledge to feed back to each client for updating the local model; The distillation process includes generating soft labels: the teacher model calculates the softmax output for the shared pseudo data set, and the temperature parameter t softens the probability distribution, specifically: , wherein, denotes the softening probability of the i-th class of soft labels; t denotes a temperature parameter; denotes the score of the j-th class in the logits vector, j denotes the class index in the sum; Then, the student model minimizes the difference with the teacher soft label through KL divergence loss, specifically: , where LOSS denotes the total loss function; CE denotes cross-entropy, KL denotes divergence, and a denotes a balancing factor, denotes a real label; denotes a predicted probability distribution of a student model; denotes a soft label probability distribution of a teacher model; S7, iterative optimization: close-loop iteration is performed on the above steps until the global model converges.
2. The privacy-preserving federated learning method based on Transformer and knowledge distillation according to claim 1, wherein: In S2, the specific implementation process is: The encoder layer of the Transformer model uses the multi-head self-attention mechanism to calculate the dependency weight between input sequences. Specifically, it uses the query, key, and value matrices for dot product attention calculation, specifically: , wherein softmax() represents a normalization function; Q represents a query matrix, K represents a key matrix, denotes a key dimension, and T represents a transpose operation of a matrix.
3. The privacy-preserving federated learning method based on Transformer and knowledge distillation according to claim 1, characterized in that: In S4, the specific implementation process is: The Paillier encryption protocol is applied to encrypt the model weights. The specific formula for calculating the ciphertext c is: , wherein r denotes a random number; denotes a plaintext message; n denotes a modulus; g denotes a base number; denotes an m-th power of the base number g, for encoding the plaintext message m; denotes an n-th power of the random number r, for introducing randomness.
4. The privacy-preserving federated learning method based on Transformer and knowledge distillation according to claim 1, characterized in that: In S5, the specific implementation process is: Using the homomorphic property of Paillier, the encrypted weights are executed for weighted average, and the specific calculation formula is: , wherein, denotes a client weight coefficient; denotes a global model weight; denotes a local model weight of the i-th client; denotes an encryption function; Then initialize the global model, iteratively accumulate the encrypted contribution, and finally output the global model in encrypted form.
5. The privacy-preserving federated learning method based on Transformer and knowledge distillation according to claim 4, characterized in that: The specific content of initializing the global model, iteratively accumulating the encrypted contribution, and finally outputting the global model in encrypted form is: Step 1, Initialize global model: global model weights initialized to zero vector; Step 2, Iterative Accumulation of Encryption Contribution: For each client, the server calculates the encryption contribution of the client, then adds noise to the encryption contribution, and finally accumulates to the global model; Step 3, Output the global model in encrypted form.
6. A privacy-preserving federated learning system based on Transformer and knowledge distillation, implementing the privacy-preserving federated learning method based on Transformer and knowledge distillation in any one of claims 1 to 5, characterized in that: Including client and central server; The client includes local sensor, feature extractor, local trainer, encryption transmitter; Collect raw data from the local sensor and preprocess it into sequence format; Use the feature extractor to extract features from the preprocessed data; use the local trainer to train the intrusion detection model based on the extracted high-dimensional features and optimize the loss function; finally, use the encryption transmitter to apply the Paillier encryption protocol for homomorphic encryption of the trained model weights and transmit them to the central server through a secure channel; The central server includes aggregator, knowledge distiller, and counter-feeding distributor. Receive encrypted weights from multiple clients, use the aggregator to perform weighted average aggregation on encrypted weights using Paillier homomorphic properties, generate global model in encrypted form; use the knowledge distiller to decrypt the aggregated global model as a teacher model, apply knowledge distillation techniques to shared pseudo data sets or anonymous samples, generate soft labels and intermediate representations, refine knowledge to improve model generalization and compression; use the counter-feeding distributor to apply the refined knowledge to the Paillier protocol after encryption, and distribute it back to each client through a secure channel to update the local model and optimize performance under Non-IID data.
Citation Information
Patent Citations
Transform-federated learning-knowledge distillation fused network attack detection method
CN118353654A
Method for constructing a vehicle networking intrusion detection model based on federated learning
US12368738B1
Federated large codeword model deep learning architecture with homomorphic compression and encryption
US20250047296A1
Cited By
Multi-algorithm fusion big data analysis method and system based on integrated collaboration
CN121301823A
Double-stage pre-training system for reading of industrial inspection instrument
CN121960652A