Network intrusion detection method and device, computer equipment and storage medium

By using VAE-GAN model and blockchain technology in the Internet of Things network, the problem of insufficient recognition capabilities of complex attack forms in the existing technology is solved, more efficient and accurate intrusion detection is achieved, and data security is enhanced.

CN120018138AActive Publication Date: 2025-05-16ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510187272.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-16
Estimated Expiration
2045-02-20

AI Technical Summary

Technical Problem

The existing IoT network intrusion detection technology lacks the ability to generalize complex and changeable attack forms, making it difficult to effectively identify and defend against unknown attack types, especially when facing distributed attacks such as DDoS.

Method used

A network intrusion detection method is adopted to obtain device data of IoT devices in a wireless sensor network, determine the optimized feature set, and train the VAE-GAN model composed of a variational autoencoder and a generation adversarial network based on the feature set. Send the trained local model parameters to the central server for parameter aggregation to form global model parameters, and authenticate and distribute them through the blockchain network to improve the accuracy of intrusion detection.

Benefits of technology

By optimizing the joint training of feature sets and VAE-GAN models, the accuracy of abnormal detection of IoT devices is improved, the accuracy of intrusion detection for unknown attack types is enhanced, and the security and credibility of data are ensured through blockchain authentication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120018138A_ABST
    Figure CN120018138A_ABST
Patent Text Reader

Abstract

The invention relates to a network intrusion detection method and device, computer equipment, a storage medium and a computer program product, and relates to the technical field of network security. The method comprises the following steps: acquiring equipment data of each piece of Internet of Things equipment in a wireless sensor network; determining an optimized feature set corresponding to the equipment data, and training a variational auto-encoder (VAE) and a generative adversarial network (GAN) model based on the optimized feature set to obtain local model parameters corresponding to the trained VAE-GAN model; sending the local model parameters to a central server, so that the central server performs parameter aggregation on the local model parameters corresponding to the node devices to obtain global model parameters, and distributing the global model parameters to the node devices and the block chain network; and acquiring the authenticated global model parameter in the block chain network, detecting the to-be-detected equipment data in the wireless sensor network according to the global model adopting the global model parameter, and determining a detection result of the to-be-detected equipment data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and in particular to a network intrusion detection method, apparatus, computer equipment, storage medium and computer program product. Background Art

[0002] IoT security technology is a comprehensive technology that provides a series of security measures such as data privacy protection, device authentication, intrusion detection, and defense mechanisms for the IoT (Internet of Things) ecosystem, aiming to cope with the diverse security threats brought about by device interconnection and data sharing. With the widespread application of IoT technology, IoT devices and networks are facing increasingly complex security threats, such as data tampering, distributed denial of service attacks, and malicious code injection.

[0003] Relevant technologies can accurately identify abnormal behaviors and potential attacks in IoT networks through innovative defense and detection mechanisms, reducing the risk of data leakage and network interruption. At present, most of the relevant technologies focus on intrusion detection mechanisms, attack and defense game strategies, distributed collaborative defense, feature extraction and identification, lightweight security protection, etc. Regarding intrusion detection technology in the field of IoT security, the existing detection algorithms have insufficient generalization capabilities for complex and changing attack forms, and it is difficult to effectively identify and defend against unknown attack types, especially when facing distributed attacks such as DDoS. The detection effect is poor, resulting in insufficient accuracy in intrusion detection of unknown attack types. Summary of the invention

[0004] Based on this, it is necessary to provide a network intrusion detection method, device, computer equipment, computer-readable storage medium and computer program product that can improve the accuracy of anomaly detection of IoT devices in response to the above technical problems.

[0005] In a first aspect, the present application provides a network intrusion detection method. The method comprises:

[0006] Acquire device data of an Internet of Things device in a wireless sensor network that is communicatively connected to the node device; the device data includes radio frequency data, fuel value, power value, speed, temperature, and transaction data;

[0007] Determine an optimized feature set corresponding to the device data, and train a VAE-GAN model composed of a variational autoencoder and a generative adversarial network based on the optimized feature set to obtain local model parameters corresponding to the trained VAE-GAN model; the local model parameters comply with the global aggregation interface specification; the optimized feature set is obtained by concatenating feature vectors of the device data through a self-attention module;

[0008] Send the local model parameters to the central server, so that the central server can aggregate the local model parameters corresponding to the multiple node devices to obtain global model parameters, and distribute the global model parameters to each node device and the blockchain network, so that each node device can retrain the local VAE-GAN model, and the blockchain network can authenticate the global model parameters and add them to the blockchain network;

[0009] The authenticated global model parameters are obtained from the blockchain network, and the device data to be detected of the IoT device to which the node device is connected for communication is detected according to the global model using the global model parameters, so as to determine the detection result of the device data to be detected; the global model is a VAE-GAN model jointly trained by a plurality of the node devices and the central server.

[0010] In one embodiment, the obtaining of device data of an Internet of Things device in a wireless sensor network that is communicatively connected to the node device includes:

[0011] Receive device data collected by each IoT device through a sensor; the device data is encrypted by AES256 and bidirectionally authenticated by a device public key and a private key; the device data is obtained in a structured format by collecting data through the sensor of the IoT device according to the data collection frequency corresponding to the device type of the IoT device; the device data in a structured format includes a device ID, a data type, a value, and a timestamp;

[0012] Performing data reception verification and data cleaning processing on the device data to obtain cleaned device data; the data reception verification includes data field integrity verification, timestamp continuity verification and data packet checksum verification; the data cleaning processing includes missing value interpolation, outlier value replacement, data standardization and data encoding;

[0013] The cleaned device data is stored in a hierarchical storage structure; the hierarchical storage structure includes a storage structure partitioned by day or device.

[0014] In one embodiment, the method further comprises:

[0015] Determine the model structure and parameter configuration of the global model; the input layer in the model structure is dynamically adjusted according to the feature dimension of the standardized data; the hidden layer in the model structure is a three-layer fully connected layer; the bottleneck layer in the model structure contains hidden variables, mean vectors and variance vectors; the hidden layer and bottleneck layer in the model structure use LeakyReLU activation function; the self-attention module includes a multi-head attention mechanism; the parameter configuration initializes the encoder and decoder weights through He initialization, uses Xavier initialization for the self-attention module weights, and initializes the bias term to 0; the loss function of the parameter configuration is L=λ1*L recon +λ2*L KL +λ3*L adv ;

[0016] Among them, the reconstruction loss L recon Using mean square error, KL divergence loss L KL Used to control the distribution of latent variables and the adversarial loss L adv The Wasserstein distance is used, and λ1, λ2, and λ3 are weight coefficients.

[0017] In one embodiment, determining the optimized feature set corresponding to the device data, and training a VAE-GAN model composed of a variational autoencoder and a generative adversarial network based on the optimized feature set to obtain local model parameters corresponding to the trained VAE-GAN model includes:

[0018] Extracting features of the device data based on the self-attention module to obtain feature vectors corresponding to the device data, and concatenating the feature vectors to obtain a unified feature vector;

[0019] The unified feature vector is weighted by a multi-head attention mechanism to obtain a feature matrix; the feature matrix is ​​added to a position encoding matrix corresponding to the relative position encoding to obtain an optimized feature set;

[0020] The pre-constructed VAE is trained by the optimized feature set until the loss value of the VAE satisfies the first training end condition, thereby obtaining a trained VAE, and a potential feature set is determined by the trained VAE and the optimized feature set;

[0021] Training the pre-constructed GAN by using the potential feature set until the discriminator loss value corresponding to the GAN satisfies the second training condition, thereby obtaining a trained GAN, and determining an enhanced feature set by using the trained GAN and the potential feature set;

[0022] According to the enhanced feature set, the parameters corresponding to the self-attention module, the trained VAE and the trained GAN are determined respectively, and the local model parameters are determined according to the global aggregation interface specification.

[0023] In one embodiment, sending the local model parameters to a central server includes:

[0024] After the local model parameters are sent to the central server for the central server to receive and verify the local model parameters corresponding to the plurality of node devices, the node weight corresponding to each node device is determined according to the number of samples in the node device, and the local model parameters are aggregated through the weighted average strategy and the node weight to obtain the global model parameters.

[0025] In one embodiment, obtaining authenticated global model parameters from the blockchain network includes:

[0026] Receiving authenticated global model parameters returned by the blockchain network;

[0027] The blockchain network is used to determine the set of transaction data blocks corresponding to the global model parameters, and broadcast each transaction block in the transaction data block set to the nodes in the blockchain network; when the nodes pass the trusted node screening and the trusted nodes complete the verification of the transaction blocks, the successfully verified transaction blocks are added to the blockchain network.

[0028] In one embodiment, the detecting of the device data to be detected of the IoT device to which the node device is connected in communication according to the global model using the global model parameters, and determining the detection result of the device data to be detected includes:

[0029] Loading the global model parameters obtained from the blockchain network into the VAE-GAN model to obtain a global model;

[0030] Determine a feature vector corresponding to the data of the device to be detected, and input the feature vector into the global model to obtain a predicted value corresponding to the data of the device to be detected;

[0031] According to the predicted value, a detection result of the device to be detected is determined; the detection result includes normal behavior and abnormal behavior.

[0032] In a second aspect, the present application also provides a network intrusion detection device. The device comprises:

[0033] A data acquisition module, used to obtain device data of an Internet of Things device in a wireless sensor network that is communicatively connected to the node device; the device data includes radio frequency data, fuel value, power value, speed, temperature and transaction data;

[0034] A local model training module, used to determine an optimized feature set corresponding to the device data, and train a VAE-GAN model composed of a variational autoencoder and a generative adversarial network based on the optimized feature set to obtain local model parameters corresponding to the trained VAE-GAN model; the local model parameters comply with the global aggregation interface specification; the optimized feature set is obtained by concatenating feature vectors of the device data through a self-attention module;

[0035] A parameter sending module, used for sending the local model parameters to a central server, so that the central server can aggregate the local model parameters corresponding to the multiple node devices to obtain global model parameters, and distribute the global model parameters to each node device and the blockchain network, so that each node device can retrain the local VAE-GAN model, and the blockchain network can authenticate the global model parameters and add them to the blockchain network;

[0036] An intrusion detection module is used to obtain authenticated global model parameters from the blockchain network, detect the device data to be detected of the IoT device to which the node device is connected in communication according to a global model using the global model parameters, and determine the detection result of the device data to be detected; the global model is a VAE-GAN model jointly trained by multiple node devices and the central server.

[0037] In a third aspect, the present application further provides a computer device, wherein the computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the method described in the first aspect are implemented.

[0038] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0039] In a fifth aspect, the present application further provides a computer program product, wherein the computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.

[0040] The above-mentioned network intrusion detection method, device, computer equipment, storage medium and computer program product obtain the device data of each IoT device in the wireless sensor network, determine the optimized feature set corresponding to the device data, and train the VAE-GAN model based on the optimized feature set to obtain the local model parameters corresponding to the trained VAE-GAN model. Since the optimized feature set is obtained by splicing the feature vectors of the device data through the self-attention module, the optimized feature set can represent the characteristics of various types of device data. The VAE-GAN model obtained by training the optimized feature set can improve the generalization ability of the global model, and when facing diversified intrusion attacks, it can improve the prediction accuracy of the local model parameters corresponding to the global model. In addition, the node device can send the local model parameters to the central server so that the central server aggregates the local model parameters of multiple node devices to obtain the global model parameters, and distributes the global model parameters to the blockchain network for data authentication to obtain the authenticated global model parameters. Based on this, each node device can load the global model parameters into the global model, and predict the detection results of the device data to be detected through the global model. Since the global model parameters are obtained by aggregating the local model parameters, and the local model parameters are learned from the device data corresponding to different node devices, they can detect multiple types of device data. The aggregated global model parameters further improve the detection accuracy and robustness of the global model, thereby improving the accuracy of intrusion detection of unknown attack types. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the related technologies, the drawings required for use in the embodiments or the related technical descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0042] Figure 1 An application environment diagram of a network intrusion detection method in an embodiment;

[0043] Figure 2 A schematic diagram of a flow chart of a network intrusion detection method in an embodiment;

[0044] Figure 3 A schematic diagram of the overall structure of a network intrusion detection system in one embodiment;

[0045] Figure 4 A schematic diagram of the structure of a federated learning architecture of a self-attention time-varying VAE-GAN in one embodiment;

[0046] Figure 5is a flowchart of a network intrusion detection method in another embodiment;

[0047] Figure 6 A logarithmic distribution diagram of a confusion matrix of a model prediction result in one embodiment;

[0048] Figure 7 A comparative analysis diagram of performance evaluation between federated learning and non-federated learning in one embodiment;

[0049] Figure 8 A schematic diagram of comparative analysis of architecture performance based on federated learning in one embodiment;

[0050] Fig. 9 is a structural block diagram of a network intrusion detection device in one embodiment;

[0051] Fig.10 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0053] The network intrusion detection method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the node device 102 communicates with the central server 104 through the network. The node device 102 can obtain the device data of each IoT device in the wireless sensor network, and splice the feature vectors of the device data through the self-attention module to obtain an optimized feature set. The node device 102 can train the variational autoencoder VAE (Variational Auto-Encoder) and the generative adversarial network GAN (Generative Adversarial Networks) model by optimizing the feature set to obtain a trained VAE-GAN model, thereby obtaining local model parameters. Each node device 102 can send the local model parameters to the central server 104. After the central server 104 receives the local model parameters of multiple node devices 102, it can aggregate the local model parameters to obtain the global model parameters, and distribute the global model parameters to each node device 102, so that the node device performs multiple rounds of training based on the global model parameters. In addition, the central server 104 can send the global model parameters to the blockchain network, and perform data authentication on the global model parameters, and add the authenticated global model parameters to the blockchain network, so that each node device can subsequently obtain the authenticated global model parameters from the cloud, and load the global model parameters into the global model, and detect the data of the device to be detected in the wireless sensor network through the global model to determine the intrusion detection result. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the central server 104, or it can be placed on the cloud or other network servers. The node device 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted device can be a virtual reality (VR) device, an augmented reality (AR) device, smart glasses, etc. The central server 104 can be implemented with an independent server or a server cluster consisting of multiple servers.

[0054] In an exemplary embodiment, Figure 2 As shown, a network intrusion detection method is provided, which is applied to Figure 1 The node device in is taken as an example to illustrate, including the following steps S100 to S500. Among them:

[0055] Step S100, obtaining device data of an Internet of Things device that is communicatively connected to a node device in a wireless sensor network.

[0056] Among them, device data includes radio frequency data, fuel value, power value, speed, temperature and transaction data.

[0057] Specifically, the node device must first collect and initialize data from each IoT device. The node device is based on the device set {D1, D2, D3, …, D x} The collected environmental data and network behavior data. And with the goal of generating high-quality, standardized data sets, through the operations of device data collection and initialization, data cleaning and standardization, and global model initialization, the cleaned and standardized data sets and the initialized global model G0 are constructed to provide data and model support for feature extraction, model training and parameter aggregation of the subsequent federated learning framework.

[0058] Step S200, determining an optimized feature set corresponding to the device data, and training a VAE-GAN model composed of a variational autoencoder and a generative adversarial network based on the optimized feature set to obtain local model parameters corresponding to the trained VAE-GAN model.

[0059] Among them, the local model parameters comply with the global aggregation interface specifications; the optimized feature set is obtained by concatenating the feature vectors of the device data through the self-attention module.

[0060] Specifically, the node device can train the local model. Based on the cleaned and standardized data set outputted in step S100, with the goal of building a high-quality, localized feature model, through operations such as data loading, verification and partitioning, self-attention feature extraction, variational autoencoder VAE training, generative adversarial network GAN training, and local model parameter generation, an optimized feature set, a potential feature set, an enhanced feature set, and local model parameters in a standard format are constructed to provide support for global model aggregation and federated learning.

[0061] Step S300, the local model parameters are sent to the central server, so that the central server can aggregate the local model parameters corresponding to multiple node devices to obtain global model parameters, and distribute the global model parameters to each node device and the blockchain network, so that each node device can retrain the local VAE-GAN model, and the blockchain network can authenticate the global model parameters and add them to the blockchain network.

[0062] Specifically, the node device may send the local model parameters to the central server. For step S400, the central server may generate a local model parameter set {LM1, LM2, LM3, ..., LM x}, with the goal of generating a unified global model {GM}, through operations such as parameter reception and verification, weight calculation, global model parameter aggregation, model optimization and constraint checking, and global model distribution, a global model parameter set containing model structure information, weight matrix and bias vector, and weight calculation records is constructed. This step is used to improve the generalization ability of the model and support the next round of local model training.

[0063] Optionally, after step S400, the method further includes: sending the global model parameters to the blockchain network, and the blockchain network is based on the global model parameter set {GM}, the IoT application set {I1, I2, I3, ..., I n}, a network consisting of multiple trusted nodes and private keys {Prkey} and public keys {Pubkey}, with the goal of ensuring the data authenticity and security of global model parameters, through transaction block initialization, private key allocation and block broadcasting, trusted node screening, block signature verification and authentication result processing, the authenticated trusted data block set {Tx1, Tx2, Tx3,…, Tx i}, and store the authentication results through the blockchain network to ensure the integrity and credibility of the data.

[0064] Step S500, obtaining authenticated global model parameters from the blockchain network, detecting the device data to be detected of the IoT device to which the node device is connected for communication based on the global model using the global model parameters, and determining the detection result of the device data to be detected.

[0065] Among them, the global model is a VAE-GAN model trained jointly by multiple node devices and a central server.

[0066] Specifically, the node device is based on the authentication global model parameter set {GM} stored in the cloud and the real-time collected IoT data set {R f ,F l ,E l ,S p ,T p ,T x} new , with the goal of identifying potential intrusion behaviors and ensuring the security and stability of the IoT wireless sensor network, through operations such as global model loading, real-time data feature extraction, model prediction, anomaly determination, and result storage and output, a structured record containing data batch identification, device identification, timestamp, prediction probability and detection results is constructed for subsequent analysis and alarm response.

[0067] In the above network intrusion detection method, by obtaining the device data of each IoT device in the wireless sensor network, and determining the optimized feature set corresponding to the device data, and training the VAE-GAN model based on the optimized feature set, the local model parameters corresponding to the trained VAE-GAN model are obtained. Since the optimized feature set is obtained by splicing the feature vectors of the device data through the self-attention module, the optimized feature set can represent the characteristics of various types of device data. The VAE-GAN model trained by the optimized feature set can improve the generalization ability of the global model, and when facing diversified intrusion attacks, it can improve the prediction accuracy of the local model parameters corresponding to the global model. In addition, the node device can send the local model parameters to the central server, so that the central server aggregates the local model parameters of multiple node devices to obtain the global model parameters, and distributes the global model parameters to the blockchain network for data authentication to obtain the authenticated global model parameters. Based on this, each node device can load the global model parameters into the global model, and predict the detection results of the device data to be detected through the global model. Since the global model parameters are obtained by aggregating the local model parameters, and the local model parameters are learned from the device data corresponding to different node devices, they can detect multiple types of device data. The aggregated global model parameters further improve the detection accuracy and robustness of the global model, thereby improving the accuracy of intrusion detection of unknown attack types.

[0068] In an exemplary embodiment, the specific implementation process of the step of “obtaining device data of each IoT device in the wireless sensor network” includes:

[0069] Step S101, receiving device data collected by each IoT device through a sensor.

[0070] Among them, device data is data encrypted by AES256 and bidirectionally authenticated by the device public key and private key; device data is data in a structured format obtained by collecting data through the sensors of the IoT device based on the data collection frequency corresponding to the device type of the IoT device; device data in a structured format includes device ID, data type, value and timestamp.

[0071] Specifically, step S101 is device data collection, and the input includes an IoT device set {D1, D2, D3, ..., D x} and the types of data collected, such as radio frequency data (R f ), fuel value (F l ), power value (E l ), speed (S p ), temperature (T p ), transaction data (T x). The processing process includes: defining the frequency of data collection according to the device type (fixed sensors collect data every 5 minutes, mobile sensors collect data every 1 minute, and key devices collect data in real time and the frequency does not exceed 1 second / time); formatting the collected data into a structured format (device ID, data type, value, timestamp), the timestamp format follows ISO8601, and the accuracy is in milliseconds; setting up an exception handling mechanism, including recording the device offline time and automatic reconnection attempts, caching data locally when communication is interrupted and uploading it in batches after recovery, and marking abnormal data according to the valid value range of the device type; maintaining the collection success / failure counter for each device; using AES256 encryption to achieve data encryption transmission, using the device public key and private key for two-way authentication, and using MD5 to generate a checksum to ensure data integrity. The final output is a structured raw data set {(device ID, data type, value, timestamp)} that records all collected environmental and network behavior data Dx .

[0072] Step S102: perform data reception verification and data cleaning processing on the device data to obtain cleaned device data. The cleaned device data is stored in a hierarchical storage structure.

[0073] Among them, data reception verification includes data field integrity verification, timestamp continuity verification and data packet checksum verification; data cleaning processing includes missing value interpolation, abnormal value replacement, data standardization and data encoding. The hierarchical storage structure includes a storage structure partitioned by day or device.

[0074] Specifically, step S102 mainly performs data cleaning and standardization, and the input is a structured raw data set, including device ID, data type, value and timestamp {(device ID, data type, value, timestamp)} Dx. First, perform data reception verification: generate a unique BatchID (in the format of BATCH-YYYYMMDD-ID), check the field integrity of each record, mark missing fields as abnormal and record logs; check the timestamp continuity by device ID, record the abnormal time period; verify the packet checksum to ensure that the data has not been tampered with. Secondly, perform data cleaning: for missing values, use linear interpolation to fill in data with a continuous missing time of less than 2 hours, fill in discrete data with the mode, and mark data with a continuous missing time of more than 2 hours as abnormal and record logs; for abnormal values, calculate the Z score, and replace mild abnormalities (3<|Z|≤5) with the mean of the data in the past 1 hour, and severe abnormalities (|Z|>5) are marked as abnormal but retain the original value. Next, perform data standardization: use Min-Max to standardize numerical data to the interval [0,1], and use Z-Score to standardize time series data; use One-hot encoding for device type and integer encoding for data type for categorical data. Then, the original data, cleaned data, and abnormal data are stored in a hierarchical storage structure, and the index is optimized to quickly locate batch data. At the same time, the storage strategy of partitioning by day and sub-partitioning by device is adopted. Finally, the BatchID field and version number (for example, "v1.0") are added to the output data structure. The final output is a cleaned and standardized data set {(BatchID, device ID, data type, value, timestamp)} cleaned .

[0075] In this embodiment, data cleaning and standardization can improve the stability and standardization of device data, which is convenient for subsequent model training, reducing the training error rate and improving training efficiency.

[0076] In an exemplary embodiment, the method further comprises:

[0077] Determine the model structure and parameter configuration of the global model.

[0078] Among them, the input layer in the model structure is dynamically adjusted according to the feature dimension of the standardized data; the hidden layer in the model structure is a three-layer fully connected layer; the bottleneck layer in the model structure contains hidden variables, mean vectors and variance vectors; the hidden layer and bottleneck layer in the model structure use the LeakyReLU activation function; the self-attention module includes a multi-head attention mechanism; the parameter configuration initializes the encoder and decoder weights through He initialization, uses Xavier initialization for the self-attention module weights, and initializes the bias term to 0; the loss function of the parameter configuration is L=λ1*L recon +λ2*L KL +λ3*L adv ;

[0079] Among them, the reconstruction loss L reconUsing mean square error, KL divergence loss L KL Used to control the distribution of latent variables and the adversarial loss L adv The Wasserstein distance is used, and λ1, λ2, and λ3 are weight coefficients.

[0080] Specifically, the purpose of global model initialization is to complete the initialization of the global model through model structure definition and parameter initialization rules, to ensure that the federated learning framework provides a unified initial model structure and parameter configuration, and to support the consistency of subsequent feature extraction, training, and parameter aggregation processes. The processing process includes: model structure definition, the input layer of the encoder structure is dynamically adjusted according to the feature dimension of the standardized data, and the hidden layer is designed with three fully connected layers, with the number of nodes being n1=d modle ×2, The bottleneck layer generates hidden variables z and outputs the mean vector μ and variance vector σ 2 The decoder structure adopts a symmetrical design, the number of hidden layer nodes is [n3,n2,n1], the output layer is restored to the original feature dimension, and all hidden layers and bottleneck layers use LeakyReLU activation function. The self-attention module includes a multi-head attention mechanism with n heads. heads The value is {4,8,16}, and the dimension of each head is d head =d modle / n heads , and by the formula Calculate attention and use relative position encoding to enhance time series features. Parameter initialization uses He initialization to initialize the encoder and decoder weights, self-attention module weights use Xavier initialization, and the bias term is initialized to 0. The configured loss function L = λ1·L recon +λ2·L KL +λ3·L adv is, where the reconstruction loss L recon Use mean square error (MSE), KL divergence loss L KL Controlling the distribution of latent variables and adversarial loss L adv The Wasserstein distance is used, and the weight coefficients are λ1=1.0, λ2=0.1, and λ3=0.01. The final output is the initialized global model G0, which includes the model architecture and parameter configuration, for each node to download for local training.

[0081] In this embodiment, by initializing the global model, it is possible to ensure that the federated learning framework has a unified initial model result and parameter configuration, thereby improving the efficiency and consistency of feature extraction, training, and parameter aggregation of subsequent device nodes and central servers.

[0082] In an exemplary embodiment, the specific implementation process of the step of “determining an optimized feature set corresponding to the device data, and training a VAE-GAN model composed of a variational autoencoder and a generative adversarial network based on the optimized feature set to obtain local model parameters corresponding to the trained VAE-GAN model” includes:

[0083] Step S202, extract features from the device data based on the self-attention module to obtain feature vectors corresponding to the device data, and concatenate the feature vectors to obtain a unified feature vector; perform weighted processing on the unified feature vector through a multi-head attention mechanism to obtain a feature matrix; add the feature matrix to the position coding matrix corresponding to the relative position coding to obtain an optimized feature set.

[0084] Specifically, the purpose of self-attention feature extraction is to extract important features through the self-attention mechanism and enhance the correlation and temporal dependency between features. The input is the segmented feature data set {(BatchID, device ID, data type, value, timestamp, data set type)} output by the subsequent step S201 split The processing process includes: first, feature generation is performed. For radio frequency data, fuel value, power value, speed, temperature and transaction data, features related to their characteristics are extracted, such as frequency domain features, volatility, cumulative energy consumption, acceleration, temperature range and transaction frequency, and all features are spliced ​​into a unified feature vector X fusion Then, the features are weighted through the multi-head attention mechanism, and the feature vector is divided into multiple subspaces. The query matrix Q, key matrix K and value matrix V are calculated for each subspace respectively, and the scaled dot product formula is used Calculate the attention score and concatenate the weighted features of all heads to generate the matrix Z. Next, inject the relative position code, generate the position code matrix PE through the sine-cosine function, and add it to the output feature matrix Z of the attention mechanism to generate the feature vector X containing the timing information fusion′ =Z+PE. Finally, for the eigenvector X fusion′ = Z + PE to normalize and ensure that the eigenvalues ​​have a uniform scale. The final output is the optimized feature set

[0085] Step S203, training the pre-constructed VAE by optimizing the feature set until the loss value of the VAE meets the first training end condition, thereby obtaining a trained VAE, and determining a potential feature set by using the trained VAE and the optimized feature set.

[0086] Specifically, the variational autoencoder training aims to use the variational autoencoder VAE to extract the latent features x latent, to provide support for the subsequent generation of the GAN model, by optimizing the distribution of the latent space and the quality of feature reconstruction, to ensure the effectiveness and consistency of the latent features. The input is the optimized feature set output by step S202 The processing process includes: first, the model structure is designed, and the encoder generates the mean vector μμ and the logarithmic variance vector logσ of the latent space according to the optimized feature vector 2 , sampling latent features x through the reparameterization technique latent =μ+σ·∈,∈~N(0,I); the decoder reconstructs the latent features into feature vectors Both the encoder and decoder use fully connected networks, and the number of nodes in the hidden layer is and The latent space dimension is calculated based on the feature compression ratio, which defaults to 0.25. Then the loss function is optimized, including the reconstruction loss and KL divergence loss The final total loss is Where β = 1. During the training process, the model parameters are initialized with a batch size of B = 64, a learning rate of η = 0.001, and a maximum number of iterations N. epochs = 100, use Adam optimizer to optimize model parameters, and set early stopping strategy, stop training when the validation set loss does not decrease for 10 consecutive rounds. The final output potential feature set {(BatchID, device ID, time window, x latent , dataset type)} latent .

[0087] Step S204, training the pre-constructed GAN through the potential feature set until the discriminator loss value corresponding to the GAN meets the second training condition, obtaining a trained GAN, and determining an enhanced feature set through the trained GAN and the potential feature set.

[0088] Specifically, the purpose of generative adversarial network training is to use the generative adversarial network GAN model to generate more diverse feature samples and improve the robustness and generalization ability of the model. The input includes the potential feature set {(BatchID, device ID, time window, x latent , dataset type)} latent and potential dimensions The processing process includes: first, design the GAN model, the generator uses random noise n~N(0,I) and potential features x latent As input, the feature x is generated through two layers of ReLU activated hidden layers gan , the output layer uses Tanh activation; the discriminator uses the real feature x′ latent and generate features x ganAs input, the true and false classification is performed through two layers of LeakyReLU activated hidden layers, and the output layer uses Sigmoid activation. Then the loss function is designed, using Wasserstein distance and gradient penalty (WGANGP), the discriminator loss Calculate the distribution difference between real samples, generated samples and linear interpolation samples, and add the gradient penalty term, the generator loss Maximize the probability that the generated sample is judged as real. The training process includes initializing the generator and discriminator weights, setting the generator learning rate η G =0.0001, discriminator learning rate η D =0.0004, batch size B = 64, using Adam optimizer and proportional n D :n G =5:1 Iteratively update the weights. In each iteration, the generator is fixed to train the discriminator, and then the discriminator is fixed to train the generator, minimizing the corresponding loss function, and finally generating enhanced features. The final output is the enhanced feature set {(BatchID, device ID, time window, x gan , dataset type)} gan .

[0089] Step S205, according to the enhanced feature set, determine the parameters corresponding to the self-attention module, the trained VAE and the trained GAN respectively, and determine the local model parameters according to the global aggregation interface specification.

[0090] Specifically, the purpose of local model parameter generation is to generate local model parameters in a standard format based on the trained model to provide input for global model aggregation (step S300). Through a clear parameter extraction and standardization process, seamless connection with the global aggregation interface is ensured, and a parameter importance evaluation mechanism is added to improve model quality. The input is the enhanced feature set {(BatchID, device ID, time window, x gan , dataset type)} gan The processing process includes: firstly, parameter extraction, and then the attention weight matrix is ​​extracted from the attention layer. VAE extracts encoder weights W enc , bias b enc and decoder weight W dec , bias b dec , GAN extracts the generator weight matrix W gen and the discriminator weight matrix W disc Then the parameters are normalized and the hierarchical representation of the unified parameters is P layer ={W,b,activation}, calculate the importance weight of the parameter where ∥P i ∥ Fis the Frobenius norm, Var(P i ) is the variance, and the parameters of different layers are normalized to eliminate scale differences. Finally, parameter verification is performed to check the integrity and consistency of each layer parameter to ensure that the parameter format complies with the global aggregation interface specification (step S300). Finally, the local model parameter set {LM1, LM2, LM3, …, LM x}, each local model parameter contains the model architecture description, hierarchical parameter matrix and parameter importance weight, and is guaranteed to be consistent with the global model aggregation rules, providing support for the efficient construction of the global model.

[0091] Optionally, before step S202, step S201 is further included, wherein:

[0092] Step S201 is data loading, verification and partitioning. The purpose is to load the cleaned and standardized data set in step S102, verify its integrity and consistency, and partition it into training set, verification set and test set according to the model training requirements, while evaluating the quality and balance of the partitioned data. The input is the cleaned and standardized data set {(BatchID, device ID, data type, value, timestamp)} cleaned . The processing process includes: first, verify the input data, check the field name, data type and timestamp format of the data, ensure that the timestamp is legal and continuous, meet the time coverage requirements, check the legality of key data types and value ranges, and count the missing value ratio to ensure that the missing rate is lower than the set threshold (such as 2%), and use interpolation to process missing values. Then the data set is split, with the training set accounting for 80%, the validation set accounting for 10%, and the test set accounting for 10%, and ensure the balance of device data, and the deviation of the proportion of each device in each data set does not exceed 5%. Finally, evaluate the quality of the split data, use Kolmogorov-Smirnov (KS) to test the distribution consistency of the training set, validation set and test set, and calculate the time coverage to ensure the integrity of the time range of the split data. The final output is the split data set {(BatchID, device ID, data type, value, timestamp, data set type)} split , where the data set types include train (training set), valid (validation set) and test (test set).

[0093] In this embodiment, the prediction accuracy of the global model can be improved through self-attention feature extraction, VAE training and GAN training. The generated local model parameters can represent the features implied by the training data of the node device, and can improve the prediction accuracy of the global model in the node device.

[0094] In an exemplary embodiment, the specific implementation process of the step of "sending the local model parameters to the central server" includes:

[0095] The local model parameters are sent to the central server for the central server to receive and verify the local model parameters corresponding to multiple node devices. Then, the node weight corresponding to each node device is determined according to the number of samples in the node device, and the local model parameters are aggregated through the weighted average strategy and the node weight to obtain the global model parameters.

[0096] Specifically, the node device may send the local model parameters to the central server. The central server may perform the following steps:

[0097] Step S301: parameter reception and verification. The central server receives local model parameters {LM1, LM2, LM3, ..., LM x}, verify the uploaded data, including integrity verification, ensuring that the uploaded parameters include weight matrix, bias vector and activation function; consistency verification, ensuring that the local model parameter structure is consistent with the global model structure, including weight dimension and network hierarchy; data source audit, recording the upload time and participation status of each node to ensure data traceability.

[0098] Step S302: weight calculation, based on the number of samples of each node |D i |, calculate the weight w of weighted aggregation i , the formula is Where |D i | is the number of samples of the ith node, and N is the total number of participating nodes. The calculation of weights reflects the importance of the sample distribution of each node, ensuring that nodes with large sample sizes contribute more to the global model.

[0099] Step S303: Global model parameter aggregation, using the weighted average method to calculate the global model parameters layer by layer, including the weight matrix and the bias vector in is the weight matrix of the ith node, is the bias vector of the i-th node.

[0100] Step S304: Model optimization and constraint checking: Gradient stability check of the aggregated global model parameters, using gradient clipping technology Where C is the gradient threshold (such as 1.0 to 5.0) to avoid gradient vanishing or exploding problems. At the same time, the parameter range is corrected to ensure that all weights and biases are within a reasonable range to prevent outliers from affecting model performance.

[0101] Step S305, global model distribution, encrypt the aggregated global model {GM} and distribute it to each participating node for the next round of local model training. At the same time, add version control information (such as "v1.0") to the global model to support historical backtracking and version management of the model. The final output global model parameter set includes model structure information, weight matrix and bias vector, and weight calculation records to support subsequent analysis and auditing.

[0102] In this embodiment, by sending the local model parameters to the central server, the central server performs weighted aggregation on the local model parameters, thereby improving the accuracy of the global model parameters and improving the prediction accuracy of the global model that applies the global model parameters.

[0103] In an exemplary embodiment, the specific implementation process of the step of "obtaining authenticated global model parameters from the blockchain network" includes:

[0104] Receive the authenticated global model parameters returned by the blockchain network.

[0105] Among them, the blockchain network is used to determine the set of transaction data blocks corresponding to the global model parameters, and broadcast each transaction block in the transaction data block set to the nodes in the blockchain network; when the nodes pass the trusted node screening and the trusted nodes complete the verification of the transaction blocks, the successfully verified transaction blocks are added to the blockchain network.

[0106] Specifically, the node device can receive the authenticated global model parameters returned by the blockchain network. In addition, the specific process of the blockchain network receiving and verifying the global model parameters includes:

[0107] Step S401, initialize the transaction block set, convert the input global model parameter {GM} into a transaction data block set {Tx1, Tx2, Tx3, ..., Tx i}, where each transaction block contains a unique number (Block ID), a global model parameter summary (GM Hash) generated by the SHA256 algorithm, and a timestamp of block generation.

[0108] Step S402, distribute private keys and broadcast blocks, distribute private keys Prkey to each transaction block, sign the block and broadcast it to all nodes in the network, and prepare for trusted node authentication and block verification.

[0109] Step S403, trusted node screening, traverse all nodes in the network, check their credibility, and screen out the trusted node set {Trusted Nodes}. The trusted nodes must meet the trust value threshold (tr≥th, where th=5). If the node fails the screening, the block is rebroadcasted.

[0110] Step S404, block signature verification, uses a trusted node to verify the signature in the transaction block set, calculates the block hash value through the SHA256 algorithm and verifies the matching of the signature and the public key. The transaction block that has been successfully authenticated passes the signature verification to ensure data integrity and authenticity.

[0111] Step S405, authentication result processing, according to the authentication result, the successfully authenticated transaction blocks are added to the blockchain network and broadcast to update the ledger, and the blocks that fail to authenticate will be marked as untrustworthy and discarded, and finally the authenticated data block set {Tx1, Tx2, Tx3, ..., Tx i} and the corresponding authentication result (success / failure).

[0112] In this embodiment, by receiving the authenticated global model parameters returned by the blockchain network, since the global model parameters are verified by the blockchain network, the security and stability of the global model parameters can be improved.

[0113] In an exemplary embodiment, the step of “detecting the device data to be detected of the IoT device to which the node device is connected in communication according to the global model using the global model parameters, and determining the detection result of the device data to be detected” includes:

[0114] The global model parameters obtained from the blockchain network are loaded into the VAE-GAN model to obtain a global model; the feature vector corresponding to the data of the device to be detected is determined, and the feature vector is input into the global model to obtain the predicted value corresponding to the data of the device to be detected; based on the predicted value, the detection result of the device to be detected is determined; the detection result includes normal behavior and abnormal behavior.

[0115] Specifically, the real-time collected IoT data is anomaly detected through the certified global model, and the global model parameters {GM}, including the weight matrix W, are loaded from the cloud. GM and the bias vector b GM , initialize the detection model. f ,F l ,E l ,S p ,T p ,T x} new Perform feature extraction and generate feature vectors Calculate the predicted value y using the global model pred =sigmoid(W GM X new +b GM ), determine whether the data is abnormal based on the anomaly detection threshold τ (such as τ = 0.8): if y pred>τ, it is considered abnormal behavior; otherwise, it is considered normal behavior. The detection results are stored in structured records, including {BatchID, device ID, timestamp, y pred , detection results} for subsequent analysis or alarm.

[0116] In this embodiment, anomaly detection is performed on IoT data through a global model, which can improve the efficiency and accuracy of anomaly detection and enhance the security and stability of IoT devices.

[0117] The following is a detailed description of the specific implementation process of the above-mentioned network intrusion detection method in conjunction with a specific embodiment. The embodiment of the present application provides a wireless sensor network intrusion detection method based on a self-attention federated learning architecture, which aims to achieve efficient and accurate intrusion detection while enhancing privacy protection and system security. By designing a layered distributed architecture, combining the self-attention mechanism with the time-varying VAE-GAN model to improve detection accuracy and robustness, and using the PoAh-based blockchain consensus mechanism to reduce authentication overhead, it can effectively resist complex attack forms such as DDoS attacks in a resource-constrained wireless sensor network environment, ensuring the stability and reliability of IoT applications.

[0118] Step S1, data acquisition and initialization.

[0119] Step S101, device data collection. Purpose: Collect environmental data and network behavior data from devices in the wireless sensor network to generate a complete and structured set of raw data to provide data support for subsequent processing. Input: IoT device set:; Device type and the type of data it collects: radio frequency data, fuel value, power value, speed, temperature, transaction data. Processing process: 1. Definition of data collection frequency. Fixed sensors: collect once every 5 minutes; Mobile sensors: collect once every 1 minute; Key equipment (such as transformer monitoring equipment): real-time collection, with a frequency not exceeding 1 second / time. 2. Data packaging and formatting. Structured format: (device ID, data type, value, timestamp); Timestamp format: follow ISO8601; Timestamp accuracy: millisecond level. 3. Abnormal situation handling. Device offline: record the last time the device was online; set up an automatic reconnection mechanism, retry every 10 seconds, with a maximum of 6 retries; communication interruption: cache data locally; after restoring communication, upload cached data in batches; data anomaly: set the valid value range, define the upper and lower limits according to the device type; mark abnormal data for subsequent analysis and processing. 4. Device status record maintains the collection success / failure counter of each device: Collection success counter: records the total number of data packets successfully collected by the device; Collection failure counter: records the number of collection failures caused by abnormalities of the device. 5. Data transmission security mechanism. Data encryption transmission: use AES256 encryption; device authentication mechanism: two-way authentication through device public key and private key; transmission checksum: use MD5 to generate a checksum to ensure data integrity. Output: a structured set of raw data: Record all collected environmental and network behavior data.

[0120] Step S102: Data cleaning and standardization. Purpose: Clean, standardize, format and store the original data set output from step S101 to provide a high-quality input data set for subsequent training steps. Input: Structured original data set: Contains all raw data collected from the device. Processing process:

[0121] 1. Data reception verification. Generate BatchID: Generate a unique BatchID based on the time range or reception order of the data, in the format of BATCH-YYYYMMDD-ID. Field integrity verification: Check whether each record contains the device ID, data type, value, and timestamp fields; records with missing fields will be marked as abnormal and added to the log. Timestamp continuity check: Check the continuity of data timestamps by device ID and record abnormal time periods. Checksum verification: Verify the checksum of the data packet to ensure that the data has not been tampered with.

[0122] 2. Data cleaning. (1) Missing value processing: For data with consecutive missing values ​​of less than 2 hours, use linear interpolation to fill in; for discrete data, use mode filling; for data with consecutive missing values ​​of more than 2 hours, mark it as anomaly and record it in the log. (2) Outlier processing: Calculate the Z score, 3<|Z|≤5 is a mild anomaly, and use the mean of the data in the past hour to replace it. N is the number of data points in the past hour; |Z|>5 is a serious anomaly, which is marked as an anomaly but the original value is retained.

[0123] 3. Data standardization. (1) Numerical data standardization: use Min-Max to standardize continuous values ​​to the interval [0,1]; use Z-Score to standardize time series data. (2) Categorical data encoding: use One-hot encoding for device types; use integer encoding for data types.

[0124] 4. Data storage. (1) Hierarchical storage structure. Original data layer: stores complete original data for retrospective analysis. Cleaned data layer: stores cleaned and standardized data. Abnormal data layer: stores data sets marked as abnormal. (2) Index optimization: BatchID is the unique identifier of the entire batch of data, which is used to quickly locate batch data. (3) Partition strategy: partition by day and store by device sub-partition.

[0125] 5. Data output configuration. (1) Add BatchID to the output data structure: Add a BatchID field to each record to quickly track the source of the data batch. (2) Record version information: Add a version number to each batch of data, such as "v1.0". Output: Cleaned and standardized data set.

[0126] Among them, BatchID is the number that uniquely identifies each batch of data; Device ID is the unique identifier of the device from which the data comes; Data type is the indicator collected by the sensor, including RF data (R f ), fuel value (F l ), power value (E l ), speed (S p ), temperature (T p ), transaction data (T x ). ; Value is the numerical value corresponding to the data type (floating point type). ; Timestamp records the data collection time in ISO8601 format, accurate to milliseconds.

[0127] Step S103: Initialize the global model. Purpose: Initialize the global model through model structure definition and parameter initialization rules to ensure that the federated learning framework provides a unified initial model structure and parameter configuration, and supports the consistency of subsequent feature extraction, training, and parameter aggregation processes. Processing process:

[0128] 1. Model structure definition: (1) Encoder structure: Input layer: According to the feature dimension d of the standardized data modle Dynamic adjustment. Hidden layer: Design three fully connected layers, with n1=d nodes. modle ×2, Bottleneck layer: Generates hidden variables z and outputs mean vector μ and variance vector σ 2 (2) Decoder structure: Symmetric design: the number of hidden layer nodes is [n3, n2, n1] in sequence. Output layer: restored to the original feature dimension d modle (3) Activation function: All hidden layers and bottleneck layers use the LeakyReLU activation function (α=0.1).

[0129] 2. Self-attention module: (1) Multi-head attention mechanism: Number of attention heads: n heads ∈{4,8,16}. Dimension of each head: d head =d modle / n heads . Attention calculation formula: (2) Position encoding: Use relative position encoding to enhance time series features.

[0130] 3. Parameter initialization. (1) Weight initialization: Encoder and decoder weights are initialized using He; self-attention module weights are initialized using Xavier. (2) Bias initialization: All are set to 0.

[0131] 4. Loss function configuration. The total loss function is: L = λ1·L recon +λ2·L KL +λ3·L adv . Weight coefficients: λ1=1.0, λ2=0.1, λ3=0.01.

[0132] Among them, L recon For the reconstruction loss term, the mean square error (MSE) is minimized; L KL is the KL divergence loss term, and the β-VAE form is introduced to control the distribution of latent variables; L adv To combat the loss term, the Wasserstein distance is used to evaluate the similarity between the generated and real data distributions; each loss term has a corresponding weight coefficient (λ) to balance their contributions, λ1 = 1.0 is the reconstruction loss weight; λ2 = 0.1 is the KL divergence loss weight; λ3 = 0.01 is the adversarial loss weight. Output: The initialized global model G0, including the model architecture and parameter configuration, is downloaded by each node for local training.

[0133] Step S2: local model training.

[0134] Step S201, data loading, verification and division. Purpose: Load the data cleaned and standardized in step S102, verify its integrity and consistency, and divide it into training set, verification set and test set according to the model training requirements, and evaluate the quality and balance of the data after segmentation. Input: Data cleaned and standardized in step S102 {(BatchID, device ID, data type, value, timestamp)} cleaned . Processing process: 1. Input data verification. Check the integrity and consistency of the data format, including field name, data type and timestamp format. Verify the legitimacy and continuity of the timestamp to ensure that the time coverage meets the requirements. Check the legitimacy of key data types and value ranges. Count the missing value ratio to ensure that the missing rate is within the threshold range (such as within 2%); for missing values, use interpolation method. 2. Data set segmentation. (1) Time series segmentation. Training set ratio: N train =0.8·N total . Validation set ratio: N valid =0.1·N total . Test set ratio: N test =0.1·N total (2) Device data balance. For each device, the proportion in each data set must satisfy: Where ∈ is the allowable proportional deviation. 3. Quality assessment after segmentation. (1) Distribution consistency test. Use the Kolmogorov-Smirnov (KS) test to compare the distribution consistency of the training set, validation set, and test set, and calculate the KS statistic: D = sup x |F1(x)-F2(x)|. Where, and are the cumulative distribution functions of the two data sets. (2) Time coverage calculation.

[0135] Output: Split data set {(BatchID, device ID, data type, value, timestamp, data set type)} split Data set type: train means training set; valid means validation set; test means test set.

[0136] Step S202: Self-attention feature extraction. Purpose: Extract important features through the self-attention mechanism to enhance the correlation and temporal dependency between features. Input: The segmented feature data set output by S201 {(BatchID, device ID, data type, value, timestamp, data set type)} split .

[0137] Processing process: 1. Feature generation.

[0138] For each data type (R f 、F l 、El , S p , T p , T x ), adopting the following detailed feature generation strategy:

[0139] (1) RF data (R f ): frequency feature fusion. Original features: extract spectrum-related features, such as power spectrum density, bandwidth features, etc. Feature processing: calculate R f The mean of the data in the frequency domain Standard Deviation and peak value; extract frequency domain energy distribution features through fast Fourier transform (FFT); extract high frequency signal features through high pass filtering, and smooth low frequency signal features through low pass filtering. Fusion method: splice time domain features (mean, variance) and frequency domain features (spectral density) into a unified feature vector

[0140] (2) Fuel value (F l ): Amplitude feature fusion. Original features: Extract the volatility and outlier features of fuel value. Feature processing: Calculate volatility Mark outliers that exceed the upper and lower quartiles and calculate the density of outlier distribution. Fusion method: Normalize the volatility feature and the outlier distribution feature and then splice them

[0141] (3) Electricity value (E l ): Energy feature fusion. Original feature: Calculate the accumulated energy and energy consumption rate. Feature processing: Calculate the accumulated energy consumption and the energy change per unit time (energy consumption rate) Fusion method: concatenate the accumulated energy consumption and rate features into a vector

[0142] (4) Speed ​​(S) p ): Motion feature fusion. Original features: Extract velocity change trend and acceleration information. Feature processing: Calculate acceleration And use sliding window to calculate the speed growth or decay trend. Fusion method: Combine speed and acceleration features into trend vector

[0143] (5) Temperature (T p ): Environmental feature fusion. Original features: Extract temperature fluctuation features and extreme point features. Feature processing: Get the maximum temperature within the time window and minimum value Calculate the fluctuation range Mark outliers that exceed physical limits (e.g., T p >85℃ or T p<-40℃). Fusion method: extreme values, fluctuation ranges and abnormal markers are used as unified features.

[0144] (6) Transaction data (T x ): Time series feature fusion. Original features: Extract transaction frequency, time interval and sequence correlation features. Feature processing: Count the number of transactions per unit time (i.e. transaction frequency) Calculate the mean and variance of the transaction time interval, and use the autocorrelation function to analyze the periodic characteristics of the transaction time series. Fusion method: Combine the transaction frequency, time interval characteristics and autocorrelation characteristics.

[0145] Concatenate the features of all data types to generate a unified feature vector:

[0146] 2. Multi-head attention calculation.

[0147] (1) Initialization of attention head: Input feature vector X fusion Divide into n heads subspaces, the dimension of each head is: The remaining dimension (d model mod n heads ) are assigned to some heads to ensure the total dimension is consistent. Initialize the query (Q), key (K), and value (V) matrices: Q = XW Q ,K=XW K ,V=XW V Among them, W Q ,W K ,W V is the learnable weight matrix, and its dimensions are

[0148] (2) Attention score calculation: Use the scaled dot product attention formula: Filter below attention threshold The attention weights ensure that effective features participate in fusion.

[0149] (3) Output feature calculation: Output the weighted feature matrix of each attention head: head i =Attention(Q,K,V),i=1,2,…,n heads . Concatenate the outputs of all attention heads: Among them, is the output weight matrix, the dimension is:

[0150] 3. Position coding injection. (1) Relative position coding: Generate the relative position matrix PE according to the time window, using the sine-cosine position coding formula: Among them, pos represents the time position index of the feature in the sequence; i is the index of the current feature dimension; d model is the total dimension of the feature vector, and the dimension of the position encoding matrix PE is ΔT is the length of the time window. (2) Position encoding fusion: Add the generated relative position encoding matrix to the output features of the attention mechanism to generate a feature vector that injects position information: Among them, X fusion′ The feature vector injected with temporal information has the dual expression ability of time position and feature correlation; The feature matrix generated by the multi-head attention mechanism contains the correlation information between features; PE is the position encoding matrix used to provide time series information. (3) Feature normalization: fusion′ Perform LayerNormalization to ensure that the eigenvalues ​​have a uniform scale: Among them, μ is the feature mean; σ is the feature standard deviation; ∈ is the smoothing factor to prevent the denominator from being zero. Output: Optimized feature set

[0151] Step S203: Variational Autoencoder Training. Purpose: Use variational autoencoder (VAE) to extract latent features to support the generation of subsequent generative adversarial network (GAN) models. Ensure the effectiveness and consistency of latent features by optimizing the distribution of latent space and the quality of feature reconstruction. Input: The optimized feature set output by S202

[0152] Processing process: 1. Model structure design: (1) Encoder: Input: Optimized feature vector

[0153] Output: Use the fully connected network (Fully Connected Layers) to generate the mean vector μ and logarithmic variance vector logσ of the latent space 2 : Network configuration: 1. Input layer: feature dimension d modle 2. Hidden layer 1: ReLU activation, number of nodes 3. Hidden layer 2: ReLU activation, number of nodes 4. Output layer: Generate μ and logσ 2 , the number of nodes is all potential dimension k. Potential dimension k: according to the feature compression ratio The default Compression Ratio is 0.25. (2) Latent space sampling: sampling latent features x by reparameterization according to mean and variance latent=μ+σ·∈,∈~N(0,I). (3) Decoder: Input: Sampled latent vector Output: Generate reconstructed feature vector Network configuration: 1. Input layer: potential dimension k. 2. Hidden layer 1: ReLU activation, number of nodes 3. Hidden layer 2: ReLU activation, number of nodes 4. Output layer: feature dimension d model , no activation.

[0154] 2. Loss function optimization. (1) Reconstruction loss: measures the difference between input features and reconstructed features. (2) KL divergence loss: constrains the distribution of the latent space to be close to the standard normal distribution. (3) Total loss: the weighted sum of the two parts of loss. Wherein, β is the weight coefficient (the default value is β=1).

[0155] 3. Training process. (1) Initialize model parameters. Randomly initialize the weights of the encoder and decoder. Training parameter configuration: Batch size: B = 64. Learning rate: η = 0.001. Maximum number of iterations: N epochs =100. Optimizer: Adam (β1 = 0.9, β2 = 0.999). Early stopping strategy: Stop training when the validation set loss does not decrease for 10 consecutive rounds. (2) Loss optimization and weight update. Use the Adam optimizer to update the model parameters and minimize the total loss: In each batch, the reparameterization technique is used to sample potential features and update the model weights. Validation set usage: Calculate the validation set loss after each iteration for early stopping judgment. Output: Potential feature set {(BatchID, device ID, time window, x latent , dataset type)} latent .

[0156] Step S204: Generative Adversarial Network training. Purpose: Generate more diverse feature samples using the Generative Adversarial Network (GAN) model to improve the robustness and generalization ability of the model. Input: 1. The potential feature set output by S203 {(BatchID, device ID, time window, x latent , dataset type)} latent 2. Potential Dimensions Processing process:

[0157] 1. GAN model design. (1) Generator: Input: Random noise Noise Dimension Potential feature x latent . Output: Generated features Network configuration: 1. Input layer: noise dimension n=100 and potential dimension 2. Hidden layer 1: ReLU activation, number of nodes 3. Hidden layer 2: ReLU activation, number of nodes 4. Output layer: feature dimension d model , using Tanh activation. (2) Discriminator: Input: 1. True feature x′ latent : Directly from the latent features x latent 2. Generate feature x gan : The production features generated by the generator are used as the generated data input of the discriminator. Output: classification probability, distinguishing true and false samples. Network configuration: 1. Input layer: feature dimension d model 2. Hidden layer 1: LeakyReLU activation, number of nodes 3. Hidden layer 2: LeakyReLU activation, number of nodes 4. Output layer: 1 node, using Sigmoid activation.

[0158] 2. Loss function design. Wasserstein distance and gradient penalty (WGANGP) is used. (1) Discriminator loss: Among them, λ is the gradient penalty coefficient (default λ=10), is the real feature x′ latent and generate features x gan Linear interpolation of . (2) Generator loss:

[0159] 3. Training process. (1) Initialization: Randomly initialize the weights of the generator and discriminator. Training parameter configuration: Generator learning rate: η G =0.0001. Discriminator learning rate: η D =0.0004. Batch size: B=64. Optimizer: Adam (β1=0.5, β2=0.999). Discriminator and generator update ratio: n D :n G =5:1. (2) Iterative training: Each iteration includes the following steps: 1. Discriminator training. Fix the generator weights to ensure that the optimization target of the discriminator is independent of the generator. Sample x′ from the true feature distribution latent ~P r . Get the generated sample x from the generator output gan ~P g . Generate linear interpolation samples: Calculate the discriminator loss using real samples, generated samples, and interpolated samples: minimize Update the discriminator weights. 2. Generator training. Fix the discriminator weights to ensure that the generator optimization target is independent of the discriminator weight update. Use random noise and the latent feature x latent , generate enhanced features: Generate a sample x using gan ~P g Calculate the generator loss: minimize Update the generator weights. Output: Enhanced feature set {(BatchID, device ID, time window, x gan , dataset type)} gan .

[0160] Step S205: Generate local model parameters. Purpose: Generate local model parameters in a standard format based on the trained model to provide input for global model aggregation (step S3). Through a clear parameter extraction and standardization process, ensure seamless connection with the global aggregation interface, and add a parameter importance evaluation mechanism to improve model quality. Input: The enhanced feature set output by S204 {(BatchID, device ID, time window, x gan , dataset type)} gan . Processing process:

[0161] 1. Parameter extraction. (1) Self-attention layer: extract the attention weight matrix

[0162] in, are the query and key matrices, d head is the attention head dimension. (2)VAE: Extracting encoder weights and bias b enc . Extract decoder weights and bias b dec (3) GAN: Extracting the generator weight matrix Extract the discriminator weight matrix

[0163] 2. Parameter standardization. (1) Unify the format of all parameters: use hierarchical representation: in, is the weight matrix, and b is the bias vector. Add hierarchical information to ensure consistency with the interface specification of step S3. (2) Calculate the importance weight of the parameters: Use the Frobenius norm and variance to comprehensively evaluate the importance of the parameters: in, is the Frobenius norm of the parameter matrix, is the matrix element, M, N are the matrix dimensions; is the element variance of the parameter matrix, is the mean of the matrix elements; ζ, η are weight coefficients used to balance the importance of norm and variance (e.g. ζ = 0.7, η = 0.3). (3) Adjustment of inter-layer scale differences: Parameters of different layers may have scale differences (e.g. weight matrices or bias vectors of different sizes). In order to avoid these differences affecting the calculation of importance weights, normalization is required: in, is the normalized parameter matrix, and are the minimum and maximum values ​​of the matrix respectively.

[0164] 3. Parameter verification. (1) Integrity check: Ensure that the parameters of each layer (2) Consistency verification: Check whether the parameter dimensions match the model structure definition. Verify whether the parameter format conforms to the global aggregation interface specification (requirements of step S3). Output: 1. A set of local model parameters {LM1, LM2, LM3, ..., LM x}. Each local model parameter includes: model architecture description; hierarchical parameter matrix; parameter importance weight. 2. Local parameters are consistent with the global model aggregation rules: parameter matrix distribution must satisfy the global model aggregation formula: Ensure the consistency of parameter importance weights and sample distribution weights in order to efficiently participate in the construction of the global model.

[0165] Step S3: Global model aggregation. Purpose: Aggregate model parameters from all local nodes to generate a unified global model GM to improve the generalization ability of the model while protecting data privacy. Input: Local model parameter set {LM1, LM2, LM3, …, LM x}. Processing process:

[0166] Step S301: Parameter reception and verification. The central server receives local model parameters {LM1, LM2, LM3, ..., LM x}, perform the following verifications on the uploaded data: Integrity verification: Check whether all necessary parameters (weight matrix, bias vector and activation function) are included. Consistency verification: Ensure that the local model parameter structure is consistent with the global model structure, including weight dimension and network hierarchy. Data source audit: Record the upload time and participation status of each node to ensure traceability.

[0167] Step S302: weight calculation. According to the number of samples of each node |D i |, calculate the weight w of weighted aggregation i : Among them, |D i | is the number of samples of the ith node, and N is the total number of participating nodes. The calculation of weights ensures that nodes with large sample sizes contribute more to the global model, fully reflecting the importance of data distribution at each node.

[0168] Step S303: Global model parameter aggregation. The global model parameters are calculated layer by layer using the weighted average method: For the weight matrix: For the bias vector: in, is the weight matrix of the ith node, is the bias vector of the ith node.

[0169] Step S304: Model optimization and constraint checking. The aggregated global model parameters must meet the following constraints: Gradient stability: To avoid gradient vanishing or exploding problems, use gradient clipping technology: Where C is the gradient threshold (e.g., 1.0 to 5.0). Parameter range correction: Ensure that all weights and biases are within a reasonable range to prevent outliers from affecting model performance.

[0170] Step S305: Global model distribution. Encrypt and distribute the aggregated global model GM to each participating node for the next round of local model training. Add version control information (e.g., "v1.0") to support historical backtracking and version management of the model.

[0171] Output: Global model parameter set {GM}: 1. Model structure information: describes the structural information such as network hierarchy and activation function. 2. Weight matrix and bias vector: stored by hierarchy to ensure compatibility with the local model of each node. 3. Weight calculation record: stores the sample weight and calculation formula of each node to support subsequent analysis and audit. Among them, W GM is the weight matrix of the global model, b GM is the bias vector of the global model; ‖g‖ is the gradient norm, and C is the clipping threshold.

[0172] Step S4: Data authentication. Purpose: To ensure the data authenticity and security of global model parameters, based on the consensus algorithm PoAh (Proof of Authentication) mechanism, the data block is authenticated, a trusted data block set is generated, and the authentication results are stored through the blockchain network. Input: Global model parameter set {GM}, IoT application set {I1, I2, I3, …, I n}, a network containing multiple trusted nodes, private key {Prkey}, public key {Pubkey}. Processing process:

[0173] Step S401: Initialize the transaction block set. Purpose: Pack the global model parameters and generate transaction data blocks. Operation: Convert the input global model parameters {GM} into a transaction data block set {Tx1, Tx2, Tx3, ..., Tx i}. Data block definition: Each block contains a unique number, summary information of global model parameters and a timestamp. Tx i =(BlockID,GMHash,Timestamp). Block ID is the block number; GMHash is the global model parameter summary generated by the SHA256 algorithm; Timestamp is the timestamp of the block generation. Output: transaction data block set {Tx1, Tx2, Tx3,…, Tx i}.

[0174] Step S402: Allocate private keys and broadcast blocks. Purpose: Allocate private keys to blocks and send them to nodes in the network for authentication. Operation: Allocate private keys to each transaction block: Prkey → {Tx1, Tx2, Tx3, …, Tx i}. Where {Prkey} is the set of private keys assigned to the block. The private key-signed block is broadcast to all nodes in the network: Output: The broadcasted transaction block is ready for receiving node authentication.

[0175] Step S403, trusted node screening. Purpose: Screen trusted nodes from the network to participate in the authentication process. Operation: Traverse all nodes in the network and check their credibility. If the node is a trusted node Node=trusted, and its trust value tr satisfies tr≥th (trust value threshold th=5), the node passes the authentication. If the node is an ordinary node, check whether the trust value is received. If there is no trust value, jump back to the rebroadcast step. Output: The screened trusted node set {TrustedNodes}.

[0176] Step S404: Block signature verification. Purpose: To authenticate the transaction data block to ensure data integrity and authenticity. Operation: Use a trusted node to verify the signature in the transaction data block set: Use the SHA256 algorithm to calculate the hash value of the data block and verify the matching of the signature and the public key: H i =SHA256(Tx i ),Verify(H i ,Signature). Among them, H iis the hash value of the ith transaction data block, and Verify is the authentication function used to verify the consistency between the signature and the hash value. Output: A set of transaction data blocks that have been successfully authenticated.

[0177] Step S405: Process the authentication result. Purpose: Process the block according to the authentication result. Operation: If the authentication succeeds, add the block to the blockchain network and broadcast it to all nodes for account book update; if the authentication fails, mark the transaction data as untrustworthy and discard it. Output: The authenticated data block set {Tx1, Tx2, Tx3, ..., Tx i}, Authentication result (success / failure).

[0178] Step S5, real-time intrusion detection. Purpose: To detect anomalies in real-time collected IoT data through the authenticated global model, identify potential intrusion behaviors (such as DDoS, DoS attacks, etc.), and output the detection results to ensure the security and stability of the IoT wireless sensor network (WSN). Input: 1. The authenticated global model parameter set {GM} stored in the cloud. 2. The real-time collected IoT data set {R f ,F l ,E l ,S p ,T p ,T x} new . Processing process:

[0179] 1. Global model loading. Load the authenticated global model parameters {GM} from the cloud, including the weight matrix W GM and the bias vector b GM , initialize the detection model. 2. Real-time data feature extraction. For real-time input data {R f ,F l ,E l ,S p ,T p ,T x} new Perform feature extraction and generate a unified feature vector X new For the specific process, refer to the feature generation process in step S202. The final feature vector X new It is expressed as: 3. Model prediction. Use the global model to predict the feature vector X new Make prediction: y pred =sigmoid(W GM X new +b GM ). Among them, y pred is the predicted value (probability). 4. Abnormality determination. According to the set abnormality detection threshold τ (such as τ = 0.8), determine whether the data is abnormal: if y pred>τ, it is judged as abnormal behavior “abnormal”; if y pred ≤τ, it is judged as normal behavior "normal". 5. Result storage and output. The test results are stored as structured records for subsequent analysis or alarm. The record format is: {BatchID, device ID, timestamp, y pred , test results}. Among them, BatchID is the unique identifier of the data batch; device ID is the data source device; timestamp is the data collection time; pred The model predicts the probability; the detection result is "abnormal" or "normal". Output: {BatchID, device ID, timestamp, y pred , test results}.

[0180] The embodiments of the present application have the following advantages: First, the embodiments of the present application are based on the self-attention mechanism, and enhance the time series feature extraction capability through a multi-head attention module, and combine the deep feature fusion and correlation analysis of multi-source heterogeneous data such as radio frequency, power, and speed, which significantly improves the detection accuracy of abnormal behaviors and potential attacks in wireless sensor networks, and effectively solves the problem of insufficient generalization ability of traditional methods in complex scenarios.

[0181] Second, the embodiment of the present application combines a time-varying variational autoencoder (VAE) with a generative adversarial network (GAN). VAE extracts potential features for abnormal pattern recognition, and GAN generative adversarial training is combined to enhance sample diversity, thereby effectively improving the robustness and adaptability of the model in the face of complex and changeable attacks such as DDoS, and overcoming the shortcomings of existing methods in their limited ability to detect unknown attacks.

[0182] Third, the embodiment of the present application is based on the PoAh consensus mechanism. Through the efficient data and model authentication process of trusted nodes, while ensuring the security and credibility of the distributed system, it significantly reduces the computing and communication overhead of the blockchain, and solves the problem of excessive computing burden of traditional consensus algorithms in resource-constrained wireless sensor network environments.

[0183] Fourth, the embodiment of the present application adopts a privacy-preserving federated learning architecture, which avoids the centralized transmission of raw data by performing local training on each node and sharing only model parameters. This not only effectively protects user privacy, but also significantly reduces the computing load of the central server and network communication resource consumption through distributed computing.

[0184] Fifth, the four-layer distributed architecture of device layer, federation layer, authentication layer and cloud layer proposed in the embodiment of the present application realizes an end-to-end framework from data collection, model training to anomaly detection. The architecture is clear and easy to expand, can adapt to a variety of IoT application scenarios, and supports efficient and real-time intrusion detection and data management, further improving the stability and security of the IoT network.

[0185] In an exemplary embodiment, the security analysis of this embodiment is studied for three types of attacks: distributed denial of service attack (DDoS), denial of service attack (DoS) and abnormal attack. The simulation of the proposed architecture is run on the following hardware environment: Intel Core i7 processor (3.2GHz), 16GB memory, 64-bit Windows 7 operating system. The network architecture contains 3×33×3 hidden layers of local model and global model. The node performance and communication rounds of successful data authentication and verification of miner nodes in the IoT blockchain network are evaluated using Node.js v8.9.1, and the PoAh consensus algorithm is adopted.

[0186] For the quantitative analysis of the proposed architecture, this embodiment uses an existing open source DDoS attack dataset. The dataset contains 809,361 log records with 78 attributes. Under the network analysis performance of the proposed architecture, the logs are classified into two categories: "normal" and "DoS attack". Among them, "normal" is related to legitimate traffic, and "DoS attack" is related to "slow attack" and "flood attack". The description of the dataset attributes is shown in Table 1, and the number of samples corresponding to each category is shown in Table 2.

[0187] Table 1 Dataset attributes

[0188] variable describe Destination Port Destination port number Stream duration Duration of total flow Total forwarded packets Total number of forwarded packets Total returned packets Returns the total number of packets Forwarded packet length The total length of the forwarded data packet Stream IAT-Maximum Maximum inter-arrival time Stream IAT-minimum Minimum arrival interval Forward IAT-Total Total inter-arrival time Flow packets / second Number of packets per second Stream Bytes / Sec The number of bytes per second

[0189] Table 2 Sample size

[0190]

[0191] Figure 3 The overall architecture diagram of the wireless sensor network (WSN) system with self-attention mechanism. This architecture is used in the hierarchical structure of DDoS attack detection in the IoT network, aiming to solve the problems of privacy protection, security, data authentication and data processing delay. The proposed architecture consists of four layers: device layer, federation layer, authentication layer and cloud layer. Each layer realizes efficient intrusion detection function through layer-by-layer data transmission and verification.

[0192] The figure shows the main functions of the architecture, including data collection, data aggregation and data authentication, and completes these functions layer by layer. The following are the specific functions and processes of each layer:

[0193] 1. Device layer: The device layer is the main data collection layer of the IoT, which contains various IoT sensors and devices used to obtain and collect data from various IoT applications. The device layer includes three categories: IoT devices, sensor devices, and industrial devices, such as personal computers, smartphones, cameras, cars, wearable devices, and drones. This layer generates raw data (including RF, fuel, electricity, speed, temperature, traffic, entertainment, etc.) according to application requirements and transmits the data to the federation layer. The set of IoT devices is represented by {D1, D2, D3, …, D x}, the IoT application set is represented as {I1,I2,I3,…,I n}, where represents the number of IoT and sensor devices; represents the number of IoT applications; the data generated by each IoT device includes radio frequency R f Fuel F l 、Electricity l , speed S p Temperature I p , transaction data T x wait.

[0194] 2. Federal layer: As an intermediate layer, the federal layer is responsible for data processing and data aggregation, connecting the device layer and the authentication layer. The federal layer consists of base stations and access points for temporary storage and transmission of data. In this layer, decentralized learning and privacy protection are achieved through a federated learning network (including local models and global models). The local model (such as the self-attention time-varying VAE-GAN model) trains local data, generates gradients and weights, and transmits them to the global model for aggregation. In this layer, the local model and the global model in the proposed architecture are represented as {LM1, LM2, LM3} and GM, respectively. The local model will be trained before being uploaded to the global model, and data aggregation will be performed on the central server. After training, it will be transferred to the authentication layer.

[0195] 3. Authentication layer: The authentication layer provides a bridge function between the federation layer and the cloud layer. Its main purpose is to transmit verified data to the cloud layer. The authentication layer uses the PoAh consensus algorithm to implement data authentication and block verification through a distributed blockchain network. In the authentication layer, network nodes are responsible for data authentication and verification operations. Each node maintains the security and privacy of data in the blockchain network and transmits the data to the cloud layer after verification. The authentication process is performed through a private key {Pr key} and the public key {Pub key}conduct.

[0196] 4. Cloud layer: The cloud layer is the final storage layer of the architecture, which contains data centers for storing and analyzing data. The cloud layer has a large amount of storage capacity and high-performance computing devices, which can provide management, configuration, distribution and prediction services for IoT applications. The cloud layer also faces security and performance issues such as session hijacking and MITM attacks. After the authentication layer completes the data verification, the data is transferred to the cloud layer for storage and further analysis.

[0197] Figure 3 The overall architecture of the wireless sensor network (WSN) system with self-attention mechanism is demonstrated, and the application of this architecture in DDoS attack detection in the Internet of Things is explained. Through layer-by-layer data collection, processing, authentication and storage processes, this architecture effectively solves the security and privacy issues in the Internet of Things and realizes an efficient intrusion detection and protection mechanism.

[0198] Figure 4 The federated learning architecture of self-attention time-varying VAE-GAN is presented, which is designed for intrusion detection tasks in wireless sensor networks (WSNs). The framework is divided into multiple functional modules, including data collection, local model training, global model aggregation, and data authentication and storage, as follows:

[0199] 1. Data collection: The framework first collects data from devices in the wireless sensor network, such as radio frequency, fuel value, power value, speed, temperature, transaction data, etc. These data will be transmitted to the local model training module to support intrusion detection tasks.

[0200] 2. Local model training: Each node uses the self-attention time-varying VAE-GAN model for local training. This model uses the advantages of the self-attention mechanism and the generative adversarial network (GAN) to perform deep learning on time series data features to identify normal and abnormal patterns. After training, the model parameters (such as weights and gradients) will be uploaded to the data aggregation module.

[0201] 3. Data aggregation and global model training: Local model parameters are aggregated on the central server to generate a global model. This process uses a federated learning approach to ensure that the original data is not transmitted and protect the data privacy of each node.

[0202] 4. Data authentication and storage: At the authentication layer, the system uses the PoAh mechanism to verify the aggregated model to ensure the authenticity and security of the data. The authenticated model blocks are added to the blockchain and eventually stored in the cloud for download and use by each node.

[0203] The framework uses the self-attention time-varying VAE-GAN model for intrusion detection and combines PoAh authentication to enhance the security and privacy protection capabilities of the system. Through the architecture of distributed data processing and federated learning, the framework achieves efficient and reliable intrusion detection in resource-constrained wireless sensor network environments.

[0204] Figure 5 This is a flowchart of a wireless sensor network intrusion detection method based on a self-attention federated learning architecture. The main steps include:

[0205] Step S1, data collection and initialization, based on the environmental data and network behavior data collected by the device set in the wireless sensor network, aims to generate a high-quality, standardized data set. Through the operations of device data collection and initialization, data cleaning and standardization, and global model initialization, a cleaned and standardized data set and an initialized global model are constructed to provide data and model support for feature extraction, model training and parameter aggregation of the subsequent federated learning framework.

[0206] Step S2, local model training, is based on the cleaned and standardized data set output by step S1, with the goal of building a high-quality, localized feature model. Through operations such as data loading, verification and partitioning, self-attention feature extraction, variational autoencoder training, generative adversarial network training, and local model parameter generation, it builds optimized feature sets, potential feature sets, enhanced feature sets, and local model parameters in a standard format to provide support for global model aggregation and federated learning.

[0207] Step S3, global model aggregation, based on the local model parameter set, aims to generate a unified global model. Through operations such as parameter reception and verification, weight calculation, global model parameter aggregation, model optimization and constraint checking, and global model distribution, a global model parameter set containing model structure information, weight matrix and bias vector, and weight calculation records is constructed to improve the generalization ability of the model and support the next round of local model training.

[0208] Step S4, data authentication, is based on the global model parameter set, the IoT application set, the network containing multiple trusted nodes, and the private key and public key, with the goal of ensuring the data authenticity and security of the global model parameters. Through operations such as transaction block initialization, private key allocation and block broadcasting, trusted node screening, block signature verification, and authentication result processing, a set of authenticated trusted data blocks is constructed, and the authentication results are stored through the blockchain network to ensure the integrity and credibility of the data.

[0209] Step S5, real-time intrusion detection, based on the authenticated global model parameter set stored in the cloud and the IoT data set collected in real time, aims to identify potential intrusion behaviors and ensure the security and stability of the IoT wireless sensor network. Through operations such as global model loading, real-time data feature extraction, model prediction, anomaly determination, and result storage and output, a structured record containing data batch identification, device identification, timestamp, prediction probability and detection results is constructed for subsequent analysis and alarm response.

[0210] Figure 6 The figure shows the logarithmic distribution of the confusion matrix of the model prediction results. This figure shows the quantitative distribution of the four key indicators predicted by the model. The number of samples of true positive (TP) and true negative (TN) reached 65,996 and 55,315 respectively, indicating that the model has good recognition ability for both positive and negative samples. The number of samples of false positive (FP) and false negative (FN) is 85 and 9 respectively, which are relatively small values, indicating that the model has a low misclassification rate. The logarithmic coordinate axis is used in the figure to better display the indicator data with significant differences in order of magnitude, which intuitively reflects the prediction performance of the model. From the data distribution, it can be seen that the model has high accuracy and reliability, and there are fewer misclassifications.

[0211] Figure 7 This is a comparative analysis chart of the performance evaluation of federated learning and non-federated learning, showing the comparative analysis of federated learning and non-federated learning on five key performance indicators. The left side is a radar comparison chart, which intuitively shows the overall performance difference between the two learning methods in terms of accuracy, precision, recall, F1 score and efficiency; the right side is a performance improvement analysis chart, which shows the specific numerical changes through a broken line and annotates the performance improvement with a bar chart. It can be seen from the figure that federated learning is better than non-federated learning schemes in all evaluation indicators, including: the accuracy rate increased by about 4.92 percentage points (99.92% vs 95%); the precision rate increased by about 7.87 percentage points (99.87% vs 92%); the recall rate increased by about 5.99 percentage points (99.99% vs 94%); the F1 score increased by about 6.93 percentage points (99.93% vs 93%); and the efficiency increased by about 7.92 percentage points (99.92% vs 92%).

[0212] Figure 8The figure is a schematic diagram of the performance comparison analysis of the architecture based on federated learning, which shows the performance comparison analysis of the proposed architecture and the traditional architecture under different transaction numbers. The left figure is a direct performance comparison diagram, which shows the computing cost change trend of the two architectures in the range of transaction numbers from 20 to 100 through broken lines and dotted lines. The blue curve represents the proposed architecture and the orange curve represents the traditional architecture. The right figure is a performance difference and improvement analysis diagram. The blue bar chart represents the absolute difference in computing cost, and the orange curve represents the relative performance improvement percentage. It can be seen from the figure that: as the number of transactions increases, the performance difference between the two architectures gradually becomes significant; when the number of transactions is 60, the performance difference reaches 300, and the corresponding performance improvement reaches 42.9%; when the number of transactions reaches 100, the proposed architecture still maintains a good performance advantage, and the computing cost is reduced by about 50 units compared with the traditional architecture; overall, the proposed architecture shows better scalability and efficiency advantages when processing a large number of transactions.

[0213] In the PoAh-based authentication process, six nodes are used in the blockchain network, including three trusted nodes (miner nodes). The block size of these nodes is 35 bytes. All network nodes encrypt transactions using public key cryptography and sign certificates using private keys. The trusted nodes verify the original signature using the public key to verify the block. After successful verification, the trusted node broadcasts the block to all nodes in the blockchain network, ensuring that each node keeps a copy of the block in its ledger. The average time for block authentication using the PoAh consensus algorithm is 3.33 seconds, including nine iterations, and following all the steps in the proposed architecture, see Table 3.

[0214] Table 3 Block verification schedule based on PoAh

[0215] Iterations (number of transactions) 1 2 3 4 5 6 7 8 9 Block verification time (seconds) 3.4 3.6 2.8 4.02 2.9 3.23 3.42 3.4 3.26

[0216] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0217] Based on the same inventive concept, the embodiment of the present application also provides a network intrusion detection device for implementing the network intrusion detection method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more network intrusion detection device embodiments provided below can refer to the limitations on the network intrusion detection method above, and will not be repeated here.

[0218] In an exemplary embodiment, Fig. 9 As shown, a network intrusion detection device 900 is provided, including: a data collection module 901, a local model training module 902, a parameter sending module 903 and an intrusion detection module 904, wherein:

[0219] The data acquisition module 901 is used to obtain device data of the IoT device in the wireless sensor network that is in communication with the node device; the device data includes radio frequency data, fuel value, power value, speed, temperature and transaction data;

[0220] The local model training module 902 is used to determine the optimized feature set corresponding to the device data, and train the VAE-GAN model composed of the variational autoencoder and the generative adversarial network based on the optimized feature set to obtain the local model parameters corresponding to the trained VAE-GAN model; the local model parameters comply with the global aggregation interface specification; the optimized feature set is obtained by concatenating the feature vectors of the device data through the self-attention module;

[0221] The parameter sending module 903 is used to send the local model parameters to the central server, so that the central server can aggregate the local model parameters corresponding to multiple node devices to obtain global model parameters, and distribute the global model parameters to each node device and the blockchain network, so that each node device can retrain the local VAE-GAN model, and the blockchain network can authenticate the global model parameters and add them to the blockchain network;

[0222] The intrusion detection module 904 is used to obtain the authenticated global model parameters from the blockchain network, detect the device data to be detected of the IoT device connected to the node device communication according to the global model using the global model parameters, and determine the detection results of the device data to be detected; the global model is a VAE-GAN model jointly trained by multiple node devices and a central server.

[0223] Furthermore, the data acquisition module 901 is specifically used to: receive device data collected by each IoT device through sensors; the device data is encrypted by AES256 and is bidirectionally authenticated by the device public key and private key; the device data is obtained by collecting through the sensors of the IoT device according to the data collection frequency corresponding to the device type of the IoT device; the device data in the structured format includes device ID, data type, value and timestamp; perform data reception verification and data cleaning processing on the device data to obtain cleaned device data; data reception verification includes data field integrity verification, timestamp continuity verification and data packet checksum verification; data cleaning processing includes missing value interpolation, outlier replacement, data standardization and data encoding; the cleaned device data is stored in a hierarchical storage structure; the hierarchical storage structure includes a storage structure partitioned by day or device.

[0224] Furthermore, the device also includes a model initialization module for determining the model structure and parameter configuration of the global model; the input layer in the model structure is dynamically adjusted according to the feature dimension of the standardized data; the hidden layer in the model structure is a three-layer fully connected layer; the bottleneck layer in the model structure contains hidden variables, mean vectors and variance vectors; the hidden layer and the bottleneck layer in the model structure use the LeakyReLU activation function; the self-attention module includes a multi-head attention mechanism; the parameter configuration initializes the encoder and decoder weights through He initialization, uses Xavier initialization for the self-attention module weights, and initializes the bias term to 0; the loss function of the parameter configuration is L=λ1*L recon +λ2*L KL +λ3*L adv ; Among them, the reconstruction loss L recon Using mean square error, KL divergence loss L KL Used to control the distribution of latent variables and the adversarial loss L adv The Wasserstein distance is used, and λ1, λ2, and λ3 are weight coefficients.

[0225] Furthermore, the local model training module 902 is specifically used to: extract features from device data based on the self-attention module to obtain feature vectors corresponding to the device data, and concatenate the feature vectors to obtain a unified feature vector; perform weighted processing on the unified feature vector through a multi-head attention mechanism to obtain a feature matrix; add the feature matrix to the position encoding matrix corresponding to the relative position encoding to obtain an optimized feature set; train the pre-constructed VAE through the optimized feature set until the loss value of the VAE meets the first training end condition, obtain the trained VAE, and determine the potential feature set through the trained VAE and the optimized feature set; train the pre-constructed GAN through the potential feature set until the discriminator loss value corresponding to the GAN meets the second training condition, obtain the trained GAN, and determine the enhanced feature set through the trained GAN and the potential feature set; determine the parameters corresponding to the self-attention module, the trained VAE and the trained GAN respectively according to the enhanced feature set, and determine the local model parameters according to the global aggregation interface specification.

[0226] Furthermore, the parameter sending module 903 is specifically used to: send the local model parameters to the central server so that the central server receives and verifies the local model parameters corresponding to multiple node devices, and then determines the node weight corresponding to each node device according to the number of samples in the node device, and aggregates each local model parameter through a weighted average strategy and node weight to obtain a global model parameter.

[0227] Furthermore, the intrusion detection module 904 is specifically used to: receive the authenticated global model parameters returned by the blockchain network; wherein the blockchain network is used to determine the transaction data block set corresponding to the global model parameters, and broadcast each transaction block in the transaction data block set to the nodes in the blockchain network; when the node passes the trusted node screening and the trusted node completes the verification of the transaction block, the successfully verified transaction block is added to the blockchain network.

[0228] Furthermore, the intrusion detection module 904 is specifically used to: load the global model parameters obtained from the blockchain network into the VAE-GAN model to obtain a global model; determine the feature vector corresponding to the data of the device to be detected, and input the feature vector into the global model to obtain the predicted value corresponding to the data of the device to be detected; determine the detection result of the device to be detected based on the predicted value; the detection result includes normal behavior and abnormal behavior.

[0229] Each module in the above network intrusion detection device can be implemented in whole or in part by software, hardware and their combination. Each module can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0230] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Fig.10 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store device data of each Internet of Things device. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a network intrusion detection method is implemented.

[0231] Those skilled in the art will understand that Fig.10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0232] In an exemplary embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0233] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0234] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0235] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0236] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0237] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A network intrusion detection method, characterized in that: Applied to a node device, the method comprises: Acquire device data of an Internet of Things device in a wireless sensor network that is communicatively connected to the node device; the device data includes radio frequency data, fuel value, power value, speed, temperature, and transaction data; Determine an optimized feature set corresponding to the device data, and train a VAE-GAN model composed of a variational autoencoder and a generative adversarial network based on the optimized feature set to obtain local model parameters corresponding to the trained VAE-GAN model; the local model parameters comply with the global aggregation interface specification; the optimized feature set is obtained by concatenating feature vectors of the device data through a self-attention module; Send the local model parameters to the central server, so that the central server can aggregate the local model parameters corresponding to the multiple node devices to obtain global model parameters, and distribute the global model parameters to each node device and the blockchain network, so that each node device can retrain the local VAE-GAN model, and the blockchain network can authenticate the global model parameters and add them to the blockchain network; The authenticated global model parameters are obtained from the blockchain network, and the device data to be detected of the IoT device to which the node device is connected for communication is detected according to the global model using the global model parameters, so as to determine the detection result of the device data to be detected; the global model is a VAE-GAN model jointly trained by a plurality of the node devices and the central server.

2. The method according to claim 1, characterized in that The step of obtaining device data of an Internet of Things device in a wireless sensor network that is communicatively connected to the node device includes: Receive device data collected by each IoT device through a sensor; the device data is encrypted by AES256 and bidirectionally authenticated by a device public key and a private key; the device data is obtained in a structured format by collecting data through the sensor of the IoT device according to the data collection frequency corresponding to the device type of the IoT device; the device data in a structured format includes a device ID, a data type, a value, and a timestamp; Performing data reception verification and data cleaning processing on the device data to obtain cleaned device data; the data reception verification includes data field integrity verification, timestamp continuity verification and data packet checksum verification; the data cleaning processing includes missing value interpolation, outlier value replacement, data standardization and data encoding; The cleaned device data is stored according to a hierarchical storage structure; the hierarchical storage structure includes a storage structure partitioned by day or device.

3. The method according to claim 1, characterized in that The method further comprises: Determine the model structure and parameter configuration of the global model; the input layer in the model structure is dynamically adjusted according to the feature dimension of the standardized data; the hidden layer in the model structure is a three-layer fully connected layer; the bottleneck layer in the model structure contains hidden variables, mean vectors and variance vectors; the hidden layer and bottleneck layer in the model structure use LeakyReLU activation function; the self-attention module includes a multi-head attention mechanism; the parameter configuration initializes the encoder and decoder weights through He initialization, uses Xavier initialization for the self-attention module weights, and initializes the bias term to 0; the loss function of the parameter configuration is L=λ1*L recon +λ2*L KL +λ3*L adv ; Among them, the reconstruction loss L recon Using mean square error, KL divergence loss L KL Used to control the distribution of latent variables and the adversarial loss L adv The Wasserstein distance is used, and λ1, λ2, and λ3 are weight coefficients.

4. The method according to claim 1, characterized in that: The step of determining an optimized feature set corresponding to the device data, and training a VAE-GAN model composed of a variational autoencoder and a generative adversarial network based on the optimized feature set to obtain local model parameters corresponding to the trained VAE-GAN model includes: Extracting features of the device data based on the self-attention module to obtain feature vectors corresponding to the device data, and concatenating the feature vectors to obtain a unified feature vector; The unified feature vector is weighted by a multi-head attention mechanism to obtain a feature matrix; the feature matrix is ​​added to a position encoding matrix corresponding to the relative position encoding to obtain an optimized feature set; The pre-constructed VAE is trained by the optimized feature set until the loss value of the VAE satisfies the first training end condition, thereby obtaining a trained VAE, and a potential feature set is determined by the trained VAE and the optimized feature set; Training the pre-constructed GAN by using the potential feature set until the discriminator loss value corresponding to the GAN satisfies the second training condition, thereby obtaining a trained GAN, and determining an enhanced feature set by using the trained GAN and the potential feature set; According to the enhanced feature set, the parameters corresponding to the self-attention module, the trained VAE and the trained GAN are determined respectively, and the local model parameters are determined according to the global aggregation interface specification.

5. The method according to claim 1, characterized in that The sending the local model parameters to the central server comprises: The local model parameters are sent to the central server for the central server to receive and verify the local model parameters corresponding to the multiple node devices, and then the node weight corresponding to each node device is determined according to the number of samples in the node device, and the local model parameters are aggregated through a weighted average strategy and the node weight to obtain the global model parameters.

6. The method according to claim 1, characterized in that The obtaining authenticated global model parameters from the blockchain network includes: Receiving authenticated global model parameters returned by the blockchain network; The blockchain network is used to determine the set of transaction data blocks corresponding to the global model parameters, and broadcast each transaction block in the transaction data block set to the nodes in the blockchain network; when the nodes pass the trusted node screening and the trusted nodes complete the verification of the transaction blocks, the successfully verified transaction blocks are added to the blockchain network.

7. The method according to claim 6, characterized in that The detecting, based on the global model using the global model parameters, the device data to be detected of the IoT device to which the node device is communicatively connected, and determining the detection result of the device data to be detected comprises: Loading the global model parameters obtained from the blockchain network into the VAE-GAN model to obtain a global model; Determine a feature vector corresponding to the data of the device to be detected, and input the feature vector into the global model to obtain a predicted value corresponding to the data of the device to be detected; According to the predicted value, a detection result of the device to be detected is determined; the detection result includes normal behavior and abnormal behavior.

8. A network intrusion detection device, characterized in that: The device comprises: A data acquisition module, used to obtain device data of an Internet of Things device in a wireless sensor network that is communicatively connected to the node device; the device data includes radio frequency data, fuel value, power value, speed, temperature and transaction data; A local model training module, used to determine an optimized feature set corresponding to the device data, and train a VAE-GAN model composed of a variational autoencoder and a generative adversarial network based on the optimized feature set to obtain local model parameters corresponding to the trained VAE-GAN model; the local model parameters comply with the global aggregation interface specification; the optimized feature set is obtained by concatenating feature vectors of the device data through a self-attention module; A parameter sending module, used for sending the local model parameters to a central server, so that the central server can aggregate the local model parameters corresponding to the multiple node devices to obtain global model parameters, and distribute the global model parameters to each node device and the blockchain network, so that each node device can retrain the local VAE-GAN model, and the blockchain network can authenticate the global model parameters and add them to the blockchain network; An intrusion detection module is used to obtain authenticated global model parameters from the blockchain network, detect the device data to be detected of the IoT device to which the node device is connected in communication according to a global model using the global model parameters, and determine the detection result of the device data to be detected; the global model is a VAE-GAN model jointly trained by multiple node devices and the central server.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Distributed Internet of Things intrusion detection method and system based on block chain and federated learning

    CN113794675A

  • Multi-domain DDoS attack detection method and device based on trusted federated learning

    CN115102763A

  • Federal learning-based model training method and device

    CN116957103A

  • Smart power grid federal learning method driven by block chain

    CN118747541A