Satellite network anomaly detection method based on federated learning

The federated learning-based satellite network anomaly detection method using CvT-LSTM models addresses the challenge of balancing model training and communication overhead by selecting high-quality clients, enhancing convergence speed and accuracy in satellite network security.

CN120321036AActive Publication Date: 2025-07-15CHANGCHUN UNIV OF SCI & TECH

Patent Information

Application Number
CN202510788090.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-15
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

The existing satellite network anomaly detection methods cannot effectively learn the feature patterns between data, resulting in the inability to achieve the balance between optimal model training and communication overhead.

Method used

A federated learning-based method is adopted, combined with the CvT-LSTM model for training, and a spatiotemporal feature extraction model is constructed by selecting high-quality clients to participate in model aggregation, and the tournament selection method is used to optimize client selection to reduce invalid data transmission.

Benefits of technology

It improves the convergence speed and accuracy of the global model, reduces communication costs, improves the modeling and identification capabilities of complex attack sequences, and enhances the generalization capabilities of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120321036A_ABST
    Figure CN120321036A_ABST
Patent Text Reader

Abstract

The invention discloses a satellite network anomaly detection method based on federated learning, relates to the technical field of satellite network communication, and solves the problems that an existing detection system cannot effectively learn a feature model between data and cannot realize balance of optimal model training and communication overhead. According to the method, an open source data set is preprocessed, a CvT-LSTM model is constructed, a general framework with a single server and K participants is designed, an aggregation strategy based on local data volume and anomaly detection model precision weighting is used, anomaly detection model parameters of a selected client are uploaded to a global model, and after being aggregated by the global model, the anomaly detection model parameters of the selected client are obtained. And issuing to all clients, and the like. According to the method, interference of low-quality data is avoided, transmission and calculation of invalid data are reduced from the source, the communication cost is remarkably reduced, only the high-representativeness client participates in global model aggregation, and the communication burden is reduced to the maximum extent while the model performance is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of satellite network communication, and specifically relates to a satellite network anomaly detection method based on federated learning. When this detection method is used to randomly select clients for training, it can achieve the balance between optimal model training and communication overhead. Background Art

[0002] With the continuous expansion of satellite Internet application scenarios and the increasing complexity of network connections, the network security threats it faces are also intensifying. The construction and protection of satellite Internet will become a core issue in the "space-air-ground integration" system in the 6G era. Satellite Internet uses advanced satellite communication technology to connect various terminals to the global network, as if setting up mobile base stations in the air. These base stations can cover areas that cannot be reached by ground base stations and provide uninterrupted and efficient network interaction. To protect the normal communication of satellite networks, it is necessary to design an anomaly detection scheme in advance.

[0003] Traditional intrusion detection systems often rely on continuous data exchange and analysis, which may lead to increased bandwidth usage and operating costs, bringing additional challenges to satellite communication systems. The satellite network anomaly detection methods based on machine learning either only consider spatial or temporal features, or do not fully extract the spatio-temporal features between data, so they cannot effectively learn the feature patterns between data. The satellite network anomaly detection methods based on federated learning randomly select clients for training and cannot achieve the balance between optimal model training and communication overhead. To solve the above problems, the present invention designs a satellite network anomaly detection method based on federated learning. Summary of the Invention

[0004] The present invention provides a satellite network anomaly detection method based on federated learning to solve the problems that existing detection systems cannot effectively learn the feature model between data and cannot achieve the balance between optimal model training and communication overhead.

[0005] The satellite network anomaly detection method based on federated learning is implemented by the following steps:

[0006] Step 1: Preprocess the data set and use the sliding window method to construct a message sequence with a fixed length;

[0007] Step 2: Construct a CvT-LSTM model as an anomaly detection model for the anomaly detection task of the message sequence;

[0008] The CvT-LSTM model includes a spatial feature extraction module and a temporal feature extraction module;

[0009] The spatial feature extraction module is used to extract spatial features from the input message sequence, and further extract temporal features from the spatial features by using the temporal feature extraction module, and implement multi-class anomaly detection and binary-class anomaly detection through a fully connected layer, and take the class corresponding to the maximum value as the final detection result;

[0010] Step 3: Set up a federated learning framework with a single server and clients; Select clients for each round of federated training;

[0011] Step 4: Adopt an aggregation strategy based on the weighted combination of local data volume and local model accuracy, upload the anomaly detection model parameters of the selected clients to the global model, and after aggregation by the global model, distribute them to all clients.

[0012] Advantages of the present invention:

[0013] The method of the present invention dynamically optimizes and screens according to the training performance and data quality of the clients, improving the convergence speed and accuracy of the global model; constructs a spatio-temporal feature extraction model integrating CvT and LSTM, enhancing the modeling and recognition ability for complex attack sequences.

[0014] The method of the present invention, aiming at the communication overhead problem in federated learning, uses the tournament selection method to intelligently screen high-quality clients, avoiding the interference of low-quality data, reducing the transmission and calculation of invalid data from the source, and significantly reducing the communication cost. Only allowing highly representative clients to participate in the global model aggregation minimizes the communication burden while ensuring the model performance.

[0015] In terms of performance, the method of the present invention improves the convergence speed of the global model and enhances the accuracy of the final model. On the UNSW-NB15 dataset, compared with the existing DFL-ID, Precision, Recall, and F1-Score are respectively improved by 1.9%, 2.8%, and 2.9%. Then, by integrating CvT and LSTM, the anomaly detection ability of time series data is enhanced, strengthening the modeling ability of the model for complex time series data. On the STIN satellite network security dataset, the multi-class accuracy rate is improved by 3%-25% compared with all methods. Compared with the Centralized DAE, the method of the present invention still improves by 2.8% in the multi-class classification task even in the federated learning scenario, and also shows better generalization ability. Description of the Drawings

[0016] Figure 1 It is a schematic diagram of the satellite communication architecture of the satellite network anomaly detection method based on federated learning described in the present invention;

[0017] Figure 2 Schematic diagram of the model of the satellite network anomaly detection method based on federated learning according to the present invention;

[0018] Figure 3 Schematic diagram of the structure of the spatial feature extraction module in the CvT-LSTM model;

[0019] Figure 4 Schematic diagram of the principle of the LSTM time feature extractor;

[0020] Figure 5 Schematic diagram of the federated learning architecture based on genetic optimization in satellite communication according to the present invention;

[0021] Figure 6 Schematic diagram of the principle of the federated learning process according to the present invention;

[0022] Figure 7 Experimental result graph of the comparison of the binary classification model of the present invention on the UNSW-NB15 dataset;

[0023] Figure 8 Experimental effect graph of the comparison of the multi-classification model of the present invention on the UNSW-NB15 dataset; wherein, (a) is the comparison effect graph of the CvT-LSTM model of the present invention and other models in terms of Accuracy, (b) is the comparison effect graph of the CvT-LSTM model of the present invention and other models in terms of Precision, (c) is the comparison effect graph of the CvT-LSTM model of the present invention and other models in terms of Recall, and (d) is the comparison effect graph of the CvT-LSTM model of the present invention and other models in terms of F1-Score;

[0024] Figure 9 Experimental effect graph of the comparison of the multi-classification model of the present invention on the STIN dataset; wherein, (a) is the comparison effect graph of the CvT-LSTM model of the present invention and other models in terms of Accuracy, (b) is the comparison effect graph of the CvT-LSTM model of the present invention and other models in terms of Precision, (c) is the comparison effect graph of the CvT-LSTM model of the present invention and other models in terms of Recall, and (d) is the comparison effect graph of the CvT-LSTM model of the present invention and other models in terms of F1-Score. Detailed implementation manners

[0025] Detailed implementation manner 1. In combination with Figures 1 to 6This embodiment describes a satellite network anomaly detection method based on federated learning. In this method, a time series anomaly detection model (CvT-LSTM model) that fuses CvT and LSTM is first proposed, that is, the Convolutional vision Transformer (CvT) is combined with the Long Short-Term Memory network (LSTM) for anomaly detection of time series data in the federated learning framework. This model can ensure data security while using the genetic algorithm to select high-quality clients to participate in aggregation, thereby improving the convergence speed of the global model and enhancing the final performance of the model.

[0026] As Figure 1 shown, the satellite network anomaly detection method described in this embodiment is applicable to a three-layer satellite communication system architecture, which can be divided into a satellite network, a ground station, and end users. In this architecture, the satellite network part includes inter-satellite communication, communication between the satellite and the ground station, forwarding user instructions to the ground station communication, and communication between the satellite and the end user; the ground station part includes inter-ground station communication, communication between the ground station and the satellite network, and communication between the ground station and the end user; the end user part includes communication between the user and the satellite and communication between the user and the ground station. First, the CvT-LSTM model is deployed on each client and the central server. Using the federated learning method can protect the data privacy of each client. Through the CvT-LSTM model, the spatial features and time features between data can be fully extracted to generate richer feature representations for higher accuracy. Then, each participant (client) uses the local dataset to train the model locally, and finally uploads the weights to the cloud. After being aggregated by the global model, each participant downloads the updated model.

[0027] The specific steps of the satellite network anomaly detection method in this embodiment are as follows:

[0028] Step 1: Preprocess the UNSW-NB15 open-source dataset and the STIN open-source dataset, including data cleaning, missing value filling, normalization, and sampling processing;

[0029] Step 1-1: Data cleaning. Convert non-numeric items in the data into corresponding numeric forms through a mapping method, including the label column; use 0 to represent "normal" and 1 to represent "abnormal", so as to convert the categorical data into numerical data;

[0030] Step 1-2: Missing value filling. Check whether there are missing values or invalid characters in the entire dataset. If so, replace them with 0;

[0031] Steps 1-3, Normalization: Converting all processed data to a specific range, i.e., between 0 and 1, can help the classifier eliminate biases in the source data. Normalization is completed using the min-max normalization algorithm.

[0032] Steps 1-4, Sampling: Judging according to the ratio of normal and abnormal labels. If the ratio is greater than 5:1, data oversampling is performed to maintain data balance.

[0033] Steps 1-5, Constructing a message sequence of fixed length using the sliding window method , as shown in the following formula: ; where , represents the sequence length, represents the feature dimension (number of channels) at each time step.

[0034] To maintain the sample quantity and have a low repetition rate, the sequence length (sliding window size) is set to 20, and the sliding length is set to 10. The last sequence with a length less than 20 is discarded. Each sequence corresponds to a label SeqType. If all data points in the sequence are normal, the entire sequence is labeled as normal (SeqType = 0). As long as there is attack information in the sequence, the sequence is labeled as 1 (SeqType = 1).

[0035] Step 2, Model Design; as Figures 2 to 4 shown.

[0036] Figure 2 There is a CvT-LSTM model in a federated learning framework, and this model serves as the local model (anomaly detection model). The central server first initializes the global model and distributes the initialized model parameters to each client. Each client uses the local dataset to train the CvT-LSTM model, extracts spatial and temporal features, and then uploads the weight parameters of each model, i.e.: , to the global model, and updates the global model through an aggregation strategy weighted by the local data volume and the local model accuracy, where r represents the communication round. Finally, the parameters of the updated global model are distributed to each client, i.e., 。Each client updates the local anomaly detection model according to the parameters sent by the global model, and then trains the updated anomaly detection model using the local dataset to update the parameters of the anomaly detection model. Iterate continuously according to the above process, and end the iteration when the specified communication round r = 50. Each layer of the anomaly detection model, such as the convolutional projection layer, multi-head self-attention layer, LSTM layer, etc., is used to extract different types of features, and the flow and optimization of information are ensured through residual connections and weight matrices.

[0037] Figure 3 In, the convolutional projection layer is used to extract the spatial features of the input. After feature mapping through 1×1 convolution, the mapped data is split into h parts, where h is the number of heads of the multi-head self-attention mechanism. Each head generates query ( ), key ( ), and value ( ) matrices through independent linear transformations, that is: , and then calculate the attention scores on each head. Through the similarity calculation, scaling, and softmax normalization of and , the weight distribution is obtained. After weighted fusion with the matrix, a single-head output is generated. Finally, the outputs of all heads, that is, , are concatenated and fused through a linear transformation (multiplied by the weight matrix ) to form the final multi-head attention representation Z. This result will be passed to the downstream feed-forward layer to enhance the non-linear features through the ReLU activation function, and finally output the prediction result of the anomaly detection model.

[0038] In this embodiment, the CvT-LSTM model (anomaly detection model) is used for the anomaly detection task of sequence data; the CvT-LSTM model includes a spatial feature extraction module (CvT) and a temporal feature extraction module (LSTM); The spatial feature extraction module is composed of a convolutional projection layer, a multi-head self-attention layer, a feed-forward layer, and a residual connection layer; the temporal feature extraction module is implemented by an LSTM model; The convolutional projection layer operates on the input sequence as follows;

[0039] First, a one-dimensional convolutional layer (Conv1D) is used to perform feature projection on the input message sequence to extract the local features of the time series and generate a higher-dimensional representation using the correlation between time steps. Specifically, taking the message sequence as the input sequence, after convolutional transformation of the dimension, the transformed feature representation is obtained, is the new feature dimension, and is the new input sequence length.

[0040] Then, a batch normalization layer is introduced to normalize the feature distribution, improve the training stability, and accelerate convergence. The feature map is obtained after the convolution operation , and the calculation formula is as follows:

[0041] Among them, is the weight matrix of the convolutional layer, is the weight of the convolution kernel, that is: it represents the output channel , the input channel and the position of the convolution kernel weight, represents the value of the input sequence at time step and the input channel , among which, is the stride, is the padding size, is the bias of the output channel , is the window size of the convolution kernel.

[0042] Finally, non-linearity is introduced into the network so that complex patterns between data can be learned. The ReLU activation function is used for non-linear transformation, which is defined as: ; Among them, is the local feature of the output obtained after applying the ReLU activation function, that is, the feature map after convolution after being activated by ReLU. It will convert the negative part of the convolution result to zero, thereby strengthening the network's response to positive features and avoiding the vanishing gradient problem. The local features extracted by convolution are projected into a new feature space, which is expressed by the following formula: ; Among them, is the feature after the projection transformation, representing the input local feature representation in the new space. is a weight matrix that controls the linear transformation from the input feature space to the output feature space, that is: maps the feature dimension to the model dimension .

[0043] The multi-head self-attention layer uses the feature after the convolution projection as the input of the multi-head attention mechanism , maps the feature through convolution to obtain rich intermediate features, and generates queries ( ), keys (K), and values (V), and the formula is as follows: ; Among them , , is a learnable weight matrix, is the dimension of the query, key, and value.

[0044] Calculate the attention weights, the formula is as follows: ; Among them, calculate the similarity between the query and the key, is the scaling factor to ensure numerical stability.

[0045] In this embodiment, multiple parallel self-attention mechanisms (heads) are adopted. Each head captures information in different representation subspaces by learning different weights. The outputs of multiple heads will be concatenated together and linearly transformed; the calculation formula is: ; Among them, each is the output of the multi-head self-attention mechanism, , where h is the number of attention heads, is the projection matrix of the output.

[0046] Adopt the said feed-forward layer (Feed Forward laye, FFL) for further feature extraction.

[0047] First, perform a linear transformation, then perform a non-linear activation on the result of the linear transformation, and finally, through a second linear transformation, map the feature dimension back to the original dimension. The formula is as follows: ; Among them, is the dimension of the middle layer of FFL, and the final output feature of FFL , maintains the same dimension as the input message sequence X.

[0048] The said residual connection layer performs a residual connection on the output of the multi-head self-attention mechanism and the input (message sequence X): ; Among them, represents the output of the residual connection after the attention mechanism, that is, the result after adding the output of the attention mechanism, which is used to enhance the gradient flow and stabilize the model training. Perform a residual connection on the output feature of the FFL layer as well: ; Among them, Represents the output of the residual connection after the feed-forward layer, that is, the result after adding the output of the feed-forward layer, which is used to further enhance the gradient flow and stabilize the model training.

[0049] In this embodiment, the features output by the residual connection layer are normalized, that is: for and two residual operations to obtain the output features , and the output features are normalized to obtain the spatial features ; ; Among them, LayerNorm normalizes the features of each input sample;

[0050] The time feature extraction module uses the spatial features output by CvT as the input of LSTM, and further extracts time features from the data after fully extracting the context relationship.

[0051] As Figure 4 shown, in this embodiment, Figure 4 is the structure of an LSTM (Long Short-Term Memory Network) cell. LSTM is a recurrent neural network (RNN) structure that can capture long-term dependencies. The main parts in the figure include: Input gate ( ): Determines how much of the input candidate state at the current moment is added to the current cell state; Forget gate ( ): Controls how much of the cell state at the previous moment is retained; Cell state ( ): Through the action of the forget gate and the input gate, updates the cell state at the current moment;

[0052] Output gate ( ): Determines how much of the cell state at the current moment flows to the hidden state . LSTM combines three gates (input gate, forget gate, and output gate) with the tanh activation function to effectively manage and update information, solving the problem of vanishing gradients faced by traditional RNNs in long sequence learning. In the figure, , are the input vectors at the current time step t and the previous time step t - 1, usually represented as feature inputs; , , is the hidden state of the LSTM cell, which can also be called short-term memory, and it is also one of the network outputs. , , is the cell state, which is used for long-term memory storage and is one of the core advantages of LSTM. is the candidate cell state, the intermediate state generated by the tanh function, which is used to update . The symbol represents element-wise multiplication, which is often used to control the flow of information in the gating mechanism. The symbol represents element-wise addition, which is used to combine the old state and the new candidate state.

[0053] In this embodiment, the forget gate has an output range of [0, 1], where 0 means complete forgetting and 1 means complete retention: ; where and represent the weights and biases of the forget gate, is the output of the LSTM neuron at the previous time step, represents the sigmoid activation function to ensure that remains between [0, 1].

[0054] The input gate determines how much of the input information at the current time step needs to be retained in the memory cell state. The memory state of the feature vector is: ; where and are the weights and biases of the storage unit network. The weight of the temporary memory state relative to the entire memory stream is represented by as: ; where and are the weights and biases of the input gate. Combining the outputs of the forget gate and the input gate, the memory cell state is updated to obtain the memory state of the LSTM at this moment : ;

[0055] The output gate determines the hidden state (output) at the current time step, controls the output part, and at the same time maps the memory cell state to between [-1, 1] through the tanh function: ; ; where and are the weights and biases of the output gate. Obtain the unactivated output through the fully connected layer ; ; Among them, and are the weight vector and bias vector of the fully connected layer.

[0056] Finally, the fully connected layer is used to complete multi-class anomaly detection and binary-class anomaly detection respectively, that is, M = 2 or M . The calculation formula of the SoftMax function is as follows: ; Among them, M is the number of categories, is the unnormalized score (logit) of the model for the t-th class, which is transformed into a positive weight through the exponential function . The denominator is the sum of the score exponents for all categories m = 1, 2,..., M, which is used for normalization to ensure that the sum of all predicted probabilities is 1; where t represents the target category for which the probability is being calculated, and m is the index used to iterate over all categories. Finally, represents the predicted probability that the input belongs to the t-th class, and the category corresponding to the maximum value is used as the final detection result.

[0057] Use for the loss function, and the cross-entropy loss function is adopted. Its calculation formula is as follows: ; Among them, is the sign function (0 or 1), which takes 1 if the true category of sample i is equal to m, otherwise takes 0, represents the predicted probability that sample i belongs to category m, and B represents the number of samples.

[0058] Step 3, Set up a general federated learning (FL) framework with a single server and participants (clients), and select clients for each round of federated training; the specific process is as follows:

[0059] Step 3-1, Randomly generate multiple client selection schemes as the initial population, and set the population size to Q. Each individual is a binary vector of length , which is used to represent the client selection scheme, where K represents the total number of clients, that is, the total number of anomaly detection models; , indicating the selection status of the k-th client. Specifically, indicates that the k-th client is selected to participate in the training, indicates that it is not selected to participate in the global aggregation;

[0060] Step 3-2: In each iteration, select individuals with higher fitness to generate the next generation; the tournament selection method is adopted, and the formula for tournament selection is as follows: ; is the best client selection scheme selected from the tournament, represents examining each individual in the tournament candidate set in turn, where each individual represents a complete client selection scheme. To distinguish different individuals, the client selection scheme can also be represented as , that is, the q-th client selection scheme; , each is a binary vector of length K, composed of multiple elements which represents whether the k-th client is selected to participate in the training, and is a set of multiple . f( ) is the fitness function value of the client selection scheme, used to evaluate the quality of this scheme. Each time, randomly select 3 individuals from the population, and select the one with the highest fitness to enter the next generation. The selected individuals perform a crossover operation, randomly select two crossover points and exchange the middle part to explore more solution spaces. The crossover formula is as follows: ; where, to are the client selection information after crossover. By performing a mutation operation on some bits of the individual (such as changing from 1 to 0, or from 0 to 1), the diversity of the population is increased to avoid the algorithm falling into a local optimal solution. The mutation formula is: ;

[0061] Randomly flip some bits in the client selection scheme. After iterating 10 times or meeting the convergence condition, output the optimal scheme. The fitness of the individual is calculated by the evaluation function, and output the client selection scheme with the maximum fitness function. The formula is: ; where, S = { } is the set of selected clients, satisfying , and its size is NM, where indicates that the u-th client is selected, ; is the accuracy of the selected u-th client, and the fitness function is calculated according to the accuracy of the selected clients.

[0062] Step 4: Set the total size of the training dataset to be represented as: ; is the total amount of data of the selected clients, i.e.: ; adopt an aggregation strategy based on the weighted local data volume and local model accuracy, and upload the anomaly detection model parameters of the selected clients to the global model , after being aggregated by the global model, it is sent to all clients. The aggregation formula is: ; Among them, is the anomaly detection model parameters locally trained by the selected \(u\)-th client, \(D\) u is the dataset held locally by the selected \(u\)-th client; is the data volume of the selected \(u\)-th client, is the accuracy rate of the selected \(u\)-th client, is the number of selected clients, is the aggregated global model parameters;

[0063] As Figure 5 and Figure 6 shown, in this embodiment, Figure 5 is a federated learning framework in a satellite network. The figure includes multiple satellite nodes (anomaly detection model 1 to anomaly detection model K), and each satellite conducts local model (anomaly detection model) training. High-quality clients participate in the aggregation of the global model, and the ground station is responsible for aggregating the local models and updating the global model. The upload and download traffic respectively represent the data transmission between the client and the ground station. This framework realizes distributed model training and optimization through the satellite network.

[0064] Figure 6 is the interaction process between the client and the server in federated learning. The client locally trains the anomaly detection model and uploads it. The server receives the update and aggregates it to generate the global model, and then sends the initial and updated global models back to the client.

[0065] Specific Embodiment 2. Combining Figure 7 and Figure 9 to illustrate this embodiment. This embodiment is an experimental example of the satellite network anomaly detection method based on federated learning described in Specific Embodiment 1:

[0066] In this embodiment, Figure 7The binary classification results on the UNSW-NB15 dataset are compared between the present invention and other methods, including DFL-ID (Deep Federated Learning IDS), the centralized deep autoencoder (Centralized DAE) method, and individual models such as standalone LSTM and CvT. The model of the present invention achieved an accuracy of 99.95% in the binary classification task, and the precision, recall, and F1-score were all higher than 99%. Specifically, compared with DFL-ID, the precision of the present invention was increased by at least 1.5%, the recall was increased by 2.0%, and the F1 score was increased by 2.0%. Even compared with the centralized DAE, the performance of the present invention was better. The precision was slightly lower than that of the centralized DAE, but the precision / recall was similar, indicating that the CvT-LSTM architecture of the present invention is very effective. The excellent performance of the present invention on UNSW-NB15 is attributed to its ability to capture long-range temporal dependencies (through LSTM) and global feature interactions (through Transformer-based attention), which cannot be achieved by simpler models such as standalone LSTM or CNN. The inclusion of both spatial and temporal feature extractors enables the present invention to learn complex attack patterns evolving over time, thus having an advantage in terms of accuracy. In addition, ablation experiments were conducted, and the results showed that the model proposed by the present invention had a performance improvement of about 1% compared with using the CvT and LSTM methods alone, further demonstrating that the model of the present invention can fully extract multi-dimensional features when dealing with complex datasets and improve the classification accuracy.

[0067] Figure 8For the results of multi-class classification on the UNSW-NB15 dataset, the present invention can fully learn the temporal and spatial features among the data and achieves an accuracy of 98.87%. The results show that anomaly detection methods that only consider temporal features or spatial features separately, such as Random Forest (RF), Artificial Neural Network (ANN), Gated Recurrent Unit (GRU), and Long Short-Term Memory Network (LSTM), perform unsatisfactorily, and the best result is only 75%. Even the ensemble method combined with the Sequential Forward Selection (SFS) algorithm, although considering both temporal and spatial features simultaneously, still fails to fully extract complex spatio-temporal patterns, and the best accuracy is only 79%. In contrast, the present invention captures spatio-temporal dependencies more comprehensively by effectively combining CvT and LSTM. Even compared with the centralized Deep Autoencoder (DAE), the accuracy is increased by 1.8%. When DAE is combined with the federated learning method, although an accuracy of 92% is achieved, the present invention still has a 6.8% improvement compared to DFL-ID. In addition, combining (a), (b), (c), and (d) in the figure, the present invention reaches 98.87%, 98.87%, and 98.86% in precision, recall, and F1-score respectively, demonstrating comprehensive advantages in multiple evaluation metrics. In the ablation experiment, using the CvT method alone only considering spatial features achieves an accuracy of 96.79%; while LSTM only considering temporal features has an accuracy of only 75%. The present invention achieves the highest accuracy by simultaneously learning spatio-temporal features, with a 2% and 23% improvement compared to CvT and LSTM respectively.

[0068] Figure 9 For the results of multi-class classification on the STIN dataset. The STIN security dataset contains various types of attacks from ground and satellite networks. Due to reasons such as resource limitations, differences in attack tolerance, limited computing power, and scarcity of satellite network data, the methods in existing research are mostly applied to ground networks and fail to fully adapt to the satellite network environment. However, the present invention still performs excellently on the STIN dataset. Combining (a), (b), (c), and (d) in the figure, the performance of all four metrics exceeds 96%, and compared with all the methods in the figure, there is at least a 6% improvement in accuracy, precision, recall, and F1-score.

[0069] The technical features of the above-described embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.

[0070] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A satellite network anomaly detection method based on federated learning, characterized in that: This method is implemented by the following steps: Step 1: Preprocess the dataset and construct a message sequence of fixed length using the sliding window method; Step 2: Construct a CvT-LSTM model as an anomaly detection model for the anomaly detection task of the message sequence; The CvT-LSTM model includes a spatial feature extraction module and a temporal feature extraction module; The spatial feature extraction module is used to extract spatial features from the input message sequence, and the temporal feature extraction module is used to further extract temporal features from the spatial features, and multi-class anomaly detection and binary-class anomaly detection are realized through a fully connected layer, and the class corresponding to the maximum value is used as the final detection result; Step 3. Set up a federated learning framework with a single server and clients; Select clients for each round of federated training; Step 4: Adopt a strategy based on weighted average to upload the anomaly detection model parameters of the selected clients to the global model, which are aggregated by the global model and then sent to all clients.

2. The satellite network anomaly detection method based on federated learning according to claim 1, wherein: In Step 1, the specific process of preprocessing is as follows: Step 11: Data cleaning, converting non-numeric items in the dataset into corresponding numeric forms through a mapping method; using 0 to represent normal and 1 to represent anomaly, and converting the data into numerical data; Step 12: Missing value filling, checking whether there are missing values or invalid characters in the dataset, and if so, replacing them with 0; Step 13: Normalization, converting the processed data to between 0 and 1; Step 14: Sampling, judging according to the ratio of normal labels to anomaly labels. If the ratio is greater than 5:1, data oversampling is performed to maintain data balance.

3. The satellite network anomaly detection method based on federated learning according to claim 1, characterized in that: In Step 1, a message sequence of fixed length is represented by the following formula: ; Wherein, , is the length of the message sequence, is the feature dimension of each time step; each message sequence corresponds to a tag SeqType. If all data points in the message sequence are normal, the message sequence is marked as normal, SeqType = 0; if there is attack information in the message sequence, the message sequence is marked as 1, SeqType = 1.

4. The satellite network anomaly detection method based on federated learning according to claim 1, wherein: In Step 2, the spatial feature extraction module includes a convolutional projection layer, a multi-head self-attention layer, a feed-forward layer, and a residual connection layer; The message sequence Extract local features after dimension transformation by convolution, project the local features into a new feature space, and obtain the features after projection transformation ; After Feature As the input of the multi-head self-attention layer, multiple parallel self-attention mechanisms are used to learn different weights, and finally the output of the multi-head self-attention layer is linearly transformed to obtain the output feature ; The feedforward layer is used to extract the features and output the spatial features after passing through the residual connection layer .

5. The method for anomaly detection in a satellite network based on federated learning according to claim 4, wherein: In the feedforward layer, first, linear transformation, non-linear activation, and quadratic linear transformation are performed on the output features of the multi-head self-attention layer to obtain features consistent with the dimension of the input message sequence ; the output features of the multi-head self-attention layer are subjected to residual connection with the input message sequence to obtain the output features of the residual connection ; ; The features output by the feedforward layer and the output features are subjected to a residual connection, and the output features For the said output features and perform a residual operation and normalization to obtain spatial features .

6. The satellite network anomaly detection method based on federated learning according to claim 1, characterized in that: In Step 3, the federated learning framework realizes distributed model training and optimization through a satellite network; it includes multiple satellite nodes, each satellite node conducts local anomaly detection model training, high-quality clients participate in the aggregation of the global model, the ground station is responsible for aggregating the local anomaly detection models and updating the global model, and data transmission is realized by uploading and downloading traffic between the client and the ground station.

7. The satellite network anomaly detection method based on federated learning according to claim 1, wherein: In Step 3, NM clients are selected for each round of federated training. The specific process is as follows: Step 3.1: Randomly generate multiple client selection schemes as the initial population, and set the population size to Q; each individual is a binary vector of length for representing the client selection scheme; K represents the total number of clients, , indicating the selection status of the k-th client; indicating that the k-th client is selected to participate in training indicating not being selected to participate in global aggregation; ​ Step 32: Adopt the tournament selection method to select the optimal client selection scheme, that is, output the client selection scheme with the largest fitness function value; it is expressed by the following formula: ; wherein, { } is the set of selected clients, satisfying , and its size is NM, indicates that the u-th client is selected, ; is the accuracy of the u-th selected client, and the fitness function is calculated according to the accuracy of the selected clients.

8. The method for anomaly detection in a satellite network based on federated learning according to claim 1, wherein: In Step 4, the aggregation formula based on the local data volume and the local model accuracy weighting is: ; Among them, are the anomaly detection model parameters locally trained by the u-th client, and D u is the dataset held locally by the selected u-th client, is the data volume of the u-th client, is the accuracy rate of the u-th client, is the number of selected clients, are the aggregated global model parameters.

9. The method for anomaly detection of satellite network based on federated learning according to claim 1, characterized in that: The detection method is applied to a three-layer satellite communication system architecture, which is divided into a satellite network, a ground station, and terminal users; in the system architecture, the satellite network part includes inter-satellite communication, satellite-ground station communication, forwarding user instructions to ground station communication, and satellite-terminal user communication; The ground station part includes ground station-ground station communication, ground station-satellite network communication, and ground station-terminal user communication; The terminal user part includes user-satellite communication and user-ground station communication; First, an anomaly detection model is deployed on each client and the central server, and the federated learning method is used to protect the data privacy of each client, and the spatial features and temporal features between the data are extracted through the anomaly detection model; Then, each participant locally trains an anomaly detection model using the local dataset; Finally, the weights of the trained anomaly detection model are uploaded to the cloud. After aggregation by the global model, each participant downloads the updated anomaly detection model.

Citation Information

Patent Citations

  • Transform anomaly detection method based on spatio-temporal characteristics

    CN115618196A

  • Edge anomaly traffic detection method based on multi-scale aggregation Transform

    CN119202840A

  • Unsupervised GAN-based intrusion detection system using temporal convolutional networks, self-attention, and transformers

    US20240250963A1

  • Large language models for predictive modeling and inverse design

    WO2025075756A1

Cited By

  • Federal learning-based terminal Agent knowledge collaborative updating method and system, terminal, medium and product

    CN121902918A

  • Terminal agent knowledge collaborative updating method and system based on federated learning, terminal, medium and product

    CN121902918B