Unmanned aerial vehicle intrusion detection method and system, storage medium and program product

By establishing a cross-modal feature extraction model, the network and physical features of the drone are extracted using pyramid convolution layers, expanded convolution layers and Transformer encoder, and the features are fused through the Cross Attention mechanism, and finally classification is used by KAN Layer, which solves the problem of insufficient drone intrusion detection performance in the existing technology, achieving higher detection accuracy and robustness.

CN120032176APending Publication Date: 2025-05-23UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510200681.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

现有的无人机入侵检测系统无法有效利用无人机的物理和网络特征,导致检测性能不足,且受限于过时的数据集,缺乏更全面和准确的特征表示。

Method used

A drone intrusion detection method is proposed. By establishing a cross-modal feature extraction model for network and physical data, the features are extracted using pyramid convolutional layer, expanded convolutional layer and Transformer encoder, and the network and physical modal features are fused through the Cross Attention mechanism, and finally classification is used using KAN Layer.

Benefits of technology

The detection performance of drone intrusion detection has been significantly improved, with F1 score reaching 99.57%, Accuracy reaching 99.74%, Precision reaching 99.32%, Recall reaching 99.82%, and effectively responding to Byzantine attacks to enhance the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032176A_ABST
    Figure CN120032176A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle intrusion detection method and system, a storage medium and a program product, and relates to the technical field of intrusion detection of an unmanned aerial vehicle group, and the method comprises the steps: S1, carrying out the parameter initialization of a network data branch and a physical data branch through a model, and generating an initial weight; s2, setting hyper-parameters to train the model; s3, model parameters are updated, verification loss values are calculated, and then the verification loss values are sent to a parameter server; s4, comparing the received verification loss values by the parameter server, and selecting the minimum verification loss value and the model parameter corresponding to the minimum verification loss value; s5, if the models of all the nodes pass verification, the parameter server updates the global model to be a current iteration model; and if no model passes the verification, keeping the global model of the previous round unchanged. According to the method, novel convolution Transform networks are respectively established for two kinds of heterogeneous data so as to carry out feature extraction; the KAN Layer is used for replacing a traditional mlp classifier to better process nonlinear mapping.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intrusion detection of drone swarms, and in particular to a drone intrusion detection method, system, storage medium and program product. Background Art

[0002] Unmanned aerial vehicles are rapidly gaining popularity in many fields such as smart agriculture, military operations, and transportation logistics, and have received tremendous attention. UAVs are usually equipped with various sensors, actuators, and processing units to assist in performing specific tasks. They use sensors to collect measurement data, receive command signals from ground controllers, and then perform corresponding physical actions, which makes the UAV system have both cyber information attributes and physical information attributes. As a cyber-physical system (CPS), UAV networks face serious security vulnerabilities due to their inherent mobility and temporary connectivity.

[0003] In order to address the security challenges of unmanned aerial vehicle systems, a large number of researchers have conducted research on intrusion detection systems (IDS). The developed IDS are designed to detect replay attacks, hijacking attempts, spoofing attacks, false data injection (FDI) attacks, and denial of service (DoS) attacks. However, existing IDSs have limitations, namely, they do not consider drones as network and physical information systems, and only rely on the physical characteristics of drones (e.g., roll, pitch, yaw angle, speed, temperature, etc.) or network characteristics (e.g., IP address, MAC address, port number, etc.), resulting in insufficient detection performance. In addition, the literature on IDS is also limited by the use of outdated datasets such as NLS KDD99 and CICIDS2017. It was not until 2024 that Samuel Chase Hassler et al. used both physical and network heterogeneous data in IDS for the first time, and contributed a new dataset Dataset-ITS. However, this work simply combined the heterogeneous data and input them into the model, uniformly extracted features, and output the detection results after training. The detection performance of the final model still has a lot of room for improvement. Therefore, we consider whether we can build a better network model to integrate two kinds of heterogeneous feature information, so as to obtain a more comprehensive and accurate feature representation, while significantly improving the generalization ability of the model. Summary of the invention

[0004] The purpose of the present invention is to overcome the deficiencies of the prior art and to provide a drone intrusion detection method, system, storage medium and program product.

[0005] The objective of the present invention is achieved through the following technical solutions:

[0006] In a first aspect, the present invention discloses a method for detecting intrusion of a drone, comprising the following steps:

[0007] S1, model initialization, the model initializes the parameters of the network data branch and the physical data branch to generate initial weights;

[0008] S2, hyperparameter setting, setting hyperparameters for training the model;

[0009] S3, distributed training, updating model parameters and calculating validation loss values, and sending the updated model parameters and calculated validation loss values ​​to the parameter server;

[0010] S4, aggregation and verification of parameter servers. The parameter server compares the received verification loss values ​​and selects the minimum verification loss value and its corresponding model parameter.

[0011] S5. Global model update and verification. If the models of all nodes pass the verification, the parameter server updates the global model to the model of the current iteration; if no model passes the verification, the global model of the previous round remains unchanged.

[0012] Based on the first aspect, in step S1, the initial weight includes a first weight and the second weight Then calculate the initial loss weight through the initial weight Where D train Represents the training data set, which includes the following steps:

[0013] S11, the network data branch is the network branch data through the pyramid convolution layer PyramidConv1d, through the formula Y cyb,i =f(W i *X cyb +b i ) extracts network modal features, where W i represents the weight of the i-th convolution kernel that needs to be initialized, b i represents the constant corresponding to the i-th convolution kernel, f is the ReLU activation function, * represents the convolution operation, X cyb is the network modality data, Y cyb,i is the first network modality feature;

[0014] The physical data branch is the physical modal data. The first physical modal feature is extracted through the dilated convolution layer DilatedConv1d. The formula of the dilated convolution is Y phy =f(W *d X phy +b), where W is the convolution kernel weight that needs to be initialized, *d is the dilation convolution operation, d represents the dilation rate, and X phy is the physical modal data, Y phy is the first physical mode characteristic, b represents a constant;

[0015] S12, process the first network modality feature Y through the Transformer encoder cyb,i and the first physical mode characteristic Y phy , obtain the second network modality feature and the second physical modality feature, where the first network modality feature and the first physical modality feature are respectively applied with the self-attention mechanism, and the formula for calculating the attention weight is: Where Q is the Query matrix, K is the Key matrix, and V is the Value matrix, which are respectively input by embedding and the learnable weight matrix (W Q ,W K ,W V ) and multiply them to get, d k is the dimension of the key matrix used to scale the attention scores, T is the transpose;

[0016] S13. Use Cross Attention Mechanism The extracted second network modality features and second physical modality features are integrated, where Q phy Indicates that the physical modal data X phy As the Query matrix, K cyber Indicates that the network modal data is used as the Key matrix, V cyber Indicates that the network modality data is used as the Value matrix;

[0017] S14, through two layers of fully connected layers, the fusion features are mapped to the classification KAN Layer for output, and the calculation formula is: Among them, W base is the base linear layer weight, B i (x) is the B-spline basis function used to capture features in a finer granularity, W spline,i is the B-spline basis function B i (x), x represents the fusion feature, and y represents the output.

[0018] Based on the first aspect, step S2 specifically includes: setting hyperparameters during the training process, including the number of iterations num_epochs, the batch size batch_size, the learning rate learning_rate, the hidden layer dimension hidden_dimension and the patience value patience of early stopping.

[0019] Based on the first aspect, the step S3 specifically includes: in the tth (t∈[1,T'])th iteration, the local worker i (i=1,…,U) uses the AdamW optimization algorithm Update the first model parameters Get the second model parameters Where T' represents the number of iterations, U represents the last local worker, Represents the calculation of the minimum loss function, represents the third weight, represents the fourth weight; then local worker i calculates the first validation loss value on the validation set Where D val Represents the validation data set; then the second model parameter and the first validation loss are sent to the parameter server for aggregation.

[0020] Based on the first aspect, step S4 specifically includes: the parameter server compares the received first verification loss value Save the minimum first validation loss value and its corresponding second model parameter to obtain the second validation loss value and the third model parameter Then the parameter server calculates the third model parameters The mean S is the third model parameter The total number of samples forms the global model parameters During this process, the third model parameter is received After that, the parameter server also needs to use the validation data set D val Verify it, if It is identified as a Byzantine attack behavior, and the third model parameter is filtered out. The parameter server will query the next third model parameter until all local workers are traversed.

[0021] In a second aspect, the present invention discloses an intrusion detection system of a drone intrusion detection system, which is used in the drone intrusion detection method described in the first aspect, comprising:

[0022] The model initialization module is used to initialize the model data and measure the performance of the initial model on the training set;

[0023] The hyperparameter module is used to set hyperparameters during model training;

[0024] The training module is used to train the model, update the model parameters and calculate the validation loss value;

[0025] The server module is used to store the validation loss value in the parameter server, generate global model parameters, and traverse all local workers to identify Byzantine attack behaviors;

[0026] The update and verification module is used to update and verify the global model of the parameter server.

[0027] In a third aspect, the present invention discloses a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the instructions are executed on a processor, the processor executes any one of the methods provided in the first aspect.

[0028] In a fourth aspect, the present invention discloses a computer program product comprising instructions, which, when executed by a processor, enables the processor to execute any one of the methods provided in the first aspect.

[0029] The beneficial effects of the present invention are:

[0030] 1) In the field of UAV IDS, this invention is the first to regard physical and network data as cross-modal data, and innovatively proposes the idea of ​​sub-modal feature extraction, that is, using a specific feature extraction network for data of different modalities, and then using CrossAttention to fuse the extracted cross-modal features.

[0031] 2) Compared with the traditional UAV feature extraction network, the present invention establishes a new convolutional Transformer network for feature extraction of two types of heterogeneous data.

[0032] 3) In order to better handle nonlinear mapping, the present invention uses KAN Layer instead of the traditional MLP classifier.

[0033] 4) After comparing the model proposed in the present invention with seven comparative models in previous work, it is found that the present invention can significantly improve the detection performance. The F1 score of the model of the present invention can reach 99.57%, the Accuracy can reach 99.74%, the Precision can reach 99.32%, and the Recall can reach 99.82%. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 A schematic diagram of a process flow of a drone intrusion detection method according to an embodiment of the present invention;

[0035] Figure 2 A flowchart of a drone intrusion detection system according to an embodiment of the present invention;

[0036] Figure 3 A schematic diagram of the structure of a drone intrusion detection system according to an embodiment of the present invention;

[0037] Figure 4 It is a learning curve diagram of F1Score of the present invention and other seven comparison models under the physical network joint data set of an embodiment of the present invention;

[0038] Figure 5It is a learning curve diagram of Accuracy of the present invention and other seven comparison models under the physical network joint data set of an embodiment of the present invention;

[0039] Figure 6 It is a learning curve diagram of Precision of the present invention and other seven comparison models under the physical network joint data set of an embodiment of the present invention;

[0040] Figure 7 This is a learning curve diagram of Recall for the present invention and other seven comparison models under the physical network joint dataset of an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0042] The present invention discloses a method, system, storage medium and program product for detecting intrusion of unmanned aerial vehicles, which are used to solve the problem in the prior art that the physical characteristics and network characteristics of unmanned aerial vehicles are not used as cross-modal characteristics, nor are they fused to obtain fused characteristics, which will result in the inability to completely extract the characteristic representation of the information-physical system of the unmanned aerial vehicle group, thereby resulting in a decrease in the accuracy of intrusion detection. The method is mainly applied to unmanned aerial vehicle groups, and improves the detection accuracy and saves computing overhead while ensuring data security. Specifically, if there is inconsistency in the data received by the server, it will be discarded, which ensures the Byzantine robustness of the entire system process. At the same time, pyramid convolution and dilated convolution are used to save computing overhead, and the learning rate warm-up strategy is used to prevent overfitting, enhance its robustness, and improve detection accuracy. The present invention proposes an intrusion detection method based on the physical and network cross-modal data of the unmanned aerial vehicle group. The unmanned aerial vehicle system is a network-based measure that can realize communication and coordination between unmanned aerial vehicles. The system allows multiple unmanned aerial vehicles and unmanned aerial vehicles (UAVs) and ground control stations (GCSs) to exchange commands, data and information; the physical and network cross-modal data are two different data types containing different feature columns; the step flow diagram of the method is as follows Figure 1 As shown, the following steps are included:

[0043] S1. Model initialization: The model initializes the parameters of the network data branch (Cyber) and the physical data branch (Physical) to generate initial weights;

[0044] S2, hyperparameter setting, setting hyperparameters for training the model;

[0045] S3, distributed training, updating model parameters and calculating validation loss values, and sending the updated model parameters and calculated validation loss values ​​to the parameter server;

[0046] S4, aggregation and verification of parameter servers. The parameter server compares the received verification loss values ​​and selects the minimum verification loss value and its corresponding model parameter.

[0047] S5. Global model update and verification: If the models of all nodes pass the verification, the parameter server updates the global model to the best model of the current iteration; if no model passes the verification, the global model of the previous round remains unchanged. This strategy can effectively deal with potential Byzantine attacks while ensuring the stability and effectiveness of the global model.

[0048] Exemplarily, in step S1, the initial weight includes a first weight and the second weight Then calculate the initial loss weight through the initial weight Where D train represents a training data set, which is used to measure the performance of the initial model on the training set, and specifically includes the following steps: S11, the network data branch is the network branch data through the pyramid convolution layer PyramidConv1d, through the formula Y cyb,i =f(W i *X cyb +b i ) extracts network modal features, where W i represents the weight of the i-th convolution kernel that needs to be initialized. In this embodiment, the sizes of the i-th convolution kernels used in the network structure of this model are 3×1, 5×1, and 7×1, respectively, which are used to extract network modal features of different sizes; b i represents the constant corresponding to the i-th convolution kernel, f is the ReLU activation function, * represents the convolution operation, X cyb is the network modality data, Y cyb,i is the first network modal feature; pyramid convolution is an enhanced structure based on the traditional convolutional neural network (CNN), which convolves the input features through a variety of convolution kernel sizes to obtain feature representations of different scales; the physical data branch is the physical modal data that extracts the first physical modal feature through the dilated convolution layer DilatedConv1d, and the dilated convolution formula is Y phy =f(W *d X phy +b), where W is the convolution kernel weight that needs to be initialized, *d is the dilation convolution operation, d represents the dilation rate, and X phy is the physical modal data, Y phyis the first physical modal feature, b represents a constant; the dilated convolution increases the receptive field by inserting dilation without increasing the size of the convolution kernel.

[0049] S12, process the first network modality feature Y through the Transformer encoder cyb,i and the first physical mode characteristic Y phy , obtain the second network modality feature and the second physical modality feature, where the first network modality feature and the first physical modality feature are respectively applied with the self-attention mechanism, and the formula for calculating the attention weight is: Where Q is the Query matrix, K is the Key matrix, and V is the Value matrix, which are respectively input by embedding and the learnable weight matrix (W Q ,W K ,W V ) and multiply them to get, d k is the dimension of the key matrix used to scale the attention scores, and T is the transpose.

[0050] S13. Use Cross Attention Mechanism The extracted second network modality features and second physical modality features are integrated, where Q phy Indicates that the physical modal data X phy As the Query matrix, K cyber Indicates that the network modal data is used as the Key matrix, V cyber It means that the network modality data is used as the Value matrix. In this embodiment, the cross-modal correlation between the two modality features can be captured in the above manner.

[0051] S14. After obtaining the fusion features, the fusion features are mapped to the classification KAN Layer through two fully connected layers for output. Since this mapping is a nonlinear mapping, the classification effect of KAN Layer is theoretically better than that of the traditional LinearLayer. The core idea of ​​KAN is to replace the traditional linear mapping with the nonlinear mapping of the kernel function to improve the model's ability to capture complex features. Its calculation formula is: Among them, W base is the base linear layer weight, B i (x) is the B-spline basis function used to capture features in a finer granularity, W spline,i is the B-spline basis function B i (x), x represents the fusion feature, and y represents the output.

[0052] Specifically, step S2 specifically includes: setting hyperparameters during the training process, including the number of iterations num_epochs, batch size batch_size, learning rate learning_rate, hidden layer dimension hidden_dimension and early stopping patience value patience. These hyperparameters play a key role in the training efficiency and final performance of the model.

[0053] Specifically, step S3 specifically includes: in the tth (t∈[1,T'])th iteration, local worker i (i=1,…,U) uses the AdamW optimization algorithm Update the first model parameters Get the second model parameters Where T' represents the number of iterations, U represents the last local worker, Represents the calculation of the minimum loss function, represents the third weight, represents the fourth weight; then local worker i calculates the first validation loss value on the validation set Where D val Represents the validation data set; then the second model parameter and the first validation loss are sent to the parameter server (PS) for aggregation.

[0054] Specifically, step S4 specifically includes: the parameter server compares the received first verification loss value Save the minimum first validation loss value and its corresponding second model parameter to obtain the second validation loss value and the third model parameter To further enhance the robustness of the global model, the parameter server then calculates the third model parameters The mean S is the third model parameter The total number of samples forms the global model parameters During this process, the third model parameter is received After that, the parameter server also needs to use the validation data set D val Verify it, if The Byzantine attack behavior is identified, the third model parameter is filtered out, and the parameter server will query the next third model parameter until all local workers are traversed.

[0055] Exemplarily, the present invention is a new intrusion detection system (IDS) based on cross-modal feature extraction and feature fusion. The drone swarm is regarded as a system that integrates physical modal and network modal data. Different new convolutional Transformer feature extraction networks are established for cross-modal data. Then, CrossAttention is innovatively used to perform cross-modal feature fusion of drones. Finally, KAN is used instead of traditional MLP to perform fusion vector embedding mapping. We name this intrusion detection model TransCAIDS. The design concept of the present invention is to make full use of the complementarity between network data and physical data. CrossAttention establishes an interconnection between network data and physical data, so that the model can be more efficient and accurate in fusing information from different data sources. At the same time, the new convolutional Transformer network can perform multi-level representation of heterogeneous data separately. The self-attention mechanism enables the data at each time step to interact with other time steps, thereby capturing the global dependencies in the data. Therefore, compared with the other seven comparison models, TransCAIDS has better detection performance, with an F1 score of 99.57%, an Accuracy of 99.74%, a Precision of 99.32%, and a Recall of 99.82%.

[0056] The present invention also discloses a drone intrusion detection system, which is used for the drone intrusion detection method described above. The working flow diagram thereof is as follows: Figure 2 The specific structural diagram is shown in Figure 3 Shown; including:

[0057] The model initialization module is used to initialize the model data and measure the performance of the initial model on the training set;

[0058] The hyperparameter module is used to set hyperparameters during model training;

[0059] The training module is used to train the model, update the model parameters and calculate the validation loss value;

[0060] The server module is used to store the validation loss value in the parameter server, generate global model parameters, and traverse all local workers to identify Byzantine attack behaviors;

[0061] The update and verification module is used to update and verify the global model of the parameter server.

[0062] For example, the intrusion detection method for drone physical network information system based on the method of the present invention (hereinafter referred to as TransCAIDS) is compared with the existing drone intrusion detection method. The learning curves of the present invention and other seven comparison models for four evaluation indicators under the physical network joint dataset are shown in the figure below: Figure 4-Figure 7 As shown, Figure 4 This is a schematic diagram of the F1Score comparison results. Figure 5 This is a schematic diagram of the Accuracy comparison results. Figure 6 This is a schematic diagram of the Precision comparison results. Figure 7 The figure is a schematic diagram of the recall comparison results, in which the seven compared models are: CNN is an intrusion detection model based on the physical network joint dataset, using the classic CNN model; LSTM is an intrusion detection model based on the physical network joint dataset, using the classic LSTM model; CNN-LSTM is a model that uses convolutional layers CNN and LSTM layers to extract features from training data; CNN-LSTM-Res adds residual connections to the CNN-LSTM model to enhance feature extraction; Transformer is an intrusion detection model based on the physical network joint dataset, using the classic Transformer model; TransCAIDS lstm TransCAIDS replaces the Transformer encoder with LSTM, and the fully connected layer does not use KAN, but the traditional Linear Layer; TransCAIDS linear Based on TransCAIDS, the traditional Linear Layer is used as the fully connected layer. This embodiment will verify the effectiveness of the method TransCAIDS of the present invention based on the physical network joint dataset Dataset-ITS; the verification index used in this embodiment is: Accuracy represents the percentage of samples that describe the correct classification, and its calculation formula is: Among them, TP represents the number of true positives, TN represents the number of true negatives, FP represents the number of false positives, and FN represents the number of false negatives. Precision refers to the proportion of correctly identified positive samples to the total number of predicted positive samples, and its calculation formula is: Recall shows the proportion of positive samples accurately identified among all actual positive samples, expressed as: F1Score coordinates the precision and recall indicators and provides an effective analysis indicator. It is calculated as the harmonic mean of precision and recall: from Figures 4 to 7It can be seen that the classification performance of the present invention on the four evaluation indicators is better than that of the other seven models. Models 1-5 all treat Dataset-ITS as single-modal data and do not perform cross-modal feature extraction. It can be found that Model 7 and TransCAIDS are both cross-modal feature fusion models, and their performance evaluation indicators are significantly higher than those of Models 1-5. Among them, the four evaluation indicators of TransCAIDS are all above 99%, which further proves the superiority of using Cross Attention for cross-modal feature fusion from the simulation.

[0063] Exemplarily, the present invention further provides a computer program product, comprising computer executable instructions. In one embodiment, the computer executable instructions are used to enable a computer to execute Figure 1 The computer executable instructions may be stored in a computer readable storage medium. The present invention also provides a computer readable storage medium, wherein the computer readable storage medium stores the executable instructions. In one embodiment, the computer executable instructions are used to enable a computer to execute Figure 1 Functionality in the method embodiment shown.

[0064] Exemplarily, the computer-readable storage medium provided by the present invention can be a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a mobile hard disk, a CD-ROM, or any other form of computer-readable storage medium known in the art.

[0065] Computer executable instructions may be stored in a computer readable storage medium or transferred from one computer readable storage medium to another computer readable storage medium. For example, the computer program or instructions may be transferred from one website, computer, server or data center to another website, computer, server or data center by wired or wireless means. The computer readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it may also be an optical medium, such as a digital video disc (DVD); it may also be a semiconductor medium, such as a solid state drive.

[0066] The above is only a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein, and should not be regarded as excluding other embodiments, but can be used in various other combinations, modifications and environments, and can be modified within the scope of the concept described herein through the above teachings or the technology or knowledge of the relevant field. The changes and modifications made by those skilled in the art shall not deviate from the spirit and scope of the present invention, and shall be within the scope of protection of the claims attached to the present invention.

Claims

1. A method for detecting intrusion of a drone, characterized in that: The following steps are involved: S1, model initialization, the model initializes the parameters of the network data branch and the physical data branch to generate initial weights; S2, hyperparameter setting, setting hyperparameters for training the model; S3, distributed training, updating model parameters and calculating validation loss values, and sending the updated model parameters and calculated validation loss values ​​to the parameter server; S4, aggregation and verification of parameter servers. The parameter server compares the received verification loss values ​​and selects the minimum verification loss value and its corresponding model parameter. S5, global model update and verification, if the models of all nodes pass the verification, the parameter server updates the global model to the model of the current iteration; if no model passes the verification, the global model of the previous round remains unchanged.

2. A drone intrusion detection method according to claim 1, characterized in that: In step S1, the initial weight includes a first weight and the second weight Then calculate the initial loss weight through the initial weight Where D train Represents the training data set, which includes the following steps: S11, the network data branch is the network branch data through the pyramid convolution layer PyramidConv1d, through the formula Y cyb,i =f(W i *X cyb +b i ) extracts network modal features, where W i represents the weight of the i-th convolution kernel that needs to be initialized, b i represents the constant corresponding to the i-th convolution kernel, f is the ReLU activation function, * represents the convolution operation, X cyb is the network modality data, Y cyb,i is the first network modality feature; The physical data branch is the physical modal data. The first physical modal feature is extracted through the dilated convolution layer DilatedConv1d. The formula of the dilated convolution is Y phy =f(W *d X phy +b), where W is the convolution kernel weight that needs to be initialized, *d is the dilation convolution operation, d represents the dilation rate, and X phy is the physical modal data, Y phy is the first physical mode characteristic, b represents a constant; S12, process the first network modality feature Y through the Transformer encoder cyb,i and the first physical mode characteristic Y phy , obtain the second network modality feature and the second physical modality feature, where the first network modality feature and the first physical modality feature are respectively applied with the self-attention mechanism, and the formula for calculating the attention weight is: Where Q is the Query matrix, K is the Key matrix, and V is the Value matrix, which are respectively input by embedding and the learnable weight matrix (W Q ,W K ,W V ) and multiply them to get, d k is the dimension of the key matrix used to scale the attention scores, T is the transpose; S13. Use Cross Attention Mechanism The extracted second network modality features and second physical modality features are integrated, where Q phy Indicates that the physical modal data X phy As the Query matrix, K cyber Indicates that the network modal data is used as the Key matrix, V cyber Indicates that the network modality data is used as the Value matrix; S14, through two layers of fully connected layers, the fusion features are mapped to the classification KAN Layer for output, and the calculation formula is: Among them, W base is the base linear layer weight, B i (x) is the B-spline basis function used to capture features in a finer granularity, W spline,i is the B-spline basis function B i (x), x represents the fusion feature, and y represents the output.

3. A drone intrusion detection method according to claim 2, characterized in that: The step S2 specifically includes: setting hyper parameters during the training process, including the number of iterations num_epochs, the batch size batch_size, the learning rate learning_rate, the hidden layer dimension hidden_dimension and the patience value patience of early stopping.

4. A drone intrusion detection method according to claim 3, characterized in that: The step S3 specifically includes: in the tth (t∈[1,T'])th iteration, the local worker i (i=1,...,U) uses the AdamW optimization algorithm Update the first model parameters Get the second model parameters Where T' represents the number of iterations, U represents the last local worker, Represents the calculation of the minimum loss function, represents the third weight, represents the fourth weight; then local worker i calculates the first validation loss value on the validation set Where D val Represents the validation data set; then the second model parameter and the first validation loss are sent to the parameter servers for aggregation.

5. A drone intrusion detection method according to claim 4, characterized in that: The step S4 specifically includes: the parameter server compares the received first verification loss value Save the minimum first validation loss value and its corresponding second model parameter to obtain the second validation loss value and the third model parameter Then the parameter server calculates the third model parameters The mean S is the third model parameter The total number of samples forms the global model parameters During this process, the third model parameter is received After that, the parameter server also needs to use the validation data set D val Verify it, if It is identified as a Byzantine attack behavior, and the third model parameter is filtered out. The parameter server will query the next third model parameter until all local workers are traversed.

6. A drone intrusion detection system, used in a drone intrusion detection method according to any one of claims 1 to 5, characterized in that: include: The model initialization module is used to initialize the model data and measure the performance of the initial model on the training set; The hyperparameter module is used to set hyperparameters during model training; The training module is used to train the model, update the model parameters and calculate the validation loss value; The server module is used to store the validation loss value in the parameter server, generate global model parameters, and traverse all local workers to identify Byzantine attack behaviors; The update and verification module is used to update and verify the global model of the parameter server.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed on a processor, the processor is caused to execute the method according to any one of claims 1 to 5.

8. A computer program product comprising instructions, characterized in that When the instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 5.