A defense method for deep learning signal individual recognition model based on multimodality
By utilizing the multimodal information of individual radio frequency signals, especially the fusion of time-domain and frequency-domain data, the problem of vulnerability of the signal individual identification model is solved, and higher classification accuracy and defense capabilities are achieved.
Patent Information
- Application Number
- CN202210835292.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-15
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2042-07-15
AI Technical Summary
Existing deep learning-based signal individual recognition models are vulnerable to adversarial evasion attacks, and there are few studies on attacks that cannot be migrated between signal domains, resulting in insufficient model defense performance.
Using multimodal information of individual radio frequency signals, especially time domain and frequency domain data, through multimodal data fusion and model structure improvement, adversarial samples are generated and the defense performance of the model is tested, thereby improving the classification accuracy and defense ability of individual signal recognition.
While improving the accuracy of signal individual recognition, it greatly reduces the attack success rate of traditional attack methods on the model and enhances the defense performance of the model.
Smart Images

Figure CN115392285B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a convolutional neural network (CNN), multimodal fusion of electromagnetic signals, individual identification of electromagnetic signals, and a defense method for electromagnetic signal classification models, and in particular to a method for improving the defense performance of a signal individual identification model by utilizing multimodal information of different signal domains of electromagnetic signals. Background Art
[0002] Radio signal classification has a wide range of applications in wireless communications and electromagnetic spectrum management. In recent years, with the rapid development of deep learning technology, deep learning has been widely used in image recognition, speech recognition, natural language recognition, and wireless communications. In the past few years, deep learning has also been used to solve radio signal classification problems (Reference [1]: T.J.O.Shea, T.Roy, and T.C.Clancy, “Over-the-air deep learning based radiosignal classification,” IEEE J.Sel.Topics Signal Process., vol.12, no.1, pp.168-179, Feb.2018, i.e. T.J.O.Shea, T.Roy, and T.C.Clancy, “Over-the-air deep learning based radiosignal classification,” IEEE J.Sel.Topics Signal Process., vol.12, no.1, pp.168-179, Feb.2018.).
[0003] Research on radio signal classification based on deep learning mainly focuses on two aspects: automatic modulation recognition and radio frequency identification. Radio frequency identification is to achieve individual identification by transmitting and receiving radio frequency signals to objects of different categories. In order to achieve better practical application effects of individual identification, researchers around the world have conducted a lot of research. Some researchers used CNN to perform fingerprint identification on 5 ZigBee devices (Reference [2]: K. Merchant, S. Revay, G. Stantchev, and B. Nousian, "Deep learning for RF device fingerprinting incognitive communication networks," IEEE J. Sel. Topics Signal Process., vol. 12, no. 1, pp. 160-167, Feb. 2018, ... Process., vol.12, no.1, pp.160-167, Feb.2018.), some researchers used convolutional neural networks (CNN) and commercial real-world Aircraft Communication Address and Reporting System (ACARS) signal data to identify aircraft (Reference [3]: S.Zheng, S.Chen, L.Yang, J.Zhu, Z.Luo, J.Hu, and X.Yang, "Big data processing architecture for radio signals empowered by deep learning: Concept, experiment, application and challenges," IEEE Access, vol.6, pp.55907-55922, 2018, i.e. S.Zheng, S.Chen, L.Yang, J.Zhu, Z.Luo, J.Hu, and X.Yang, "Big data processing architecture for radio signals empowered by deep learning: Concept, experiment, application and challenges," IEEE Access, vol.6, pp.55907-55922, 2018.)
[0004] Although deep learning models can achieve strong individual recognition performance, they are highly vulnerable to adversarial evasion attacks (Reference [4]: M. Sadeghi and E.G. Larsson, “Adversarial attacks on deep-learning based radio signal classification,” IEEE Wireless Commun. Lett., vol.8, no.1, pp.213-216, Feb. 2019, i.e. M. Sadeghi and E.G. Larsson, “Adversarial attacks on deep-learning based radio signal classification,” IEEE Wireless Commun. Lett., vol.8, no.1, pp.213-216, Feb. 2019.), which introduce additive wireless interference into the RF signal, thereby inducing erroneous behavior on well-trained individual recognition models. Recently, studies have shown that (reference [5]: R.Sahay, C.G.Brinton and DJLove, “A Deep Ensemble-Based Wireless Receiver Architecture for Mitigating Adversarial Attacks in Automatic Modulation Classification,” IEEE Transactions and Cognitive Communication and Network, vol.8, no.1.Mar.2022.), for attacks in the signal time domain, the attack effect will be greatly reduced in the process of converting time domain information into frequency domain information, that is, the attack on the signal is not transferable between signal domains. This guides us to use multimodal information from different signal domains to enhance the defense performance of individual recognition and classification models.
[0005] Multimodal information is a set of information that maps real individuals or phenomena. Since different modalities contain rich and complementary information, it tends to replace traditional single-modal information. In recent years, the successful application of multimodal methods in the field of biomedical imaging has brought unique value to medical applications. Pattern recognition based on multimodal information has also begun to be applied to the signal field (reference [6]: Peihan Qi, Xiaoyu Zhou, Shilian Zheng, and Zan Li, “Automatic Modulation Classification Based on Deep Residual Networks With Multimodal Information,” IEEE Transactions and Cognitive Communication and Network, Vol. 7, no. 1, Mar. 2021, i.e. Peihan Qi, Xiaoyu Zhou, Shilian Zheng, and Zan Li, “Automatic Modulation Classification Based on Deep Residual Networks With Multimodal Information,” IEEE Transactions and Cognitive Communication and Network, Vol. 7, no. 1, Ma. 2021.). However, there are few studies on using multiple signal modality data to train deep neural networks to improve the model's defense capabilities.
[0006] Therefore, in order to improve the defense capability of the signal individual recognition domain model and enable it to be better used in edge environments with limited hardware resources, this project makes full use of the multimodal information of the signal, combines the experimental research results that the attack on the signal cannot be transferred between signal domains, and proposes a multimodal deep learning signal individual recognition model defense method. While improving the accuracy of signal individual recognition, it greatly reduces the success rate of traditional attack methods on the model and improves the defense performance of the model. Summary of the Invention
[0007] In order to overcome the current situation that existing signal individual recognition research based on deep learning rarely utilizes the multimodal information of signals to improve the defense performance of classification models, the present invention provides a method that utilizes the multimodal information of radio frequency signals, gives full play to the research results that attacks are not transferable between signal domains, and through certain multimodal data fusion methods and model structure improvements, while improving the individual recognition accuracy of the model, greatly improving the defense performance of the model.
[0008] The technical solutions adopted by the present invention to solve the technical problems are as follows:
[0009] A multimodal deep learning signal individual recognition model defense method includes the following steps:
[0010] S1: The individual RF signal dataset is divided into a training set and a test set according to a certain ratio. The training set is used to train the model, and the test set is used to test the individual recognition accuracy of the trained model. The original signal is used to generate two sets of multimodal data: the real part and imaginary part of the signal and the short-time Fourier transform. The former is the mode of the signal in the time domain, and the latter is the mode of the signal in the frequency domain. The two belong to different signal domains.
[0011] S2: The two sets of modal data are normalized and input into the built convolutional neural network model suitable for multimodal input for training. The model training is completed and the individual recognition and classification accuracy of the model for the test set is obtained;
[0012] S3: Attack the original signal to generate adversarial samples. Use these adversarial samples to generate two sets of multimodal adversarial data: the real and imaginary parts, and the short-time Fourier transform. Using this multimodal adversarial data and the trained model, test the attack success rate and evaluate the defense performance of the individual recognition model for multimodal signal input.
[0013] Furthermore, the step S1 includes the following contents:
[0014] S1.1: Generate the real and imaginary parts of the individual RF signal, which is the modal data in the time domain. Extract the real and imaginary parts of the complex value of each sampling point of the signal, where the real part of the complex value is defined as I n (n=0,1,2,...,N-1), the imaginary part of the complex value is defined as Q n (n=0,1,2,...,N-1), N is the number of sampling points of each signal data;
[0015] S1.2: Generate the short-time Fourier transform of the individual RF signal, which is the frequency domain modal data. Calculate the short-time Fourier transform of the individual RF signal, which is:
[0016]
[0017] Furthermore, step S2 includes the following contents:
[0018] S2.1: To reduce the impact of the magnitude difference of different modal data in the same group on model training, each group of multimodal data is normalized separately. The normalization method used is:
[0019] x=2×(x-min) / (max-min)-1 (2)
[0020] Where max is the maximum value in the same modality data, and min is the minimum value in the same modality data. After normalization, the values of all data used for model training are within the range of -1 to 1;
[0021] S2.2: The normalized multimodal data of the individual RF signals is input into an individual identification and classification model suitable for multimodal input. Specifically, the model consists of two parallel convolutional neural networks with identical structures. The two sets of modal data are input into the two parallel convolutional neural networks respectively. The two logit soft label probability matrices output by the convolutional neural network after feature extraction are fused by voting by adding corresponding positions to obtain the model prediction results, and finally the individual identification and classification accuracy of the trained model for the test set is obtained.
[0022] Furthermore, step S3 includes the following contents:
[0023] S3.1: Use a traditional iterative attack method to attack the original signal to obtain an adversarial sample of the original signal. Using the original signal adversarial sample, organize and calculate the real and imaginary parts of the signal in the time domain and the short-time Fourier transform of the signal in the frequency domain;
[0024] S3.2: Input adversarial examples from the time and frequency domains of the signal into the trained multimodal input convolutional neural network model to obtain predicted labels. Compare the predicted labels with the true labels to determine the attack success rate against the model. This is used to evaluate the defense performance of the multimodal input individual recognition model. The attack success rate is the ratio of the number of samples where the predicted label differs from the true label to the total number of samples.
[0025] The working principle of the present invention is: utilizing the multimodal data of individual RF signals, especially the complementary information of data in two different signal domains, time domain and frequency domain, richer signal characteristics, and experimental research results that signal attacks cannot be transferred between signal domains. While improving the classification accuracy of signal individual recognition, it greatly improves the defense performance of the multimodal input individual recognition model.
[0026] The advantages of the present invention are: by utilizing the multimodal information of individual radio frequency signals and the experimental research results that the attacks on the signals cannot be transferred between signal domains, the classification accuracy of the deep learning individual recognition model with multimodal input and the defense performance of the model are effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A block diagram of the method of the present invention.
[0028] Figure 2 A framework diagram of a convolutional neural network with multimodal input for the method of the present invention. DETAILED DESCRIPTION
[0029] The following is attached with the instruction manual Figure 1 The present invention is further described in detail by taking an aircraft dataset containing radio frequency signals emitted by individual aircraft of 14 categories as an example.
[0030] The defense method of the present invention based on a multimodal deep learning aircraft radio frequency signal individual recognition model is specifically divided into the following steps:
[0031] Step 1: Use the original data set to divide it into training set and test set, and generate signal multimodal data;
[0032] Step 2: Normalize the multimodal data and input it into the multimodal input deep neural network model we built for training to obtain the individual recognition accuracy of the model;
[0033] Step 3: Using the original time-domain data and the time-domain individual recognition model from the trained model, attack the time-domain data using traditional iterative attack methods. Time-domain adversarial samples are then used to generate two sets of multimodal adversarial sample data, including the real and imaginary parts and short-time Fourier transforms. Using the multimodal adversarial sample data and the trained multimodal input individual recognition model, the attack success rate of the model is tested to evaluate the defense performance of the signal multimodal input individual recognition model.
[0034] Step 4: Apply the defense method based on the multimodal deep learning aircraft RF signal individual identification model to real-world scenarios.
[0035] In step 1, the specific operation process is as follows: Figure 1 First, the aircraft individual signal dataset is divided into a training set and a test set with a 4:1 ratio. The training set is used to train the deep neural network individual recognition model for multimodal input, while the test set is used to test the model and evaluate its classification performance. Specifically, each sample in the aircraft dataset contains 15,000 samples, each of which includes a real (I) value and an imaginary (Q) value, meaning that each sample has a data shape of 1×15,000×2. Using the training and test sets, two sets of multimodal data are generated, representing the real and imaginary parts of the signal and the short-time Fourier transform (SFT), which can represent the original signal information. The time domain data of the real and imaginary parts of each sample signal have a data shape of 1×15,000×2. For the SFT, we use a window size of 15,000, the same as the length of the original signal, and a Hanning window type, so the frequency domain data of the SFT of each sample signal has a data shape of 1×15,000×3.
[0036] In step 2, the specific operation process is as follows: Figure 1and attached Figure 2 The calculated time and frequency domain data of the individual RF signals are normalized to achieve better model training results. Specifically, we use the normalization method x = 2 × (x-min) / (max-min) - 1, where max is the maximum value in each sample data and min is the minimum value in each sample data. After normalization, the values of all data involved in model training are between -1 and 1. The normalized time and frequency domain modal data of the training and test sets are used as inputs to two identical, parallel convolutional neural networks for our multimodal input individual recognition and classification network, respectively. This results in logit soft label probability matrices after feature extraction by the two convolutional neural networks. Specifically, because our aircraft individual dataset contains 14 different aircraft individual classes, the resulting logit soft label probability matrix for each path is 14 × 1, with the value at each position representing the probability of the sample belonging to the corresponding class. The two logit soft label probability matrices belonging to the time domain and frequency domain are fused by adding the corresponding positions, and the prediction results of each sample are obtained by using them, and finally the individual recognition and classification accuracy of the model for all test sets is obtained.
[0037] In step 3, the specific operation process is as follows: Figure 1 and attached Figure 2 Using the original time-domain data and the time-domain individual recognition model from the trained model, we attack the time-domain data using PGD's traditional iterative attack method. Because our model performs feature fusion after the fully connected layer, the time-domain and frequency-domain models are essentially trained separately before voting, resulting in two classification models (one in the time domain and one in the frequency domain). Here, we attack the trained time-domain individual recognition model to generate time-domain adversarial examples, with the perturbation size epsilon set to 0.01. Next, we use the time-domain adversarial examples to generate adversarial examples in the frequency domain of the short-time Fourier transform (SFT) signal. These adversarial examples in the time and frequency domains are input into the trained multimodal input convolutional neural network model to obtain predicted labels. The predicted labels are compared with the true labels to determine the attack success rate, which is used to evaluate the defense performance of the multimodal input individual recognition model. The attack success rate is the ratio of the number of samples whose predicted labels differ from the true labels to the total number of samples.
[0038] In step 3, the specific operation process is as follows: the defense method based on the multimodal deep learning aircraft radio frequency signal individual identification model is applied to real-world scenarios. The airport tower converts the received aircraft signal into time-frequency, uses the signal's time-frequency domain data to individually identify the aircraft signal's source, and determines the source's individual number. Using this method for aircraft individual identification greatly reduces the probability of prediction failure due to interference and attacks during signal transmission, greatly improving the defense performance of the aircraft individual identification system. In the civil aviation field, it strengthens the ground's real-time monitoring capabilities of in-flight flights, improving aviation safety, including flight safety, flight ground safety, and air defense safety. In military aviation, it can more accurately monitor and determine the entry of foreign aircraft into the airspace.
[0039] The above is an introduction to an embodiment of the present invention for enhancing the defensive performance of a deep neural network individual signal recognition model using multimodal information from two signal domains, time and frequency, using the original data set to generate multimodal information in two different signal domains, time and frequency, and using the multimodal information to train the multimodal input convolutional neural network we built. Feature fusion is performed by adding the corresponding positions of the logit soft label probability matrix. The multimodal information of different signal domains of individual RF signals is used to give full play to the complementary advantages between the modalities, as well as the experimental research results that the attack of the signal cannot be transferred between signal domains. While improving the classification accuracy of individual signal recognition based on deep learning, the defensive performance of the classification model is greatly improved. This is merely illustrative and not restrictive of the invention. Those skilled in the art understand that many changes, modifications, and even equivalents can be made to it within the spirit and scope defined by the claims, but they will all fall within the scope of protection of the present invention.
Claims
1. A multimodal deep learning signal individual recognition model defense method, characterized by: The steps include: S1: The individual RF signal dataset is divided into a training set and a test set according to a certain ratio. The training set is used to train the model, and the test set is used to test the individual recognition accuracy of the trained model. The original signal is used to generate two sets of multimodal data: the real part and imaginary part of the signal and the short-time Fourier transform. The former is the mode of the signal in the time domain, and the latter is the mode of the signal in the frequency domain. The two belong to different signal domains. S2: Normalize the two sets of modal data generated and input them into the built convolutional neural network model suitable for multimodal input for training. Complete the model training and obtain the individual recognition and classification accuracy of the model for the test set. Specifically, it includes: S2.1: To reduce the impact of the magnitude difference of different modal data in the same group on model training, each group of multimodal data is normalized separately. The normalization method used is: x=2×(x-min) / (max-min)-1 (2) Where max is the maximum value in the same modality data, and min is the minimum value in the same modality data; after normalization, the values of all data used for model training are within the range of -1 to 1; S2.2: Input the normalized multimodal data of the individual RF signal into an individual identification and classification model suitable for multimodal input. Specifically, the model consists of two parallel convolutional neural networks with identical structures. The two sets of modal data are input into the two parallel convolutional neural networks respectively. The two logit soft label probability matrices output after feature extraction by the convolutional neural networks are fused by voting by adding corresponding positions to obtain the model prediction results. Ultimately, the individual identification and classification accuracy of the trained model for the test set is obtained. S3: Perturb the original signal to generate adversarial samples, and use the adversarial samples to generate two sets of multimodal adversarial sample data: real and imaginary parts and short-time Fourier transforms. Using the multimodal adversarial sample data and the trained model, test the perturbation success rate of the model and evaluate the defense performance of the signal multimodal input individual recognition model. Specifically, this includes: S3.1: Use a traditional iterative perturbation method to perturb the original signal to obtain an adversarial sample of the original signal. Using the adversarial sample of the original signal, organize and calculate the time-domain adversarial samples of the real and imaginary parts of the signal and the frequency-domain adversarial samples of the short-time Fourier transform of the signal. S3.2: Input adversarial samples of the signal in the time and frequency domains into the trained multimodal input convolutional neural network model to obtain predicted labels. Compare the predicted labels with the true labels to obtain the perturbation success rate of the model to evaluate the defense performance of the signal multimodal input individual recognition model; where the perturbation success rate is the ratio of the number of samples whose predicted labels are different from the true labels to the total number of samples.
2. The multimodal deep learning signal individual recognition model defense method according to claim 1 is characterized by: The step S1 includes the following contents: S1.1: Generate the real and imaginary parts of the individual RF signal, which are modal data in the time domain; extract the real and imaginary parts of the complex value of each sampling point of the signal, where the real part of the complex value is defined as I n , n=0,1,2,...,N-1, the imaginary part of the complex value is defined as Q n , n=0,1,2,...,N-1, N is the number of sampling points of each signal data; S1.2: Generate the short-time Fourier transform of the individual RF signal, which is frequency domain modal data; calculate the short-time Fourier transform of the individual RF signal from: