Multi-modal signal identification method based on additive attention mechanism
Through the additive attention mechanism, the three modal signals of IQ, FFT and AP are integrated, which solves the problem of insufficient accuracy and robustness of the traditional single modal signal recognition method in complex environments, and achieves higher recognition accuracy and stability under noise interference.
Patent Information
- Application Number
- CN202510379148.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-08-15
AI Technical Summary
The traditional single-modal signal recognition method lacks recognition accuracy and robustness in complex environments, and cannot effectively capture key information in complex signals.
A multimodal signal recognition method based on the additive attention mechanism is adopted, and the three modal signals of IQ, FFT and AP are fused, and the characteristic attention is dynamically adjusted using the additive attention mechanism, and the sequence timing characteristics are extracted in combination with the LSTM layer and the attention mechanism.
It improves the accuracy and robustness of signal recognition, enhances the system's performance under noise and interference conditions, and adapts to signal transformation in complex environments.
Smart Images

Figure CN120492996A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a signal recognition method, and more specifically, to a multimodal signal recognition method based on an additive attention mechanism, which belongs to the cross-application field of artificial intelligence and communications. Background Art
[0002] Signal modulation type identification is a key technology in communication systems. It primarily involves extracting and classifying information from received signals. The background of signal modulation identification technology stems primarily from the development of communication systems. With advances in wireless communications, satellite communications, and radar technology, the efficient transmission and reception of information has become a key issue. Modulation is the process of embedding information signals into a carrier signal by modifying certain characteristics of the carrier (such as amplitude, frequency, and phase) to make it resistant to interference and attenuation during transmission.
[0003] Using deep learning for signal modulation recognition is an important research area in the communications field. Traditional modulation recognition methods rely primarily on hand-crafted feature extraction and classic signal processing techniques such as Fourier transforms, time-domain analysis, and statistical property analysis. These methods often require the expertise of domain experts and are susceptible to environmental changes and interference when processing complex signals, resulting in reduced recognition performance. Deep learning methods break away from this reliance on traditional manual extraction, automatically extracting important features from the raw signal. By training models with large amounts of labeled data, deep learning can automatically learn effective feature representations, capturing subtle changes and complex patterns in the signal. This enables the model to handle nonlinear relationships and is suitable for complex modulated signals. Furthermore, it can adapt to varying channel conditions and environmental variations, demonstrating superior generalization capabilities.
[0004] In modern communications and signal processing, most signal recognition techniques utilize IQ signals to efficiently process and analyze modulated signals. However, IQ signals only provide amplitude and phase information and cannot fully reflect the signal's frequency domain characteristics and energy distribution, thus limiting the integrity of the information. Furthermore, single IQ signals are less robust to noise and interference, which can easily lead to recognition errors. Fusion of multiple vectors can enhance system stability. Multimodal signal recognition technology is gaining increasing attention. Accurately identifying and classifying signal types is crucial to system performance, particularly in applications such as wireless communications, radar systems, and electronic reconnaissance. Traditional single-modal signal recognition methods often fail to fully capture key information in complex signals, resulting in reduced recognition accuracy and robustness. Summary of the Invention
[0005] In order to solve the above-mentioned problems in the prior art, the present invention provides a multimodal signal recognition method based on the additive attention mechanism, which has the technical characteristics of being able to enhance key features to better adapt to signal transformations in complex environments, thereby not only improving the accuracy of recognition, but also enhancing the robustness of the system under noise and interference conditions.
[0006] In order to achieve the above object, the present invention is implemented through the following technical solutions:
[0007] The present invention provides a multimodal signal recognition method based on an additive attention mechanism, the method comprising the following steps:
[0008] Step 1: Data input, preprocessing of the data set;
[0009] Select the RML2016.10A dataset and perform preprocessing on it, converting it into vectors of three modes: AP, FFT, and IQ.
[0010] The conversion formula is as follows:
[0011]
[0012] Where A(t) represents the instantaneous amplitude of the IQ signal. The instantaneous amplitude (envelope) of the signal is calculated by taking the square root of the I and Q paths. I(t) and Q(t) represent the real and imaginary parts of the signal in the time domain, respectively.
[0013]
[0014] Where P(t) represents the instantaneous phase of the signal, which is obtained by taking the inverse tangent of the ratio of the I and Q paths.
[0015] X[k]=Re(X[k])+j·Im(X[k]);
[0016] Where X[k] is the frequency domain representation of the frequency domain signal, which is the frequency domain complex signal obtained after fast Fourier transform, R[·] and I[·] represent the real and imaginary parts of the signal respectively;
[0017]
[0018] Where Re(X[k]) is the real number representation of FFT, which is obtained by summing the products of the time domain signal and the cosine basis function x[n]; N is the number of FFT points, n is the time domain index, and k is the frequency domain index;
[0019]
[0020] Where Im(X[k]) is the real number representation of FFT, which is obtained by summing the product of the time domain signal and the sine basis function x[n] and taking the negative; N is the number of FFT points, n is the time domain index, and k is the frequency domain index;
[0021] Step 2: Design a multi-channel network structure;
[0022] Based on the three modes, the IQ is split into IQ channels, I channels, and Q channels; the AP is split into A channels, P channels, and AP channels. The FFT vector is divided into real and imaginary components as input. The shape of each single channel is 128×1, and then a one-dimensional convolution operation is performed on each single channel.
[0023] Step 3: multimodal fusion processing;
[0024] After convolution extraction of the three signal variables, IQ vector, AP vector, and FFT vector, the extracted features are fused in the first dimension to facilitate subsequent feature extraction. A convolution kernel with 100 channels and 2*5 kernel is used for feature extraction. The final data format is 3*124*100, which is then reshaped into a 372*100 data format for subsequent operations.
[0025] Step 4: Add LSTM layer and attention mechanism to extract sequence timing features;
[0026] After the data is fused by multiple channels and modules, it is passed through a bidirectional LSTM layer and an attention mechanism module to further extract features of the data and enhance the signal representation capability.
[0027] Step 5: Output classification;
[0028] The Flatten layer and the Dense layer convert the multidimensional features processed by the attention mechanism into a format suitable for classification. The Flatten layer is responsible for flattening the data, and the Dense layer generates the final output result through a fully connected operation.
[0029] Step 6: Verify the results and compare them;
[0030] In the framework of tensorflow, the training set and validation set are divided into 8:2; then the multimodal signal recognition (model) training is performed on the first 80% of each type of signal-to-noise ratio data in the dataset, and the multimodal signal recognition is verified on the last 20%.
[0031] Preferably, the specific operation in step 4 is: the above-mentioned fused data result is sent to the bidirectional LSTM layer after a one-dimensional convolution, and after processing the features using a fully connected layer, a weight distribution for each time step is generated, and then spliced with the original data. Through splicing, the multimodal signal recognition (model) can simultaneously utilize the original information and the important features adjusted according to the attention mechanism.
[0032] Preferably, step 5: the multimodal signal recognition (model) training optimizer adopts the Adam algorithm, the epoch is set to 100, the batch size is set to 64, and the accuracy index and the confusion matrix visualization image are used to evaluate and judge each signal classification result under different signal-to-noise ratios to achieve the classification accuracy of the evaluation signal.
[0033] Preferably, in step 2: the I and Q split channels are convolved with a convolution kernel of 8 and a channel number of 50, and then the convolution results are merged into two channels for a two-dimensional convolution operation with a channel number of 50 and a convolution kernel of 1*8. Then, the main channel data IQ vector is convolved with a 2*8 convolution kernel and a channel number of 50. Finally, feature fusion is performed with the features extracted by the above single channel, and feature extraction is performed using data with a channel number of 50 and a convolution kernel of 2*8, thereby capturing representation features at different scales; the data under the other two modes are extracted in the same way as under the IQ mode.
[0034] Preferably, the dataset RadioML 2016.10A includes 11 modulation signal types, including eight digital signals and three analog signals, with a signal-to-noise ratio range of -18 to 20 dB, a sampling rate of 200 kHz, and Gaussian white noise.
[0035] RadioML 2016.10A covers variables such as different signal-to-noise ratios and timestamps. The total number of samples is 220,000, the signal-to-noise ratio range is -18 to 20dB, the sampling rate is 200kHz, and the noise is Gaussian white noise. These features have important reference value for deep learning recognition and classification tasks in signal processing. Therefore, this application selected the open source dataset RML2016.10A as the verification dataset. Signal preprocessing is to meet the needs of analysis and processing, and the original signal is subjected to a series of transformations and modifications to improve the quality and reliability of the signal. Most signal modulation methods process data information of a single mode and ignore the correlation between the multi-modalities of the signal. This application studies the representation methods of different modes of the signal, namely amplitude, phase and frequency domain features, extracts features from multi-modalities to form interactions, and improves the recognition performance of the model.
[0036] Beneficial effects: The present invention strengthens key features by dynamically adjusting the degree of attention paid to different modal features, thereby better adapting to signal transformations in complex environments; it not only improves the accuracy of recognition, but also enhances the robustness of the system under noise and interference conditions, providing an important research direction in the field of signal processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic diagram of the model structure of the multimodal signal recognition of the present invention.
[0038] Figure 2 It is a schematic flow chart of the present invention. DETAILED DESCRIPTION
[0039] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0040] Principle / creativity of the technical solution of the present invention: IQ vector, FFT vector and AP vector are three important signal representation methods, which provide rich information from the time domain, frequency domain and phase characteristics respectively. The IQ vector can effectively describe the modulation method of the signal, the FFT vector reveals the spectral characteristics of the signal, and the AP vector integrates the changes in amplitude and phase. By fusing the IQ, FFT and AP vectors, the advantages of different features can be combined to improve the accuracy and robustness of signal recognition. By fusing the IQ, FFT and AP vectors, the advantages of different features can be combined to improve the accuracy and robustness of signal recognition.
[0041] To improve multimodal signal recognition, the additive attention mechanism has been introduced. This mechanism dynamically adjusts the focus on different modal features, enhancing key features and thus better adapting to signal changes in complex environments. This approach not only improves recognition accuracy but also enhances the system's robustness to noise and interference. Therefore, multimodal signal recognition technology using the additive attention mechanism is becoming an important research direction in the field of signal processing.
[0042] Therefore, the technical solution of this application combines the three modalities of IQ, FFT and AP, and uses the additive attention mechanism to identify and classify signals.
[0043] like Figure 1-2FIG. 1 is a specific embodiment of a multimodal signal recognition method based on an additive attention mechanism. The embodiment is a multimodal signal recognition method based on an additive attention mechanism, and the method includes the following steps:
[0044] Step 1: Signal data input: Data input, preprocessing of data set;
[0045] Select the RML2016.10A dataset and perform preprocessing on it, converting it into vectors of three modes: AP, FFT, and IQ.
[0046] The conversion formula is as follows:
[0047]
[0048] Where A(t) represents the instantaneous amplitude of the IQ signal. The instantaneous amplitude (envelope) of the signal is calculated by taking the square root of the I and Q paths. I(t) and Q(t) represent the real and imaginary parts of the signal in the time domain, respectively.
[0049]
[0050] Where P(t) represents the instantaneous phase of the signal, which is obtained by taking the inverse tangent of the ratio of the I and Q paths.
[0051] X[k]=Re(X[k])+j·Im(X[k]);
[0052] Where X[k] is the frequency domain representation of the frequency domain signal, which is the frequency domain complex signal obtained after fast Fourier transform, R[·] and I[·] represent the real and imaginary parts of the signal respectively;
[0053]
[0054] Where Re(X[k]) is the real number representation of FFT, which is obtained by summing the products of the time domain signal and the cosine basis function x[n]; N is the number of FFT points, n is the time domain index, and k is the frequency domain index;
[0055]
[0056] Where Im(X[k]) is the real number representation of FFT, which is obtained by summing the product of the time domain signal and the sine basis function x[n] and taking the negative; N is the number of FFT points, n is the time domain index, and k is the frequency domain index;
[0057] The multimodal multi-channel feature extraction module operates as follows: Step 2-Step 3:
[0058] Step 2: Design a multi-channel network structure;
[0059] Based on the three modes, IQ is split into IQ channel, I channel and Q channel; AP is split into A channel, P channel and AP channel; FFT vector is divided into real component and imaginary component as input, and the shape of each single channel is 128 × 1. Then perform one-dimensional convolution operation on each single channel respectively;
[0060] Step 3: multimodal fusion processing;
[0061] After convolution extraction of the three signal variables, IQ vector, AP vector, and FFT vector, the extracted features are fused in the first dimension to facilitate subsequent feature extraction. A convolution kernel with 100 channels and 2*5 kernel is used for feature extraction. The final data format is 3*124*100, which is then reshaped into a 372*100 data format for subsequent operations.
[0062] Step 4: Add an LSTM layer (a bidirectional LSTM layer) and an attention mechanism (using the Flatten layer and Dense layer in the traditional attention mechanism) to extract sequence timing features;
[0063] After the data is fused through multiple channels and modules, it is passed through a bidirectional LSTM layer and an attention mechanism module (Flatten layer, Dense layer) to achieve further feature extraction of the data and enhance the signal representation capability.
[0064] Step 5: Signal classification output: output classification;
[0065] The Flatten layer and the Dense layer convert the multidimensional features processed by the attention mechanism into a format suitable for classification. The Flatten layer is responsible for flattening the data, and the Dense layer generates the final output result through a fully connected operation.
[0066] Step 6: Verify the results and compare them;
[0067] In the framework of tensorflow, the training set and validation set are divided into 8:2; then the first 80% of each type of signal-to-noise ratio data in the data set are used for multimodal signal recognition (model training), and the last 20% are used for multimodal signal recognition verification. Figure 1 The designed structure as a whole constitutes the multimodal signal recognition model of this application.
[0068] Layer Kernal Shape IQ_Signal_Input / 2×128×50 Conv1D(I-Channel) 8 128×50 Conv1D(Q-Channel) 8 128×50 Conv2D(IQ-Channel) 2×8 2×128×50 Concat_IQ / 2×128×50 Conv2D(Concat_IQ) 1×8 2×128×100 AP_Signal_Input / 2×128×50 …… …… …… Conv2D(Concat_AP) 1×8 2×128×100 FFT_Signal_Input / 2×128×50 …… …… …… Conv2D(Concat_FFT) 1*8 2×128×100 Concat_IQ_AO_FFT / 6×128×100 Conv2D 2×5 3×124×100 Reshape / 372×100 LSTM / 372×128 Addictive Attention 372×100 Flatten / 37200 Dense / 128
[0069] In a preferred embodiment, the specific operation in step 4 is as follows: the fused data result is sent to a bidirectional LSTM layer after a one-dimensional convolution, and a fully connected layer is used to process the features to generate a weight distribution for each time step, which is then spliced with the original data. Through splicing, the multimodal signal recognition (model) can simultaneously utilize the original information and the important features adjusted according to the attention mechanism.
[0070] In a preferred embodiment, step 5: the multimodal signal recognition (model) training optimizer adopts the Adam algorithm, the epoch is set to 100, the batch size is set to 64, and the accuracy index and the confusion matrix visualization image are used to evaluate and judge the classification results of each signal under different signal-to-noise ratios to achieve the classification accuracy of the evaluation signal.
[0071] In a preferred embodiment, in step 2: the I and Q split channels are convolved with a convolution kernel of 8 and a channel number of 50, and then the convolution results are merged into two channels for a two-dimensional convolution operation with a channel number of 50 and a convolution kernel of 1*8. Then, the main channel data IQ vector is convolved with a 2*8 convolution kernel and a channel number of 50. Finally, feature fusion is performed with the features extracted by the above single channel, and feature extraction is performed using data with a channel number of 50 and a convolution kernel of 2*8, thereby capturing representation features at different scales; the data under the other two modes are extracted in the same way as under the IQ mode.
[0072] In a preferred embodiment, the dataset RadioML 2016.10A includes 11 modulation signal types, including eight digital signals and three analog signals, with a signal-to-noise ratio range of -18 to 20 dB, a sampling rate of 200 kHz, and Gaussian white noise.
[0073] RadioML 2016.10A covers variables such as different signal-to-noise ratios and timestamps. The total number of samples is 220,000, the signal-to-noise ratio range is -18 to 20dB, the sampling rate is 200kHz, and the noise is Gaussian white noise. These features have important reference value for deep learning recognition and classification tasks in signal processing. Therefore, this application selected the open source dataset RML2016.10A as the verification dataset. Signal preprocessing is to meet the needs of analysis and processing, and the original signal is subjected to a series of transformations and modifications to improve the quality and reliability of the signal. Most signal modulation methods process data information of a single mode and ignore the correlation between the multi-modalities of the signal. This application studies the representation methods of different modes of the signal, namely amplitude, phase and frequency domain features, extracts features from multi-modalities to form interactions, and improves the recognition performance of the model.
[0074] Finally, it should be noted that the present invention is not limited to the above embodiments and may be subject to many variations. All variations that can be directly derived or imagined by a person skilled in the art from the disclosure of the present invention should be considered within the scope of protection of the present invention.
Claims
1. A multimodal signal recognition method based on additive attention mechanism, characterized by The method comprises the following steps: Step 1: Preprocess the data set; Select the RML2016.10A dataset and perform preprocessing on it, converting it into vectors of three modes: AP, FFT, and IQ. The conversion formula is as follows: Where A(t) represents the instantaneous amplitude of the IQ signal. The instantaneous amplitude (envelope) of the signal is calculated by taking the square root of the I and Q paths. I(t) and Q(t) represent the real and imaginary parts of the signal in the time domain, respectively. Where P(t) represents the instantaneous phase of the signal, which is obtained by taking the inverse tangent of the ratio of the I and Q paths. X[k]=Re(X[k])+j·Im(X[k]); Where X[k] is the frequency domain representation of the frequency domain signal, which is the frequency domain complex signal obtained after fast Fourier transform, R[·] and I[·] represent the real and imaginary parts of the signal respectively; Where Re(X[k]) is the real number representation of FFT, which is obtained by summing the products of the time domain signal and the cosine basis function x[n]; N is the number of FFT points, n is the time domain index, and k is the frequency domain index; Where Im(X[k]) is the real number representation of FFT, which is obtained by summing the product of the time domain signal and the sine basis function x[n] and taking the negative; N is the number of FFT points, n is the time domain index, and k is the frequency domain index; Step 2: Design a multi-channel network structure; Based on the three modes, IQ is split into IQ channel, I channel and Q channel; AP is split into A channel, P channel and AP channel; FFT vector is divided into real component and imaginary component as input, and the shape of each single channel is 128 × 1. Then perform one-dimensional convolution operation on each single channel respectively; Step 3: multimodal fusion processing; After convolution extraction of the three signal variables, IQ vector, AP vector, and FFT vector, the extracted features are fused in the first dimension to facilitate subsequent feature extraction. A convolution kernel with 100 channels and 2*5 kernel is used for feature extraction. The final data format is 3*124*100, which is then reshaped into a 372*100 data format for subsequent operations. Step 4: Add LSTM layer and attention mechanism to extract sequence timing features; After the data is fused by multiple channels and modules, it is passed through a bidirectional LSTM layer and an attention mechanism module to further extract features of the data and enhance the signal representation capability. Step 5: Output classification; The Flatten layer and the Dense layer convert the multidimensional features processed by the attention mechanism into a format suitable for classification. The Flatten layer is responsible for flattening the data, and the Dense layer generates the final output result through a fully connected operation. Step 6: Verify the results and compare them; In the framework of tensorflow, the training set and validation set are divided into 8:2; then the first 80% of each type of signal-to-noise ratio data in the dataset are trained for multimodal signal recognition, and the remaining 20% are verified for multimodal signal recognition.
2. The multimodal signal recognition method based on additive attention mechanism according to claim 1, characterized in that: The specific operation in step 4 is as follows: the fused data results are sent to the bidirectional LSTM layer after a one-dimensional convolution, and the features are processed by a fully connected layer to generate the weight distribution of each time step, which is then spliced with the original data. Through splicing, multimodal signal recognition can simultaneously utilize the original information and the important features adjusted according to the attention mechanism.
3. The multimodal signal recognition method based on additive attention mechanism according to claim 1 or 2, characterized in that: Step 5: The multimodal signal recognition training optimizer uses the Adam algorithm, with the epoch set to 100 and the batch size set to 64. The accuracy index and the confusion matrix visualization image are used to evaluate the classification results of each signal under different signal-to-noise ratios to achieve the classification accuracy of the evaluation signal.
4. The multimodal signal recognition method based on additive attention mechanism according to claim 1, characterized in that: In step 2: the I and Q split channels are convolved with a convolution kernel of 8 and a channel number of 50. The convolution results are then merged into two channels for a two-dimensional convolution operation with a channel number of 50 and a convolution kernel of 1*8. The main channel data IQ vector is then convolved with a convolution kernel of 2*8 and a channel number of 50. Finally, feature fusion is performed with the features extracted from the single channel above, and feature extraction is performed using data with a channel number of 50 and a convolution kernel of 2*8, thereby capturing representation features at different scales. The data under the other two modes are extracted in the same way as under the IQ mode.
5. The multimodal signal recognition method based on additive attention mechanism according to claim 1, characterized in that: The RadioML 2016.10A dataset contains 11 modulation signal types, including eight digital signals and three analog signals. The signal-to-noise ratio range is -18 to 20 dB, the sampling rate is 200 kHz, and the noise is Gaussian white noise.
Citation Information
Cited By
Multi-modal feature fusion modulation identification method based on signal-to-noise ratio guidance
CN121009352A